Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Easy and efficient ensemble gene set testing with EGSEA.

F1000Res · 2017
L1 97/100 PQI 95
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
97/100
Reproducibility score
1.3 SD above mean
vs. all fields · 1187 studies
🎯 Scores higher than 92% of all assessed papers rank 80 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough and reproduces 1:1. EGSEA is a Bioconductor workflow article; I ran its published pipeline (EGSEA123 commit cd121b40, EGSEA/EGSEAdata/limma/edgeR) on the paper's own data (GSE63310 RNA-seq + GSE19446 microarray) on «our HPC» SLURM. All 8 deterministic anchors are byte-exact: RNA-seq DGEList dim 14165x9 (C1), all five indexed collection sizes c2/c5BP/c5CC/c5MF/kegg = 4726/4653/584/928/287 (C2), and microarray probe count 7123 (C6) - and show(gsa) independently re-prints MSigDB v5.2 (07-Mar-2017), the same snapshot the 2017 paper used. The ensemble gene-set rankings (C3/C4/C5), which involve permutation-based base methods (safe, padog) run without a seed in the vignette, still reproduce essentially exactly: KEGG top pathways for BasalvsLP (top-5) and comparative (top-4) are identical and in the same order (C4); the LIM ranking table med.ranks (36,37,92,134,180) and vote.ranks (all 5) match exactly with 2 of 5 p.adj values byte-identical (C5); the c2 top-10 contains the same 5 LIM sets the paper highlights with only a single position swap and one 10th-element difference (C3). NOT attempted (the figure/visual ~20%): heatmaps, KEGG pathview maps, MDS/summary/GO-graph plots and the HTML report (no pinnable numbers), the full microarray EGSEA ranking (only its deterministic 7123 anchor reproduced), and byte-exact p.adj for unseeded stochastic methods. No fabrication concern: every reported value is regenerated from the shipped data + published code.

💻 Code ↗ 🗄 Data: GSE63310

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 97
    assessed: 2026-06-14 ⛓ b7bb422b5909
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Given the proliferation of gene set enrichment (GSE) analysis methods in Bioconductor with no clear best choice, can combining results from multiple GSE algorithms into an ensemble (EGSEA) produce a more robust consensus ranking of biologically relevant gene sets than any single method?

Core claims
  • EGSEA combines results from up to 12 prominent gene set testing algorithms to obtain a consensus ranking of biologically relevant gene sets method
  • EGSEA calculates seven statistics to combine individual gene set statistics from base GSE methods for ranking method
  • The EGSEAdata package provides access to more than 25,000 gene sets for human and mouse organized by database source (MSigDB, KEGG, GeneSetDB) resource
  • Simulation showed that including more base methods in the ensemble gives better performance finding
  • Sorting EGSEA results by med.rank or vote.rank tends to produce fewer false positives than the default p.adj finding
  • Fry is a fast approximation of roast that produces p-values similar to roast run with infinite rotations method
  • Bioconductor hosts over 80 software packages capable of gene set analysis finding
  • EGSEA can extend limma-based differential expression analyses and integrate with existing gene expression pipelines for human and mouse RNA-seq/microarray data resource
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq mouse mammary gland, 3 sorted cell populations (Basal, Luminal Progenitor, Mature Luminal), virgin female mice none (comparison across sorted cell populations) gene-level counts / differential expression between populations Illumina HiSeq 2000, 100bp single-end; Rsubread alignment to mm10; featureCounts
microarray gene expression profiling mouse mammary gland, 3 sorted cell populations (Basal/MaSC-enriched, LP, ML) none (comparison across sorted cell populations) hybridization intensity / differential expression between populations Illumina MouseWG-6 v2 BeadChip
Key results
  • MSigDB c2 curated gene set collection indexed for mouse RNA-seq data contained 4726 gene sets (4729 originally loaded) 4726 gene sets
  • MSigDB c5 GO collection was partitioned into c5BP, c5CC and c5MF sub-collections for mouse
  • Example gene set SMID_BREAST_CANCER_LUMINAL_A_DN retrieved from the indexed c2 collection, annotated as genes down-regulated in luminal A breast cancer, with 23/24 genes present in the data 23/24 genes
  • Eleven of twelve available base GSE methods (camera, safe, gage, padog, plage, zscore, gsva, ssgsea, globaltest, ora, fry) were selected for the ensemble analysis, excluding roast in favor of its faster approximation fry
Key statistics
  • count more than 25,000 gene sets (total gene sets available in EGSEAdata for human and mouse)
  • count 4726 (indexed) / 4729 (loaded) (MSigDB c2 curated gene sets for mouse)
  • count 6166 (MSigDB c5 GO gene sets for mouse)
  • count 50 (MSigDB h Hallmark gene sets for mouse)
  • count Signaling: 132, Metabolism: 89, Disease: 67 (KEGG pathway gene set collections for mouse)
  • count 14165 genes x 9 samples (dimensions of preprocessed RNA-seq count matrix (mam.rnaseq.data))
  • count 3 replicates per population (RNA-seq) (Sheridan et al. 2015 RNA-seq experiment, triplicate samples per of 3 cell populations across 3 batches)
  • count 5 replicates per population (microarray) (Lim et al. 2010 microarray experiment, 5 replicates per cell population on 3 BeadChips)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a software/workflow article demonstrating the EGSEA R/Bioconductor package for ensemble gene set enrichment analysis, rather than a hypothesis-testing study. The statistical approach centres on a limma-based differential expression framework (TMM normalization, voom precision weights, a linear model with a design and contrasts matrix) applied to RNA-seq and microarray data from 3 sorted mammary cell populations, followed by gene set testing that runs up to twelve base GSE algorithms (e.g. camera, fry, gage, padog, globaltest, ora, gsva) and combines their per-set p-values into a consensus ranking. Results are reported as ranked gene sets visualized via heatmaps, pathway/GO graphs, scatter and bar plots and an HTML report rather than as point estimates with dispersion.

Replicationbiological Sample sizeRNA-seq: triplicate RNA samples per population obtained in 3 batches (9 total); microarray: 5 replicates per population hybridised across 3 BeadChips; no formal power/sample-size calculation described Groups3 mammary cell populations (Basal, LP, ML), pairwise Pairingunpaired Randomization/blindingnot stated Dispersionnone Multiplicity correctionadjusted p-values (p.adj) available and used as default sort; combined rank-based options (med.rank, vote.rank) recommended to reduce false positives; specific FDR method not explicitly named in this text
Statistical tests used
Test Applied to n Assumptions
limma linear model / empirical Bayes moderated differential expression (with voom precision weights) differential expression between Basal, LP and ML populations via contrasts (BasalvsLP, BasalvsML, LPvsML) RNA-seq: triplicate samples per population (9 samples); microarray: 5 replicates per population not stated
camera (competitive gene set test) base GSE method in the ensemble; comparisons defined by contrasts matrix competitive null hypothesis stated
fry (rotation gene set test, fast roast approximation) base GSE method in the ensemble self-contained null; assumes equal gene-wise variances across samples (stated)
ora (over-representation, hypergeometric test via phyper on a 2x2 table) base GSE method in the ensemble competitive null hypothesis stated
gage, gsva (competitive); safe, padog, plage, zscore, ssgsea, globaltest (self-contained) base GSE methods combined in the ensemble competitive vs self-contained null hypotheses stated
p-value combination across base methods (Wilkinson default; also fisher, average, logitp, sump, sumz, votep, median) merging individual gene-set p-values into a consensus EGSEA score not stated
Approaches that could also have been used
  • The workflow combines results from many base GSE algorithms into a consensus ranking and notes that including more methods improved performance in simulation.
    Could also: A single well-characterised method (e.g. camera alone, which accounts for inter-gene correlation) could also be applied and reported on its own. — A single method yields a simpler, more directly interpretable inferential statement and avoids the question of how heterogeneous null hypotheses (competitive vs self-contained) are merged; the ensemble in turn adds robustness across method-specific assumptions.
  • Individual gene-set p-values are merged using p-value combination functions (Wilkinson default) from the metap package.
    Could also: Rank aggregation across methods (e.g. the package's own med.rank/vote.rank) or a single FDR-controlled test could also be used as the primary readout. — Rank-based aggregation does not require the combined p-values to retain a calibrated null interpretation and, as the authors note, can produce fewer false positives; a single test keeps a clean type-I error statement.
  • Sorting and significance default to adjusted p-values (p.adj), with the specific multiple-testing method not named in this text.
    Could also: Explicitly reporting Benjamini-Hochberg FDR q-values per collection per contrast would also be standard. — Naming the correction method and reporting q-values makes the family-wise/false-discovery scope of significance calls transparent across the thousands of gene sets tested.
  • The design uses biological replicates (n=3 RNA-seq, n=5 microarray per group) without a described power or sample-size calculation.
    Could also: A brief statement of the basis for the replicate number (or its origin in the source datasets) could also be included. — Documenting the rationale for n helps readers gauge the sensitivity of downstream gene set tests, particularly for the smaller RNA-seq groups.
  • Results are presented as ranked gene sets and visualizations rather than with numeric effect sizes and dispersion.
    Could also: Reporting set-level effect summaries (e.g. avg.logfc) alongside confidence intervals or directional statistics for top sets could also accompany the rankings. — Quantitative effect and uncertainty summaries complement rank-based displays by conveying magnitude and stability of enrichment, not just ordering.
Software: R/Bioconductor EGSEA · EGSEAdata · limma · edgeR · GSVA (plage, zscore, ssgsea) · metap (p-value combination) · MSigDB gene set collections 5.2

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
78
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE19446 GEO in Article (http://semanticscience.org/resource/SIO_001029)
also used by 1 paper:
GO:0061857 Gene Ontology (GO) in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
GSE63310 GEO in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-29333246 (EGSEA workflow, Alhamdoosh et al. 2017, F1000Research)

What the paper is

A Bioconductor workflow article (tutorial). It demonstrates the EGSEA package (Ensemble of Gene Set Enrichment Analyses) by running an end-to-end gene-set testing analysis on two companion datasets of the same 3 mouse mammary cell populations (Basal, LP, ML):

  1. RNA-seq — Sheridan et al. 2015, GEO GSE63310, preprocessed into a DGEList by Law et al. 2016 (RNAseq123). Downloaded as mam.rnaseq.rdata.
  2. Microarray — Lim et al. 2010, GEO GSE19446 (Illumina MouseWG-6 v2), distributed as arraydata.zip (IDAT + BGX).

Code: companion package EGSEA123 (github.com/mritchie/EGSEA123, vignette EGSEAWorkflow.Rmd). The actual computation is done by the EGSEA + EGSEAdata + limma/edgeR Bioconductor packages.

Per BRIEF rule P16: reproducing by running these published Bioconductor tools on the paper's own data is fully valid — and is exactly what the workflow itself does.

In scope (pipeline-derived, reproduced)

All results in this workflow are pipeline-derived (every number is printed R output embedded in the article). We reproduce the clearly-specified, low-stochasticity ones first (80/20):

id result determinism
C1 RNA-seq DGEList dim 14165x9 deterministic
C2a-e indexed collection sizes (c2/c5BP/c5CC/c5MF/kegg) deterministic (EGSEAdata version)
C6 microarray probes retained = 7123 deterministic
C3 top-10 c2 comparative gene sets (LIM sets) ensemble; partly stochastic
C4 top KEGG pathways (BasalvsLP, comparison) ensemble; partly stochastic
C5 LIM gene-set ranking table (Rank/med.rank/vote.rank) ensemble; partly stochastic

Out of scope / not attempted (the hard ~20%)

  • All figure-only outputs (heatmaps, KEGG pathview maps, MDS plots, summary bubble plots, GO graphs, HTML report) — visual, not pinnable numbers.
  • The microarray EGSEA ensemble run (only its deterministic 7123-probe preprocessing anchor is reproduced; the full microarray gene-set ranking is the RNA-seq analysis repeated and adds little marginal evidence).
  • Exact med.rank values for stochastic methods: EGSEA combines 11 base methods of which several (safe, padog, and permutation-based steps) draw random permutations and the vignette sets no seed, so byte-exact med.rank/p.adj are not expected. We grade the ranking structure (which sets land in the top-10, vote.rank) rather than chasing identical floats.

Pipelines named per result

  • C1: limma/edgeR calcNormFactors+voom on the shipped DGEList.
  • C2: EGSEA::buildIdx over EGSEAdata MSigDB c2/c5 + KEGG.
  • C3-C5: EGSEA::egsea ensemble (camera, safe, gage, padog, plage, zscore, gsva, ssgsea, globaltest, ora, fry), sort.by="med.rank", then topSets.
  • C6: limma read.idat/neqc/detection-p filter + Entrez dedup.

Environment (run-time resolved)

«our HPC» SLURM, conda prefix env on «infra»: R 4.5.3, limma 3.66.0, edgeR 4.8.2, EGSEA 1.38.0, EGSEAdata 1.38.0 (newer than the 2017 paper's EGSEA>=1.5.2, but the EGSEAdata gene-set snapshot is unchanged — see exact C2 match).

Figures / tables: table
C1
Reported
14165 x 9
Reproduced
14165 x 9
exact
C2a
Reported
c2 = 4726 gene sets
Reproduced
4726
exact
C2b
Reported
c5BP = 4653
Reproduced
4653
exact
C2c
Reported
c5CC = 584
Reproduced
584
exact
C2d
Reported
c5MF = 928
Reproduced
928
exact
C2e
Reported
kegg = 287
Reproduced
287
exact
C3
Reported
5 LIM sets in top-10 c2 comparative (paper sets 1,2,4,6,8)
Reproduced
5 LIM sets in top-10 (pos 1,2,4,5,8); top-10 membership identical except 10th element
within tolerance
C4a
Reported
KEGG BasalvsLP top5: Collecting duct acid secretion, a-Linolenic acid metab, Synaptic vesicle cycle, Hepatitis C, Vascular smooth muscle contraction
Reproduced
identical top-5, same order
exact
C4b
Reported
KEGG comparative top4: Collecting duct acid secretion, Synaptic vesicle cycle, Vascular smooth muscle contraction, Axon guidance
Reproduced
identical top-4, same order
exact
C5
Reported
LIM table Rank 1,2,4,6,8; med.rank 36,37,92,134,180; vote.rank all 5
Reproduced
Rank 1,2,4,5,8; med.rank 36,37,92,134,180 (exact); vote.rank all 5 (exact); 2/5 p.adj byte-identical
within tolerance
C6
Reported
7123 probes
Reproduced
7123
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 97/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

Ran the published EGSEA Bioconductor workflow on the authors' own data (GSE63310 RNA-seq + GSE19446 microarray) with the same EGSEAdata MSigDB v5.2 (07-Mar-2017) snapshot. All 8 deterministic anchors are byte-exact (DGEList 14165x9; collection sizes 4726/4653/584/928/287; 7123 probes), the KEGG top pathways match identically and in order, and the LIM ensemble med.ranks (36,37,92,134,180) and vote.ranks reproduce exactly — with 2/5 p.adj values byte-identical. The only deviations are single-position rank swaps from unseeded stochastic base methods (safe/padog), i.e. expected RNG noise. No fabrication concern — a clean 1:1 reproduction.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

130.6 k
tokens (I/O) · 9.5 M incl. cache
16 min
runtime · 0.21 CPU-h
48.5 GB
peak RAM
2
HPC jobs
hummel
machine