Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Gene Expression Atlas update--a value-added database of microarray and sequencing-based functional genomics experiments.

Nucleic Acids Res · 2011
L1 83/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
83/100
Reproducibility score
0.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 61% of all assessed papers rank 430 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Database/resource paper (NAR Database Issue). Described well enough to reproduce the one shipped dataset, but its headline figures are live-DB content snapshots (release 11.08, Aug 2011) that are not reproducible in principle (retired Oracle+Java web app; no data dump) -> out of scope. The foregrounded public dataset E-MTAB-62 (Lukk et al.) was profiled and verified EXACTLY: 5372 samples (CELs=SDRF rows=matrix cols), 369 cell/tissue groups, U133A (22283 probesets), and a delivered 22283x5372 RMA-normalized matrix -> delivers what it promises (quality A). The described limma/t-test per-factor DE methodology was demonstrated on the matrix and yields biologically correct results (HBB/CD3D/GFAP markers correctly enriched), but the paper reports no numeric DE value to match, so that piece is graded partial (methodology, not value). No fabrication indicated. Verdict: partial; run healthy.

💻 Code ↗ 🗄 Data: E-MTAB-62

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Core claims
  • Gene Expression Atlas is an added-value database providing curated, re-annotated and statistically analysed gene expression data across cell types, organism parts, developmental stages, disease states and other biological/experimental conditions, derived from ArrayExpress Archive and the European Nucleotide Archive. resource
  • As of Release 11.08 (August 2011) the Atlas supports 19 species with expression data for 19 014 biological conditions in 136 551 assays from 5598 independent studies, a nearly 6-fold data volume increase since launch. resource
  • The ArrayExpressHTS R-based pipeline was integrated into the Atlas so that new RNA-seq submissions to ArrayExpress and ENA are processed automatically, using bowtie for read alignment and cufflinks for transcript isoform quantification. method
  • MicroRNA chip designs are re-annotated by exact sequence matching of all probes to the latest miRBase version, replacing manufacturer probe IDs with miRBase identifiers to maximize cross-platform comparability. method
  • Zooma automates mapping of submitted sample/assay annotations (type–value pairs) to Experimental Factor Ontology classes by searching prior Atlas mappings, exact EFO text matches and BioPortal/OLS via OntoCAT, using a ranking-based heuristic and writing new mappings before each release. method
  • The Atlas statistical engine uses the Bioconductor package limma with structured sample annotation curation to compute per-factor contrasts, and now reports t-statistics together with P-values and sample sizes on experiment pages and via API. method
  • The Atlas is available as standalone installable software (Oracle database required) released monthly with all content and source code free and unrestricted, plus a Distributed Atlas version federating queries across multiple Atlas servers with 'First Not Null' and 'Aggregation' (e.g. Fisher's method) conflict-resolution rules. resource
  • New user-interface features include searching for non-differentially expressed genes, filtering results by minimum number of experiments, compact EFO ontology trees, Ensembl genome browser links to BAM alignments, anatomograms with Vertebrate Bridging Ontology cross-species homology expansion, and GeneSigDB v4 gene-signature search. resource
Experimental setups
Assay System Perturbation Readout Platform
Microarray gene expression profiling (curated/re-analysed data sets in database) 19 species including human and model organisms; tissues, cell types, cell lines, disease states, developmental stages other (varied biological/experimental conditions from source studies) Differential expression per gene per experimental condition: P-values, t-statistics, sample sizes; up/down/non-differential expression calls limma (Bioconductor); Ensembl BioMart probe-to-genome mappings; Affymetrix and Illumina array designs
RNA-seq (next-generation sequencing transcriptome profiling) Human, mouse and fruit fly samples from ArrayExpress Archive and European Nucleotide Archive other (conditions from submitted studies) Transcript isoform expression levels; short-read alignments stored as BAM files ArrayExpressHTS R pipeline using bowtie (alignment) and cufflinks (isoform quantification)
MicroRNA expression microarray Multiple microRNA chip designs/platforms in the Atlas none MicroRNA expression; probes re-annotated by exact sequence match to miRBase identifiers miRBase (latest version)
Large meta-analysis of microarray gene expression (global map of human gene expression, E-MTAB-62) Human samples (~6000 microarray samples) none Re-processed and re-analysed global human gene expression profiles Affymetrix U133A GeneChip
Ontology annotation mapping (curation assay) Atlas sample/assay annotations across all species none Mapped EFO ontology class per annotation type–value pair (e.g. disease/leukemia → EFO_0000565) Zooma; EFO; BioPortal and OLS web services via OntoCAT
Key results
  • Atlas Release 11.08 (August 2011) contains expression data for 19 014 biological conditions in 136 551 assays from 5598 independent studies across 19 species 19 014 conditions; 136 551 assays; 5598 studies; 19 species
  • As of September 2011 the Atlas contains data for over 370 000 genes from nearly 6000 independent studies, more than 136 000 samples and nearly 20 000 biological conditions >370 000 genes
  • Atlas data volume increased nearly 6-fold compared with the launch report (and >6-fold in the last year and a half) nearly 6-fold
  • Twenty monthly releases were made since the previous report 20 releases
  • Of 45 human, 32 mouse and 64 fly RNA-Seq data sets in ArrayExpress in August 2011, 14 experiments were processed and loaded into the Atlas, with 20 more to be loaded before 2012 14 of 141 loaded; 20 more pending
  • Non-differentially expressed genes are defined as those with multiple-testing-adjusted P-values above the 0.05 significance threshold in simultaneous t-test comparisons with global factor means adjusted P > 0.05
  • The Lukk et al. global map of human gene expression (~6000 Affymetrix U133A samples) was loaded and integrated into the Atlas; a global mouse re-analysis and an order-of-magnitude expanded human Affymetrix data set are to be released in 2012 nearly 6000 samples
  • GeneSigDB v4 signatures were imported, allowing signature-based search (e.g. the van 't Veer breast cancer signature, ID 11823860-Figure2)
Key statistics
  • count 136 551 (Assays in Atlas Release 11.08 (August 2011))
  • count 19 014 (Biological conditions represented in Release 11.08)
  • count 5598 (Independent studies in Release 11.08)
  • count 19 (Species supported by the Atlas)
  • count over 370 000 (Genes with data in the Atlas as of September 2011)
  • fold_change nearly 6-fold (>6-fold in last year and a half) (Growth in Atlas data volume since launch)
  • count 45 human, 32 mouse, 64 fly (RNA-Seq data sets in ArrayExpress in August 2011; 14 processed and loaded into Atlas)
  • pvalue 0.05 (Multiple-testing-adjusted P-value significance threshold for differential expression calls)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a database resource paper describing the Gene Expression Atlas, which aggregates and re-analyzes microarray and RNA-seq data from thousands of independent studies. Differential expression is computed automatically for each gene against each curated experimental condition using the Bioconductor limma package to generate per-factor contrasts, with t-statistics, P-values and sample sizes reported on experiment pages and via the API. A federated 'Distributed Atlas' version additionally combines significance results across multiple Atlas server instances using meta-analytical strategies such as Fisher's method.

Replicationunclear Sample sizeSample sizes are described as being reported together with t-statistics and P-values on experiment pages and via the API, but specific values are not given in the text itself. Groupsgene expression levels across curated experimental factor values/biological conditions (e.g., disease states, tissues, developmental stages) within each independent study Pairingunclear Randomization/blindingna Dispersionunclear Exact p-valuesyes Multiplicity correctionyes
Statistical tests used
Test Applied to n Assumptions
limma moderated t-statistic (per-factor contrast) differential expression of each gene against each curated experimental condition/factor across studies sample sizes are stated to be reported alongside the statistics on experiment pages, but no specific n is given in the text not stated
simultaneous t-test comparisons against global factor means identification of 'non-differentially expressed' genes (multiple testing-adjusted P > 0.05) not stated
Fisher's method (meta-analytic combination) aggregating statistical significance results across multiple federated Atlas server instances in the Distributed Atlas not stated
Approaches that could also have been used
  • Differential expression across the database is computed uniformly with limma's moderated t-statistic, a method developed for microarray intensity data.
    Could also: For RNA-seq count data (processed via the ArrayExpressHTS/cufflinks pipeline), a count-based model such as DESeq2 or edgeR (negative binomial Wald or likelihood-ratio tests) could also be used. — Negative binomial models are designed for the discrete, over-dispersed nature of sequencing read counts, which can differ from the continuous-intensity assumptions underlying microarray-oriented moderated t-statistics.
  • The text notes that P-values used to define 'non-differentially expressed' genes are multiple testing-adjusted but does not name the specific correction method.
    Could also: Explicitly reporting the method, such as Benjamini-Hochberg FDR or Bonferroni correction, could also be provided. — Naming the specific procedure lets readers know the exact false-discovery or family-wise error control being applied, which can vary in stringency and interpretation.
  • Non-differential expression is defined using a non-significant result from a simultaneous t-test against the global factor mean.
    Could also: A formal equivalence-testing approach, such as two one-sided tests (TOST), could also be used to establish a lack of meaningful difference. — Equivalence testing is designed specifically to demonstrate that an effect falls within a pre-specified negligible range, whereas a non-significant t-test alone does not formally establish equivalence.
  • In the Distributed Atlas, statistical significance results from multiple independent Atlas servers are combined using Fisher's method.
    Could also: A random-effects meta-analytic model (e.g., DerSimonian-Laird) could also be used to combine results across servers/studies. — Random-effects models explicitly account for between-study heterogeneity (e.g., differing platforms, populations, or conditions), which can complement a fixed-effect combination approach like Fisher's method.
  • The paper notes the same core statistical approach (limma-based per-factor contrasts) is applied uniformly across increasingly complex, multifactorial experiments.
    Could also: A linear mixed-effects model or full factorial ANOVA framework explicitly modeling multiple factors and their interactions could also be used for such designs. — Explicit multifactorial models can capture interaction effects and shared variance structure across factors, which a series of per-factor contrasts may not directly represent; the authors themselves describe developing such alternative methods for future releases.
Software: R/Bioconductor limma · ArrayExpressHTS (R-based RNA-seq pipeline, using bowtie and cufflinks)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

E62-N-samples
Reported
5372 samples
Reproduced
5372 (CELs=SDRF rows=matrix cols)
exact
E62-N-groups
Reported
369 cell/tissue-type groups
Reproduced
369 distinct Factor Value[369 groups]
exact
E62-platform
Reported
Affymetrix U133A
Reproduced
22283 probesets
exact
E62-norm-matrix
Reported
normalized matrix of 5372 samples
Reproduced
22283 x 5372 RMA matrix delivered
exact
DE-method
Reported
limma per-factor contrasts, adj P<0.05
Reproduced
per-factor t-test contrasts reproduce expected marker biology; adj P<0.05 (BH); no reported numeric target to match
partial
DB-content-137k-assays-etc
Reported
136551 assays / 19014 conditions / 5598 studies / 19 species
Reproduced
not reproducible (live DB snapshot release 11.08; resource retired)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 83/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

What deviates: nothing numerically. The one shipped dataset, E-MTAB-62, reproduces 1:1 — 5372 samples confirmed by three independent counts (CELs = SDRF rows = matrix columns), 369 cell/tissue groups, Affymetrix U133A with 22283 probesets, and a delivered 22283 x 5372 RMA matrix. The described DE methodology re-runs correctly and yields the expected marker biology (HBB/HBA1 ~13.2 in blood vs ~6.5 in solid-tissue cell lines; CD3D in blood; GFAP/NEFL in brain), though the paper reports no numeric DE target, so that piece is a method reproduction only (and used Welch instead of limma's moderated t-test). Whose side: the unverifiable part — the headline totals of 136551 assays / 19014 conditions / 5598 studies / 19 species — is a live-database snapshot at release 11.08 of a service that has since been retired; the authors release-tagged it transparently, so this is a property of resource papers and of data availability, not an authors' defect and certainly not a fabrication signal. Severity: low. The verifiable core is exact and quality-A; the limitation is coverage, which is why q7 is capped at limited and q8 at solid with explainable deviations rather than green.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.