Gene Expression Atlas update--a value-added database of microarray and sequencing-based functional genomics experiments.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡Reported values were only indirectly comparable
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Database/resource paper (NAR Database Issue). Described well enough to reproduce the one shipped dataset, but its headline figures are live-DB content snapshots (release 11.08, Aug 2011) that are not reproducible in principle (retired Oracle+Java web app; no data dump) -> out of scope. The foregrounded public dataset E-MTAB-62 (Lukk et al.) was profiled and verified EXACTLY: 5372 samples (CELs=SDRF rows=matrix cols), 369 cell/tissue groups, U133A (22283 probesets), and a delivered 22283x5372 RMA-normalized matrix -> delivers what it promises (quality A). The described limma/t-test per-factor DE methodology was demonstrated on the matrix and yields biologically correct results (HBB/CD3D/GFAP markers correctly enriched), but the paper reports no numeric DE value to match, so that piece is graded partial (methodology, not value). No fabrication indicated. Verdict: partial; run healthy.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-24
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opus- ★ Gene Expression Atlas is an added-value database providing curated, re-annotated and statistically analysed gene expression data across cell types, organism parts, developmental stages, disease states and other biological/experimental conditions, derived from ArrayExpress Archive and the European Nucleotide Archive. resource
- ★ As of Release 11.08 (August 2011) the Atlas supports 19 species with expression data for 19 014 biological conditions in 136 551 assays from 5598 independent studies, a nearly 6-fold data volume increase since launch. resource
- ★ The ArrayExpressHTS R-based pipeline was integrated into the Atlas so that new RNA-seq submissions to ArrayExpress and ENA are processed automatically, using bowtie for read alignment and cufflinks for transcript isoform quantification. method
- ★ MicroRNA chip designs are re-annotated by exact sequence matching of all probes to the latest miRBase version, replacing manufacturer probe IDs with miRBase identifiers to maximize cross-platform comparability. method
- ★ Zooma automates mapping of submitted sample/assay annotations (type–value pairs) to Experimental Factor Ontology classes by searching prior Atlas mappings, exact EFO text matches and BioPortal/OLS via OntoCAT, using a ranking-based heuristic and writing new mappings before each release. method
- ★ The Atlas statistical engine uses the Bioconductor package limma with structured sample annotation curation to compute per-factor contrasts, and now reports t-statistics together with P-values and sample sizes on experiment pages and via API. method
- ★ The Atlas is available as standalone installable software (Oracle database required) released monthly with all content and source code free and unrestricted, plus a Distributed Atlas version federating queries across multiple Atlas servers with 'First Not Null' and 'Aggregation' (e.g. Fisher's method) conflict-resolution rules. resource
- ★ New user-interface features include searching for non-differentially expressed genes, filtering results by minimum number of experiments, compact EFO ontology trees, Ensembl genome browser links to BAM alignments, anatomograms with Vertebrate Bridging Ontology cross-species homology expansion, and GeneSigDB v4 gene-signature search. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Microarray gene expression profiling (curated/re-analysed data sets in database) | 19 species including human and model organisms; tissues, cell types, cell lines, disease states, developmental stages | other (varied biological/experimental conditions from source studies) | Differential expression per gene per experimental condition: P-values, t-statistics, sample sizes; up/down/non-differential expression calls | limma (Bioconductor); Ensembl BioMart probe-to-genome mappings; Affymetrix and Illumina array designs |
| RNA-seq (next-generation sequencing transcriptome profiling) | Human, mouse and fruit fly samples from ArrayExpress Archive and European Nucleotide Archive | other (conditions from submitted studies) | Transcript isoform expression levels; short-read alignments stored as BAM files | ArrayExpressHTS R pipeline using bowtie (alignment) and cufflinks (isoform quantification) |
| MicroRNA expression microarray | Multiple microRNA chip designs/platforms in the Atlas | none | MicroRNA expression; probes re-annotated by exact sequence match to miRBase identifiers | miRBase (latest version) |
| Large meta-analysis of microarray gene expression (global map of human gene expression, E-MTAB-62) | Human samples (~6000 microarray samples) | none | Re-processed and re-analysed global human gene expression profiles | Affymetrix U133A GeneChip |
| Ontology annotation mapping (curation assay) | Atlas sample/assay annotations across all species | none | Mapped EFO ontology class per annotation type–value pair (e.g. disease/leukemia → EFO_0000565) | Zooma; EFO; BioPortal and OLS web services via OntoCAT |
- – Atlas Release 11.08 (August 2011) contains expression data for 19 014 biological conditions in 136 551 assays from 5598 independent studies across 19 species 19 014 conditions; 136 551 assays; 5598 studies; 19 species
- – As of September 2011 the Atlas contains data for over 370 000 genes from nearly 6000 independent studies, more than 136 000 samples and nearly 20 000 biological conditions >370 000 genes
- ▲ Atlas data volume increased nearly 6-fold compared with the launch report (and >6-fold in the last year and a half) nearly 6-fold
- – Twenty monthly releases were made since the previous report 20 releases
- – Of 45 human, 32 mouse and 64 fly RNA-Seq data sets in ArrayExpress in August 2011, 14 experiments were processed and loaded into the Atlas, with 20 more to be loaded before 2012 14 of 141 loaded; 20 more pending
- – Non-differentially expressed genes are defined as those with multiple-testing-adjusted P-values above the 0.05 significance threshold in simultaneous t-test comparisons with global factor means adjusted P > 0.05
- – The Lukk et al. global map of human gene expression (~6000 Affymetrix U133A samples) was loaded and integrated into the Atlas; a global mouse re-analysis and an order-of-magnitude expanded human Affymetrix data set are to be released in 2012 nearly 6000 samples
- – GeneSigDB v4 signatures were imported, allowing signature-based search (e.g. the van 't Veer breast cancer signature, ID 11823860-Figure2)
- count 136 551 (Assays in Atlas Release 11.08 (August 2011))
- count 19 014 (Biological conditions represented in Release 11.08)
- count 5598 (Independent studies in Release 11.08)
- count 19 (Species supported by the Atlas)
- count over 370 000 (Genes with data in the Atlas as of September 2011)
- fold_change nearly 6-fold (>6-fold in last year and a half) (Growth in Atlas data volume since launch)
- count 45 human, 32 mouse, 64 fly (RNA-Seq data sets in ArrayExpress in August 2011; 14 processed and loaded into Atlas)
- pvalue 0.05 (Multiple-testing-adjusted P-value significance threshold for differential expression calls)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a database resource paper describing the Gene Expression Atlas, which aggregates and re-analyzes microarray and RNA-seq data from thousands of independent studies. Differential expression is computed automatically for each gene against each curated experimental condition using the Bioconductor limma package to generate per-factor contrasts, with t-statistics, P-values and sample sizes reported on experiment pages and via the API. A federated 'Distributed Atlas' version additionally combines significance results across multiple Atlas server instances using meta-analytical strategies such as Fisher's method.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma moderated t-statistic (per-factor contrast) | differential expression of each gene against each curated experimental condition/factor across studies | sample sizes are stated to be reported alongside the statistics on experiment pages, but no specific n is given in the text | not stated |
| simultaneous t-test comparisons against global factor means | identification of 'non-differentially expressed' genes (multiple testing-adjusted P > 0.05) | — | not stated |
| Fisher's method (meta-analytic combination) | aggregating statistical significance results across multiple federated Atlas server instances in the Distributed Atlas | — | not stated |
-
Differential expression across the database is computed uniformly with limma's moderated t-statistic, a method developed for microarray intensity data.↳ Could also: For RNA-seq count data (processed via the ArrayExpressHTS/cufflinks pipeline), a count-based model such as DESeq2 or edgeR (negative binomial Wald or likelihood-ratio tests) could also be used. — Negative binomial models are designed for the discrete, over-dispersed nature of sequencing read counts, which can differ from the continuous-intensity assumptions underlying microarray-oriented moderated t-statistics.
-
The text notes that P-values used to define 'non-differentially expressed' genes are multiple testing-adjusted but does not name the specific correction method.↳ Could also: Explicitly reporting the method, such as Benjamini-Hochberg FDR or Bonferroni correction, could also be provided. — Naming the specific procedure lets readers know the exact false-discovery or family-wise error control being applied, which can vary in stringency and interpretation.
-
Non-differential expression is defined using a non-significant result from a simultaneous t-test against the global factor mean.↳ Could also: A formal equivalence-testing approach, such as two one-sided tests (TOST), could also be used to establish a lack of meaningful difference. — Equivalence testing is designed specifically to demonstrate that an effect falls within a pre-specified negligible range, whereas a non-significant t-test alone does not formally establish equivalence.
-
In the Distributed Atlas, statistical significance results from multiple independent Atlas servers are combined using Fisher's method.↳ Could also: A random-effects meta-analytic model (e.g., DerSimonian-Laird) could also be used to combine results across servers/studies. — Random-effects models explicitly account for between-study heterogeneity (e.g., differing platforms, populations, or conditions), which can complement a fixed-effect combination approach like Fisher's method.
-
The paper notes the same core statistical approach (limma-based per-factor contrasts) is applied uniformly across increasingly complex, multifactorial experiments.↳ Could also: A linear mixed-effects model or full factorial ANOVA framework explicitly modeling multiple factors and their interactions could also be used for such designs. — Explicit multifactorial models can capture interaction effects and shared variance structure across factors, which a series of per-factor contrasts may not directly represent; the authors themselves describe developing such alternative methods for future releases.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
What deviates: nothing numerically. The one shipped dataset, E-MTAB-62, reproduces 1:1 — 5372 samples confirmed by three independent counts (CELs = SDRF rows = matrix columns), 369 cell/tissue groups, Affymetrix U133A with 22283 probesets, and a delivered 22283 x 5372 RMA matrix. The described DE methodology re-runs correctly and yields the expected marker biology (HBB/HBA1 ~13.2 in blood vs ~6.5 in solid-tissue cell lines; CD3D in blood; GFAP/NEFL in brain), though the paper reports no numeric DE target, so that piece is a method reproduction only (and used Welch instead of limma's moderated t-test). Whose side: the unverifiable part — the headline totals of 136551 assays / 19014 conditions / 5598 studies / 19 species — is a live-database snapshot at release 11.08 of a service that has since been retired; the authors release-tagged it transparently, so this is a property of resource papers and of data availability, not an authors' defect and certainly not a fabrication signal. Severity: low. The verifiable core is exact and quality-A; the limitation is coverage, which is why q7 is capped at limited and q8 at solid with explainable deviations rather than green.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.