Expression Atlas update--a database of gene and transcript expression from microarray- and sequencing-based functional genomics experiments.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Partial reproduction. This is a database-resource paper, not a single-numeric-result paper; scoped to the RNA-seq baseline pipeline on E-MTAB-513 (Illumina Body Map, 5/16 tissues sampled). Original TopHat1+Cufflinks1 pipeline is defunct, substituted with salmon (documented). Genome-wide TPM correlation vs Expression Atlas's own current values: log1p Pearson r=0.89-0.91 across 5 tissues (within-tol). All 7 tissue-marker-gene spot checks show exactly correct tissue-specificity direction (exact) but absolute TPM magnitude differs by a consistent, explainable 0.23-1.42x factor per gene (partial) -- attributable to quantifier choice, missing 2nd technical replicate, and no upstream QC trimming, not to fabrication. E-MTAB-513 dataset structure (16 tissues + mixture, 48 ENA runs) exactly matches ArrayExpress metadata. Not attempted: E-GEOD-26284 pipeline run (metadata-checked only), whole-database census claims (214 total experiments etc. -- not a pipeline result), remaining 11 tissues, 2nd technical replicates, differential-expression pipeline, microarray pipeline.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-29
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator headless) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opus- ★ Expression Atlas is a value-added database providing gene, protein and splice variant expression across cell types, organism parts, developmental stages, diseases and other biological/experimental conditions, built from manually curated high-quality microarray and RNA-sequencing experiments from ArrayExpress. resource
- ★ The new version introduces the concept of 'baseline' expression, i.e. gene and splice variant abundance levels in healthy or untreated conditions such as tissues or cell types. resource
- ★ Differential expression data are curated in-depth for experimental intent, yielding biologically meaningful 'contrasts' — pairwise comparisons between a 'reference' and a 'test' set of biological replicates. method
- ★ All experiments undergo strict quality control of raw data and experimental design; a minimum of three biological replicates is enforced, poor-quality/contaminated RNA-seq reads and outlier microarrays are removed. method
- ★ Sample attributes and experimental factors are systematized and mapped to Experimental Factor Ontology (EFO) terms. method
- ★ A standardized RNA-seq pipeline (FASTQC QC, contamination removal, TopHat 1 mapping, Cufflinks 1 for baseline gene/transcript quantification, HTSeq counts, DESeq differential analysis) supports reproducible analysis, with analysis methods, versions, Ensembl and miRBase releases listed per experiment. method
- ★ A more powerful search interface allows querying by gene/protein/splice-variant attributes, gene sets (e.g. REACTOME pathways, GO, InterPro), keywords, biotypes and sample attributes/experimental factors, with default ranking by condition-specific expression. resource
- Quality-driven exclusion of low-quality experiments and ongoing manual contrast curation have caused a temporary reduction in the number of experiments in Expression Atlas. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RNA-sequencing (baseline expression quantification) | human and other species tissues/cell types/cell lines (nine species; e.g. Illumina Body Map E-MTAB-513, ENCODE Cell Lines E-GEOD-26284) | none (healthy or untreated baseline conditions) | FPKM expression levels per gene and splice variant, summarized as median across first technical then biological replicates per condition | FASTQC quality control; TopHat 1 mapping to Ensembl reference genome; Cufflinks 1 quantification |
| RNA-sequencing (differential expression analysis) | multiple organisms/tissues/cell lines in ArrayExpress-derived experiments (13 species across differential experiments) | contrast-defined (e.g. test compound treatment vs untreated, mutant vs wild type, diseased vs healthy) | raw read counts; P-values, log2-fold changes at default FDR 0.05 | HTSeq for counting; DESeq for differential expression |
| Single-channel (one-colour) microarray, mainly gene arrays | multiple organisms/tissues/cell types/cell lines (e.g. Drosophila melanogaster CDK8 and Cyclin C homozygous mutants) | contrast-defined (e.g. homozygous mutant vs wild type) | normalized expression values; P-values, t-statistics, log2-fold changes; MA plots at FDR 0.05 | e.g. Affymetrix GeneChip Drosophila Genome 2.0 Array |
| Two-colour microarray | ArrayExpress-derived differential experiments across multiple species | contrast-defined (test vs reference biological replicate sets) | log2-ratios; P-values, t-statistics, log2-fold changes | — |
| MicroRNA microarray | ArrayExpress-derived experiments (e.g. E-TABM-713) | contrast-defined | differential microRNA expression; probe-set to microRNA mappings | miRBase release used for probe-set to microRNA mapping |
| Mass-spectrometry proteomics (planned integration, baseline component only) | PRIDE database datasets with EFO-matched sample descriptions | none (baseline) | protein expression shown in context of baseline expression of the coding gene per condition | PRIDE database |
- – As of 24 September 2013, Expression Atlas contains highly curated data from 214 experiments 214 experiments
- – Four baseline RNA-sequencing experiments covering nine species are included 4 experiments, 9 species
- – 210 differential experiments covering 13 species are included 210 experiments, 13 species
- – Differential experiments comprise 10 RNA-sequencing and 200 microarray experiments 10 RNA-seq / 200 microarray
- – Baseline expression search applies a default FPKM cut-off below which expression is treated as background noise; user-adjustable with a histogram of genes expressed above a given cut-off FPKM 0.5 default
- – Differential expression results are called at a default false discovery rate, shown in MA plots (differentially expressed genes in red), with user-selectable FDR FDR 0.05
- – A minimum of three biological sample replicates is enforced to ensure sufficient statistical power to detect differential expression 3 replicates
- – Clicking a non-empty baseline heatmap cell shows a breakdown of the three most abundant splice variants for that gene and condition 3 splice variants
- count 214 (total curated experiments in Expression Atlas as of 24 September 2013)
- count 210 (differential experiments (13 species))
- count 200 (differential microarray experiments, mainly single-channel gene arrays)
- count 10 (differential RNA-sequencing experiments)
- count 4 (baseline RNA-sequencing experiments (nine species))
- other FDR of 0.05 (default false discovery rate for calling differential expression / MA plots)
- other 0.5 FPKM (default baseline expression cut-off; below this treated as background noise)
- count three (minimum acceptable number of biological replicates enforced per experiment)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a database/resource paper (Expression Atlas), not a primary experimental study; differential expression is computed for curated pairwise 'contrasts' (a 'test' set vs a 'reference' set of biological replicates) using standardized pipelines that differ by platform. RNA-sequencing data are quantified with HTSeq and analyzed for differential expression with DESeq, while microarray data are summarized with P-values and t-statistics; results are reported as P-values, t-statistics (microarray only), and log2 fold-changes, and filtered in the interface using a default FDR threshold of 0.05. A minimum of three biological replicates per condition is enforced before a contrast is analyzed.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| DESeq differential expression test (negative-binomial based) on HTSeq counts | RNA-sequencing differential expression contrasts (test vs reference biological replicate sets) | minimum three biological replicates per condition (as enforced quality-control rule) | not stated |
| t-statistic-based differential expression analysis | microarray differential expression contrasts (test vs reference biological replicate sets) | minimum three biological replicates per condition (as enforced quality-control rule) | not stated |
-
Microarray contrasts are summarized with P-values and t-statistics without naming the specific test framework.↳ Could also: An explicit empirical-Bayes moderated t-test (e.g. limma), which is widely used for microarray contrasts — Moderated t-statistics borrow information across genes to stabilize variance estimates, which can be particularly useful when replicate numbers are small (as low as three here).
-
RNA-seq differential expression is computed with DESeq on HTSeq counts.↳ Could also: Newer count-based tools such as DESeq2 or edgeR, which use updated dispersion-shrinkage and independent filtering approaches — These methods can refine variance estimation for lowly expressed genes and are commonly used successors to the original DESeq method.
-
Multiple testing across genes within a contrast is controlled via an FDR threshold (default 0.05), with the specific correction procedure not named.↳ Could also: Explicitly reporting the correction method (e.g. Benjamini-Hochberg) alongside the threshold — Naming the exact procedure makes the multiplicity control fully transparent and reproducible from the text alone.
-
A minimum of three biological replicates per condition is enforced as a general quality rule to support statistical power.↳ Could also: A formal power analysis tailored to an expected effect size and desired detection sensitivity, performed per experiment — A quantitative power calculation could indicate whether three replicates are sufficient for a specific effect size of interest, rather than relying on a fixed minimum across all experiments.
-
Baseline expression (FPKM) is summarized as a median across technical then biological replicates, without an accompanying variability statistic.↳ Could also: Reporting a dispersion measure (e.g. IQR, SD, or range) alongside the median — A dispersion measure would let users gauge how consistent expression levels are across the replicates being summarized.
-
Each differential experiment is broken into pairwise (two-group) contrasts between a reference and a test set.↳ Could also: A single multi-group model (e.g. a likelihood-ratio test framework as in DESeq2, or ANOVA-like designs) when more than two conditions are studied — A joint multi-condition test can capture overall condition effects in one combined analysis rather than requiring multiple separate pairwise contrasts.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
What deviates: absolute gene-level TPM magnitudes. On identical raw ENA reads, salmon 1.10.1 recovers Expression Atlas's own current TPM table at log1p Pearson r=0.892–0.910 across ~33,800 genes in all 5 sampled tissues, and reproduces 7/7 tissue-marker enrichments in the correct tissue — but per-gene absolute values run 0.23–1.42x off (ACTA1 1689 vs 7410; GFAP 2368 vs 1667).
Whose side: ours, and expectedly so. The paper's 2013 TopHat1+Cufflinks1 pipeline is defunct and was substituted; we also skipped Atlas's adapter/QC filtering, used 1 of 2 technical replicates, applied no quantile normalization, and used Ensembl ~111 against Atlas's release 95. Nothing points at the authors, and nothing is non-derivable.
Severity: moderate. Direction, ranking and biology reproduce exactly; only the scale factor moves. Note the endpoint caveat — the comparison was against Atlas's present-day re-processed table, not against any number printed in the 2013 paper, which for a database-resource paper (headline claims are a content census: 214 experiments, 9/13 species) is the only comparable target available. Coverage is partial (5/16 tissues, no E-GEOD-26284, no differential/microarray pipeline), which keeps the overall grade at yellow rather than green.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.