Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Expression Atlas update--a database of gene and transcript expression from microarray- and sequencing-based functional genomics experiments.

Nucleic Acids Res · 2013
L1 84/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Same input data as the authors
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
84/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 63% of all assessed papers rank 392 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Partial reproduction. This is a database-resource paper, not a single-numeric-result paper; scoped to the RNA-seq baseline pipeline on E-MTAB-513 (Illumina Body Map, 5/16 tissues sampled). Original TopHat1+Cufflinks1 pipeline is defunct, substituted with salmon (documented). Genome-wide TPM correlation vs Expression Atlas's own current values: log1p Pearson r=0.89-0.91 across 5 tissues (within-tol). All 7 tissue-marker-gene spot checks show exactly correct tissue-specificity direction (exact) but absolute TPM magnitude differs by a consistent, explainable 0.23-1.42x factor per gene (partial) -- attributable to quantifier choice, missing 2nd technical replicate, and no upstream QC trimming, not to fabrication. E-MTAB-513 dataset structure (16 tissues + mixture, 48 ENA runs) exactly matches ArrayExpress metadata. Not attempted: E-GEOD-26284 pipeline run (metadata-checked only), whole-database census claims (214 total experiments etc. -- not a pipeline result), remaining 11 tissues, 2nd technical replicates, differential-expression pipeline, microarray pipeline.

💻 Code ↗ 🗄 Data: E-MTAB-513

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-29
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator headless) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Core claims
  • Expression Atlas is a value-added database providing gene, protein and splice variant expression across cell types, organism parts, developmental stages, diseases and other biological/experimental conditions, built from manually curated high-quality microarray and RNA-sequencing experiments from ArrayExpress. resource
  • The new version introduces the concept of 'baseline' expression, i.e. gene and splice variant abundance levels in healthy or untreated conditions such as tissues or cell types. resource
  • Differential expression data are curated in-depth for experimental intent, yielding biologically meaningful 'contrasts' — pairwise comparisons between a 'reference' and a 'test' set of biological replicates. method
  • All experiments undergo strict quality control of raw data and experimental design; a minimum of three biological replicates is enforced, poor-quality/contaminated RNA-seq reads and outlier microarrays are removed. method
  • Sample attributes and experimental factors are systematized and mapped to Experimental Factor Ontology (EFO) terms. method
  • A standardized RNA-seq pipeline (FASTQC QC, contamination removal, TopHat 1 mapping, Cufflinks 1 for baseline gene/transcript quantification, HTSeq counts, DESeq differential analysis) supports reproducible analysis, with analysis methods, versions, Ensembl and miRBase releases listed per experiment. method
  • A more powerful search interface allows querying by gene/protein/splice-variant attributes, gene sets (e.g. REACTOME pathways, GO, InterPro), keywords, biotypes and sample attributes/experimental factors, with default ranking by condition-specific expression. resource
  • Quality-driven exclusion of low-quality experiments and ongoing manual contrast curation have caused a temporary reduction in the number of experiments in Expression Atlas. finding
Experimental setups
Assay System Perturbation Readout Platform
RNA-sequencing (baseline expression quantification) human and other species tissues/cell types/cell lines (nine species; e.g. Illumina Body Map E-MTAB-513, ENCODE Cell Lines E-GEOD-26284) none (healthy or untreated baseline conditions) FPKM expression levels per gene and splice variant, summarized as median across first technical then biological replicates per condition FASTQC quality control; TopHat 1 mapping to Ensembl reference genome; Cufflinks 1 quantification
RNA-sequencing (differential expression analysis) multiple organisms/tissues/cell lines in ArrayExpress-derived experiments (13 species across differential experiments) contrast-defined (e.g. test compound treatment vs untreated, mutant vs wild type, diseased vs healthy) raw read counts; P-values, log2-fold changes at default FDR 0.05 HTSeq for counting; DESeq for differential expression
Single-channel (one-colour) microarray, mainly gene arrays multiple organisms/tissues/cell types/cell lines (e.g. Drosophila melanogaster CDK8 and Cyclin C homozygous mutants) contrast-defined (e.g. homozygous mutant vs wild type) normalized expression values; P-values, t-statistics, log2-fold changes; MA plots at FDR 0.05 e.g. Affymetrix GeneChip Drosophila Genome 2.0 Array
Two-colour microarray ArrayExpress-derived differential experiments across multiple species contrast-defined (test vs reference biological replicate sets) log2-ratios; P-values, t-statistics, log2-fold changes
MicroRNA microarray ArrayExpress-derived experiments (e.g. E-TABM-713) contrast-defined differential microRNA expression; probe-set to microRNA mappings miRBase release used for probe-set to microRNA mapping
Mass-spectrometry proteomics (planned integration, baseline component only) PRIDE database datasets with EFO-matched sample descriptions none (baseline) protein expression shown in context of baseline expression of the coding gene per condition PRIDE database
Key results
  • As of 24 September 2013, Expression Atlas contains highly curated data from 214 experiments 214 experiments
  • Four baseline RNA-sequencing experiments covering nine species are included 4 experiments, 9 species
  • 210 differential experiments covering 13 species are included 210 experiments, 13 species
  • Differential experiments comprise 10 RNA-sequencing and 200 microarray experiments 10 RNA-seq / 200 microarray
  • Baseline expression search applies a default FPKM cut-off below which expression is treated as background noise; user-adjustable with a histogram of genes expressed above a given cut-off FPKM 0.5 default
  • Differential expression results are called at a default false discovery rate, shown in MA plots (differentially expressed genes in red), with user-selectable FDR FDR 0.05
  • A minimum of three biological sample replicates is enforced to ensure sufficient statistical power to detect differential expression 3 replicates
  • Clicking a non-empty baseline heatmap cell shows a breakdown of the three most abundant splice variants for that gene and condition 3 splice variants
Key statistics
  • count 214 (total curated experiments in Expression Atlas as of 24 September 2013)
  • count 210 (differential experiments (13 species))
  • count 200 (differential microarray experiments, mainly single-channel gene arrays)
  • count 10 (differential RNA-sequencing experiments)
  • count 4 (baseline RNA-sequencing experiments (nine species))
  • other FDR of 0.05 (default false discovery rate for calling differential expression / MA plots)
  • other 0.5 FPKM (default baseline expression cut-off; below this treated as background noise)
  • count three (minimum acceptable number of biological replicates enforced per experiment)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a database/resource paper (Expression Atlas), not a primary experimental study; differential expression is computed for curated pairwise 'contrasts' (a 'test' set vs a 'reference' set of biological replicates) using standardized pipelines that differ by platform. RNA-sequencing data are quantified with HTSeq and analyzed for differential expression with DESeq, while microarray data are summarized with P-values and t-statistics; results are reported as P-values, t-statistics (microarray only), and log2 fold-changes, and filtered in the interface using a default FDR threshold of 0.05. A minimum of three biological replicates per condition is enforced before a contrast is analyzed.

Replicationbiological Sample sizea minimum acceptable number of biological sample replicates (three) is enforced 'to ensure sufficient statistical power to detect differential expression'; no formal power calculation is described Groups'reference' (e.g. healthy/wild-type) vs 'test' (e.g. diseased/mutant) biological replicate sets, per curated contrast Pairingunclear Randomization/blindingnot stated Dispersionunclear Effect sizesyes Multiplicity correctionfalse discovery rate (FDR) thresholding, default 0.05 (specific correction procedure, e.g. Benjamini-Hochberg, not named in text)
Statistical tests used
Test Applied to n Assumptions
DESeq differential expression test (negative-binomial based) on HTSeq counts RNA-sequencing differential expression contrasts (test vs reference biological replicate sets) minimum three biological replicates per condition (as enforced quality-control rule) not stated
t-statistic-based differential expression analysis microarray differential expression contrasts (test vs reference biological replicate sets) minimum three biological replicates per condition (as enforced quality-control rule) not stated
Approaches that could also have been used
  • Microarray contrasts are summarized with P-values and t-statistics without naming the specific test framework.
    Could also: An explicit empirical-Bayes moderated t-test (e.g. limma), which is widely used for microarray contrasts — Moderated t-statistics borrow information across genes to stabilize variance estimates, which can be particularly useful when replicate numbers are small (as low as three here).
  • RNA-seq differential expression is computed with DESeq on HTSeq counts.
    Could also: Newer count-based tools such as DESeq2 or edgeR, which use updated dispersion-shrinkage and independent filtering approaches — These methods can refine variance estimation for lowly expressed genes and are commonly used successors to the original DESeq method.
  • Multiple testing across genes within a contrast is controlled via an FDR threshold (default 0.05), with the specific correction procedure not named.
    Could also: Explicitly reporting the correction method (e.g. Benjamini-Hochberg) alongside the threshold — Naming the exact procedure makes the multiplicity control fully transparent and reproducible from the text alone.
  • A minimum of three biological replicates per condition is enforced as a general quality rule to support statistical power.
    Could also: A formal power analysis tailored to an expected effect size and desired detection sensitivity, performed per experiment — A quantitative power calculation could indicate whether three replicates are sufficient for a specific effect size of interest, rather than relying on a fixed minimum across all experiments.
  • Baseline expression (FPKM) is summarized as a median across technical then biological replicates, without an accompanying variability statistic.
    Could also: Reporting a dispersion measure (e.g. IQR, SD, or range) alongside the median — A dispersion measure would let users gauge how consistent expression levels are across the replicates being summarized.
  • Each differential experiment is broken into pairwise (two-group) contrasts between a reference and a test set.
    Could also: A single multi-group model (e.g. a likelihood-ratio test framework as in DESeq2, or ANOVA-like designs) when more than two conditions are studied — A joint multi-condition test can capture overall condition effects in one combined analysis rather than requiring multiple separate pairwise contrasts.
Software: TopHat · Cufflinks · HTSeq · DESeq · FASTQC

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

genome_wide_tpm_correlation
Reported
Expression Atlas publishes gene-level baseline TPM for E-MTAB-513 (16 human tissues), computed via iRAP 1.0.1 / HISAT2 2.1.0 (Ensembl 95) / FeatureCounts 1.6.2 / kallisto 0.42.4 (current re-processed pipeline; original 2013 paper used TopHat1+Cufflinks1)
Reproduced
salmon 1.10.1 quantification (GRCh38 cDNA, release ~111) on the same raw ENA FASTQ for 5/16 tissues; log1p Pearson r = 0.892-0.910 vs Atlas's own current TPM table across ~33,800 common genes per tissue
within tolerance
marker_gene_tissue_specificity_direction
Reported
ALB high in liver, UMOD high in kidney, MYH1/ACTA1 high in skeletal muscle, GFAP high in brain, DDX4/TNP1 high in testis (Atlas TPM table)
Reproduced
Identical tissue-specific enrichment direction confirmed for all 7 marker/tissue pairs in salmon output
exact
marker_gene_absolute_tpm
Reported
ALB=44903, UMOD=2631, MYH1=196, ACTA1=7410, GFAP=1667, DDX4=281, TNP1=7157 (Atlas TPM)
Reproduced
ALB=19894 (0.44x), UMOD=1321 (0.50x), MYH1=56.7 (0.29x), ACTA1=1689 (0.23x), GFAP=2368 (1.42x), DDX4=90.1 (0.32x), TNP1=3635 (0.51x)
partial
emtab513_dataset_structure
Reported
E-MTAB-513 = "RNA-Seq of human individual tissues and mixture of 16 tissues (Illumina Body Map)"
Reproduced
condensed-sdrf.tsv confirms 16 distinct tissues + 1 "16 tissues mixture" group = 48 total ENA runs (32 individual-tissue + 16 mixture)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator headless) · v1.0 L1 84/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

What deviates: absolute gene-level TPM magnitudes. On identical raw ENA reads, salmon 1.10.1 recovers Expression Atlas's own current TPM table at log1p Pearson r=0.892–0.910 across ~33,800 genes in all 5 sampled tissues, and reproduces 7/7 tissue-marker enrichments in the correct tissue — but per-gene absolute values run 0.23–1.42x off (ACTA1 1689 vs 7410; GFAP 2368 vs 1667).

Whose side: ours, and expectedly so. The paper's 2013 TopHat1+Cufflinks1 pipeline is defunct and was substituted; we also skipped Atlas's adapter/QC filtering, used 1 of 2 technical replicates, applied no quantile normalization, and used Ensembl ~111 against Atlas's release 95. Nothing points at the authors, and nothing is non-derivable.

Severity: moderate. Direction, ranking and biology reproduce exactly; only the scale factor moves. Note the endpoint caveat — the comparison was against Atlas's present-day re-processed table, not against any number printed in the 2013 paper, which for a database-resource paper (headline claims are a content census: 214 experiments, 9/13 species) is the only comparable target available. Coverage is partial (5/16 tissues, no E-GEOD-26284, no differential/microarray pipeline), which keeps the overall grade at yellow rather than green.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.