Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Curation of over 10 000 transcriptomic studies to enable data reuse.

Database (Oxford) · 2021
L1 80/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
80/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 56% of all assessed papers rank 484 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> reproduced (1:1 for the pipeline-derived results on the paper's named dataset). PMID 33599246 is the Gemma resource paper; its headline numbers are June-2020 snapshots of a LIVE, continuously-growing database (now 23698 vs 10811 datasets) and are therefore not exact-reproducible without the original dump - recorded as drift, not a mismatch (no June-2020 dump shipped). Reproducible, per-dataset pipeline outputs WERE reproduced on the BRIEF's named dataset GSE23579: (A) the multi-species split into GSE23579.1=human / GSE23579.2=mouse is an EXACT match to the Methods, confirmed via the live Gemma REST API; (B) Gemma's QC sample-sample correlation step was independently recomputed (pure-Python, no Gemma code) on Gemma's processed matrices for both split datasets, giving median 0.985 for both - above the paper's >=0.9 quality threshold and consistent with Gemma's stored geeq.corrMatIssues=0. Method per BRIEF P16: ran the deployed Gemma pipeline on the paper's data + an independent re-computation; the authors' full Java/Spring platform was deliberately NOT rebuilt (disproportionate hard 20%, and the deployed service already exposes the needed outputs). NOT attempted: exact June-2020 aggregate counts (time-varying), full platform ETL from raw GEO, manual-curation accuracy. No fabrication signals found.

💻 Code ↗ 🗄 Data: GSE23579

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 80
    assessed: 2026-06-15 ⛓ cdd63d7c2d57
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Public transcriptomic data in repositories like GEO are difficult to reuse due to unstructured metadata, inconsistent processing/QC, and inconsistent probe-gene mappings; the paper presents the Gemma system as a solution providing uniformly curated and reprocessed transcriptomic datasets to enable reliable data reuse.

Core claims
  • Gemma is a curated database and bioinformatics system that addresses metadata, probe annotation, and expression data inconsistencies in GEO to enable transcriptomic data reuse resource
  • As of June 2020 Gemma contains 10,811 manually curated datasets, over 395,000 samples, and hundreds of curated platforms (microarray and RNA-seq) resource
  • Gemma decouples expression data from microarray platforms and remaps probes to genes by aligning actual probe sequences to reference genomes with BLAT, ensuring consistent annotation method
  • Expression data are uniformly reprocessed (Affymetrix CEL files via RMA, RNA-seq via STAR/RSEM), log2-transformed, quantile normalized, batch corrected (ComBat) and QC-checked for outliers method
  • Dataset topics are annotated with formal ontology terms (10,215 distinct terms from 12 ontologies; 54,316 annotations; mean 5.2 topics/dataset) method
  • Gemma has broad coverage but captures a large majority of available brain-related datasets, which account for 34% of its holdings finding
  • Single-cell RNA-seq datasets are currently omitted due to minimal biological replication and inconsistent representation in GEO finding
  • Curated data and differential expression analyses are accessible via website, RESTful web service, and an R package resource
Experimental setups
Assay System Perturbation Readout Platform
Microarray (Affymetrix GeneChip) expression reprocessing human, mouse, rat (and other taxa) GEO datasets none probe set expression levels (RMA-summarized, log2, quantile normalized) Affymetrix Analysis Power Tools (apt-probeset-summarize), RMA algorithm
RNA-sequencing expression quantification pipeline human, mouse, rat GEO/SRA datasets none gene-level expression quantification Cutadapt, STAR aligner, RSEM
Probe-to-gene sequence mapping/alignment microarray platforms (Affymetrix, Agilent, Illumina, two-color) none specific probe alignments mapped to transcripts/genes BLAT against reference genomes (hg38, mm10, rn6); UCSC GoldenPath annotations
Batch effect detection and correction curated expression datasets (microarray and RNA-seq) none batch factor / batch-corrected expression matrix in-house ComBat implementation
Quality control diagnostics and outlier detection curated expression datasets none PCA, mean-variance relationship, sample-sample correlation matrix, outlier flags
Metadata/ontology curation GEO Series/Sample records none ontology topic annotations and experimental design factors
Key results
  • Gemma contains manually curated transcriptomic datasets 10,811 datasets
  • Total samples represented in Gemma over 395,000 samples
  • Brain-related datasets account for a large share of Gemma holdings 34%
  • Topic annotations applied across datasets 54,316 annotations; mean 5.2 topics/dataset
  • Distinct ontology terms used for topic annotation 10,215 terms from 12 ontologies
  • GEO DataSets (GDS) with curated design metadata are available for only a small fraction of GEO Series <5%
  • Predefined text-to-ontology-term dictionary used for automated annotation mapping 674 mappings
  • Of GEO transcriptomic studies, a majority are from human, mouse and rat 71,233 of 97,379 (73%)
Key statistics
  • count 10,811 (manually curated datasets in Gemma as of June 2020)
  • count over 395,000 (samples in Gemma)
  • count 54,316 (total topic annotations)
  • mean 5.2 (mean topics per dataset)
  • count 10,215 (distinct topic terms from 12 ontologies)
  • other 34% (proportion of Gemma holdings that are brain-related)
  • count 97,379 (transcriptomic studies in NCBI GEO as of June 2020)
  • other 73% (GEO studies generated in human, mouse and rat (71,233 of 97,379))

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a database/resource paper describing the Gemma system for curating and reprocessing transcriptomic datasets; it does not report a hypothesis-testing experimental study but rather data-processing and analysis pipelines. Reported quantitative content is largely descriptive (counts and means, e.g. 10 811 datasets, mean topics/dataset = 5.2), and the 'statistical' elements are pipeline methods such as quantile normalization, ComBat batch correction, principal component analysis and an interquartile-range-based outlier-detection heuristic. No formal between-group significance testing, sample-size justification or multiplicity correction is reported in the text provided.

Replicationunclear Sample sizeDescriptive collection totals are given (e.g. 10 811 datasets, >395 000 samples); no sample size or power calculation for a statistical comparison is described. Datasets with <10 total samples are de-prioritized and biological replication is preferred for inclusion. Groupsnone (resource description; no experimental group comparison) Pairingna Randomization/blindingna DispersionIQR
Approaches that could also have been used
  • Potential outlier samples are flagged using a heuristic based on the interquartile range of sample–sample correlations, followed by manual review.
    Could also: Model-based or formally calibrated outlier methods (e.g. Hampel/MAD-based thresholds, PCA Hotelling's T² with control limits, or array-quality metrics such as those in arrayQualityMetrics) could also be used. — Such approaches would attach an explicit, tunable false-flag rate or probability to each call, which can complement an IQR rule when one wants a documented statistical threshold.
  • Batch effects are corrected with an in-house ComBat implementation that auto-selects covariates from principal-component loadings.
    Could also: Alternatives include including batch as a covariate within the differential-expression model (e.g. limma/edgeR/DESeq2 design terms), RUV or SVA for unknown latent factors, or limma's removeBatchEffect. — Modeling batch inside the downstream test rather than adjusting the data beforehand can preserve degrees of freedom and avoid overstating precision, and SVA/RUV can capture batch structure when date-stamp information is unavailable.
  • Time-stamp-based batches are defined via a one-dimensional clustering using simple time-gap heuristics.
    Could also: A formal one-dimensional clustering criterion (e.g. Jenks natural breaks, a gap statistic, or a change-point model) could also assign batch boundaries. — A criterion with an objective cut selection can make batch boundaries reproducible and reduce sensitivity to a fixed gap threshold across datasets with differing processing cadences.
  • Expression data are log2-transformed and quantile normalized within the common post-processing pipeline.
    Could also: For RNA-seq specifically, library-size-aware normalizations (TMM, median-of-ratios/DESeq2 size factors) or variance-stabilizing transforms could also be applied. — These methods are tailored to count-based mean–variance structure and are commonly paired with the corresponding differential-expression frameworks, which the paper notes is important for RNA-seq.
  • Summary content is largely reported as counts and a single mean (mean topics/dataset = 5.2).
    Could also: Reporting an accompanying measure of spread (SD, IQR or range) or the full distribution alongside the mean would also characterize the holdings. — For a quantity like topics-per-dataset that is likely skewed, a dispersion measure or median/IQR conveys the distribution shape that a mean alone does not.
Software: ComBat (in-house implementation, for batch correction) · Affymetrix Analysis Power Tools (apt-probeset-summarize, RMA algorithm) · BLAT (probe-to-genome alignment) · STAR (RNA-seq alignment) · RSEM (RNA-seq quantification) · Cutadapt (adapter trimming)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
59
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.
Cited by (assessed papers) (3)

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GPL1261 GEO in Results (http://purl.org/orb/Results)
also used by 2 papers:
GPL1355 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
EFO_0000246 EFO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
EFO_0000324 EFO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
EFO_0000399 EFO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
EFO_0000408 EFO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
EFO_0000410 EFO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
EFO_0000513 EFO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
EFO_0000635 EFO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
EFO_0000724 EFO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
EFO_0000727 EFO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
EFO_0001185 EFO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
EFO_0001461 EFO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
EFO_0004425 EFO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
EFO_0005135 EFO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
EFO_0005168 EFO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GPL6246 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GPL6887 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE107999 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE13524 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE14499 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE1463 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE15721 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE15774 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE18162 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE2198 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE23115 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE23579 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE2426 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE3253 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE36051 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE40463 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE52022 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE8030 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE9509 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33599246 (Gemma database paper)

Paper: Lim et al. 2021, Curation of over 10 000 transcriptomic studies to enable data reuse. Database (Oxford), 10.1093/database/baab006. Named dataset (BRIEF): geo:GSE23579. Code: github.com/PavlidisLab/Gemma.

Nature of the paper

This is a database/resource description paper. Most reported numbers are aggregate descriptive statistics of the live Gemma database as of June 2020 (e.g. "10 811 curated datasets", "811 platforms"). Gemma is a continuously-updated production service, so those snapshots are not exactly reproducible by design — the live DB now holds 23 698 public datasets (checked 2026-06-15, REST API totalElements), i.e. it has more than doubled. Reproducing the June-2020 counts would require the June-2020 DB dump, which is not shipped. → out of scope for exact 1:1, recorded as drift context, not graded as mismatch.

In scope — pipeline-derived, per-dataset, reproducible on the named dataset

Gemma's per-dataset processing pipeline (curation → multi-species split → processed-data assembly → QC: sample–sample correlation, outlier flagging, batch detection, GEEQ quality scoring) produces stable, per-dataset outputs that are exposed via the public REST API and can be reproduced on the paper's own named dataset GSE23579.

# In-scope result Pipeline step How reproduced
A GSE23579 is split into GSE23579.1 (human) + GSE23579.2 (mouse) multi-species split (Methods, "Multi-platform, multi-species and overlapping datasets") live Gemma REST API: confirm two taxon-specific EEs
B Datasets pass QC when median sample–sample correlation ≥ 0.9 (82% of datasets; Results) QC / GEEQ correlation-matrix step download Gemma processed matrices for both split EEs («infra»), recompute Pearson sample–sample correlation, compare to threshold + Gemma's stored geeq.corrMatIssues

Out of scope (not attempted / not exact-reproducible)

  • Aggregate June-2020 counts (10 811 datasets; 4593/4933/894 human/mouse/rat; 811 platforms; 54 316 ontology terms; mean 5.2 topics) — time-varying live-DB snapshots, no June-2020 dump shipped.
  • The full Gemma Java/Spring platform build + ETL from raw GEO (massive enterprise app + DB backend). We reproduce pipeline-derived outputs via the deployed Gemma service + an independent re-computation of one QC step, which is a valid reproduction of the pipeline's results on the paper's data (BRIEF P16: running the tool/pipeline on the paper's data is equally valid). We did not rebuild the authors' Java codebase — noted explicitly, no completeness claim.
  • Manual curation quality, ontology-mapping accuracy — wet-lab/manual, out of scope.
A_multispecies_split
Reported
GSE23579 split into GSE23579.1 (human) + GSE23579.2 (mouse)
Reproduced
GSE23579.1 = Homo sapiens (Gemma EE 2670); GSE23579.2 = Mus musculus (Gemma EE 2669); both accession GSE23579, 4 samples each
exact
B_qc_corr_human
Reported
median sample-sample correlation >= 0.9 (QC criterion; 82% of datasets)
Reproduced
median 0.985 (6 pairs / 4 samples / 54675 probes), all >= 0.98; matches Gemma geeq.corrMatIssues=0
within tolerance
B_qc_corr_mouse
Reported
median sample-sample correlation >= 0.9 (QC criterion; 82% of datasets)
Reproduced
median 0.985 (6 pairs / 4 samples / 45101 probes), all >= 0.98; matches Gemma geeq.corrMatIssues=0
within tolerance
CTX_dataset_count
Reported
10 811 curated datasets (June 2020)
Reproduced
23 698 public datasets (live, 2026-06-15) - not exact-reproducible (live DB grew 2.2x)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 80/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

This resource paper describes the live Gemma database; its reproducible, per-dataset pipeline outputs on the named dataset GSE23579 reproduced 1:1 — the taxon split (GSE23579.1=human EE2670, GSE23579.2=mouse EE2669) is an exact structural match and the QC sample-sample correlation independently recomputed to median 0.985 for both species, comfortably above the paper's >=0.9 gate and consistent with stored geeq.corrMatIssues=0. The only deviation is the headline aggregate count (10811 in June 2020 vs 23698 live today), which is explained entirely by legitimate growth of a continuously-updated service and the absence of a shipped June-2020 dump — a data-availability/time-drift issue, not an authors' defect or fabrication. Net: a clean reproduction of all testable claims, downgraded to overall yellow only because the time-varying headline figure is not exactly reproducible.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

103.6 k
tokens (I/O) · 6.2 M incl. cache
10 min
runtime · 0 CPU-h
0.1 GB
peak RAM
2
HPC jobs
hummel
machine