Curation of over 10 000 transcriptomic studies to enable data reuse.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough -> reproduced (1:1 for the pipeline-derived results on the paper's named dataset). PMID 33599246 is the Gemma resource paper; its headline numbers are June-2020 snapshots of a LIVE, continuously-growing database (now 23698 vs 10811 datasets) and are therefore not exact-reproducible without the original dump - recorded as drift, not a mismatch (no June-2020 dump shipped). Reproducible, per-dataset pipeline outputs WERE reproduced on the BRIEF's named dataset GSE23579: (A) the multi-species split into GSE23579.1=human / GSE23579.2=mouse is an EXACT match to the Methods, confirmed via the live Gemma REST API; (B) Gemma's QC sample-sample correlation step was independently recomputed (pure-Python, no Gemma code) on Gemma's processed matrices for both split datasets, giving median 0.985 for both - above the paper's >=0.9 quality threshold and consistent with Gemma's stored geeq.corrMatIssues=0. Method per BRIEF P16: ran the deployed Gemma pipeline on the paper's data + an independent re-computation; the authors' full Java/Spring platform was deliberately NOT rebuilt (disproportionate hard 20%, and the deployed service already exposes the needed outputs). NOT attempted: exact June-2020 aggregate counts (time-varying), full platform ETL from raw GEO, manual-curation accuracy. No fabrication signals found.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 80assessed: 2026-06-15 ⛓ cdd63d7c2d57
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusPublic transcriptomic data in repositories like GEO are difficult to reuse due to unstructured metadata, inconsistent processing/QC, and inconsistent probe-gene mappings; the paper presents the Gemma system as a solution providing uniformly curated and reprocessed transcriptomic datasets to enable reliable data reuse.
- ★ Gemma is a curated database and bioinformatics system that addresses metadata, probe annotation, and expression data inconsistencies in GEO to enable transcriptomic data reuse resource
- ★ As of June 2020 Gemma contains 10,811 manually curated datasets, over 395,000 samples, and hundreds of curated platforms (microarray and RNA-seq) resource
- ★ Gemma decouples expression data from microarray platforms and remaps probes to genes by aligning actual probe sequences to reference genomes with BLAT, ensuring consistent annotation method
- ★ Expression data are uniformly reprocessed (Affymetrix CEL files via RMA, RNA-seq via STAR/RSEM), log2-transformed, quantile normalized, batch corrected (ComBat) and QC-checked for outliers method
- ★ Dataset topics are annotated with formal ontology terms (10,215 distinct terms from 12 ontologies; 54,316 annotations; mean 5.2 topics/dataset) method
- ★ Gemma has broad coverage but captures a large majority of available brain-related datasets, which account for 34% of its holdings finding
- Single-cell RNA-seq datasets are currently omitted due to minimal biological replication and inconsistent representation in GEO finding
- ★ Curated data and differential expression analyses are accessible via website, RESTful web service, and an R package resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Microarray (Affymetrix GeneChip) expression reprocessing | human, mouse, rat (and other taxa) GEO datasets | none | probe set expression levels (RMA-summarized, log2, quantile normalized) | Affymetrix Analysis Power Tools (apt-probeset-summarize), RMA algorithm |
| RNA-sequencing expression quantification pipeline | human, mouse, rat GEO/SRA datasets | none | gene-level expression quantification | Cutadapt, STAR aligner, RSEM |
| Probe-to-gene sequence mapping/alignment | microarray platforms (Affymetrix, Agilent, Illumina, two-color) | none | specific probe alignments mapped to transcripts/genes | BLAT against reference genomes (hg38, mm10, rn6); UCSC GoldenPath annotations |
| Batch effect detection and correction | curated expression datasets (microarray and RNA-seq) | none | batch factor / batch-corrected expression matrix | in-house ComBat implementation |
| Quality control diagnostics and outlier detection | curated expression datasets | none | PCA, mean-variance relationship, sample-sample correlation matrix, outlier flags | — |
| Metadata/ontology curation | GEO Series/Sample records | none | ontology topic annotations and experimental design factors | — |
- – Gemma contains manually curated transcriptomic datasets 10,811 datasets
- – Total samples represented in Gemma over 395,000 samples
- – Brain-related datasets account for a large share of Gemma holdings 34%
- – Topic annotations applied across datasets 54,316 annotations; mean 5.2 topics/dataset
- – Distinct ontology terms used for topic annotation 10,215 terms from 12 ontologies
- – GEO DataSets (GDS) with curated design metadata are available for only a small fraction of GEO Series <5%
- – Predefined text-to-ontology-term dictionary used for automated annotation mapping 674 mappings
- – Of GEO transcriptomic studies, a majority are from human, mouse and rat 71,233 of 97,379 (73%)
- count 10,811 (manually curated datasets in Gemma as of June 2020)
- count over 395,000 (samples in Gemma)
- count 54,316 (total topic annotations)
- mean 5.2 (mean topics per dataset)
- count 10,215 (distinct topic terms from 12 ontologies)
- other 34% (proportion of Gemma holdings that are brain-related)
- count 97,379 (transcriptomic studies in NCBI GEO as of June 2020)
- other 73% (GEO studies generated in human, mouse and rat (71,233 of 97,379))
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a database/resource paper describing the Gemma system for curating and reprocessing transcriptomic datasets; it does not report a hypothesis-testing experimental study but rather data-processing and analysis pipelines. Reported quantitative content is largely descriptive (counts and means, e.g. 10 811 datasets, mean topics/dataset = 5.2), and the 'statistical' elements are pipeline methods such as quantile normalization, ComBat batch correction, principal component analysis and an interquartile-range-based outlier-detection heuristic. No formal between-group significance testing, sample-size justification or multiplicity correction is reported in the text provided.
-
Potential outlier samples are flagged using a heuristic based on the interquartile range of sample–sample correlations, followed by manual review.↳ Could also: Model-based or formally calibrated outlier methods (e.g. Hampel/MAD-based thresholds, PCA Hotelling's T² with control limits, or array-quality metrics such as those in arrayQualityMetrics) could also be used. — Such approaches would attach an explicit, tunable false-flag rate or probability to each call, which can complement an IQR rule when one wants a documented statistical threshold.
-
Batch effects are corrected with an in-house ComBat implementation that auto-selects covariates from principal-component loadings.↳ Could also: Alternatives include including batch as a covariate within the differential-expression model (e.g. limma/edgeR/DESeq2 design terms), RUV or SVA for unknown latent factors, or limma's removeBatchEffect. — Modeling batch inside the downstream test rather than adjusting the data beforehand can preserve degrees of freedom and avoid overstating precision, and SVA/RUV can capture batch structure when date-stamp information is unavailable.
-
Time-stamp-based batches are defined via a one-dimensional clustering using simple time-gap heuristics.↳ Could also: A formal one-dimensional clustering criterion (e.g. Jenks natural breaks, a gap statistic, or a change-point model) could also assign batch boundaries. — A criterion with an objective cut selection can make batch boundaries reproducible and reduce sensitivity to a fixed gap threshold across datasets with differing processing cadences.
-
Expression data are log2-transformed and quantile normalized within the common post-processing pipeline.↳ Could also: For RNA-seq specifically, library-size-aware normalizations (TMM, median-of-ratios/DESeq2 size factors) or variance-stabilizing transforms could also be applied. — These methods are tailored to count-based mean–variance structure and are commonly paired with the corresponding differential-expression frameworks, which the paper notes is important for RNA-seq.
-
Summary content is largely reported as counts and a single mean (mean topics/dataset = 5.2).↳ Could also: Reporting an accompanying measure of spread (SD, IQR or range) or the full distribution alongside the mean would also characterize the holdings. — For a quantity like topics-per-dataset that is likely skewed, a dispersion measure or median/IQR conveys the distribution shape that a mean alone does not.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Brain-related transcriptomic datasets comprise 34% of all manually curated studies in the Gemma database.other gemma-curated-datasets 2021×1papers★ This paper is the founder (earliest)
-
Fewer than 5% of GEO Series have curated experimental design metadata available via GEO DataSets (GDS), indicating a major curation gap.other geo transcriptomic studies 2021×1papers★ This paper is the founder (earliest)
-
Human, mouse, and rat collectively account for 73% of transcriptomic studies deposited in GEO.other geo transcriptomic studies 2021×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
- CoINcIDE: A framework for discovery of patient... L1 87/100
- A curated collection of transcriptome datasets... L1 62/100
- An NMF-Based Methodology for Selecting Biomark... L1 84/100
- Unveiling prognostics biomarkers of tyrosine m...⚑ L1 51/100 ⚑
- Meta-analysis of gene expression profiles of l... L1 78/100
- Colorectal Cancer Prediction Based on Weighted...⚑ L1 80/100 ⚑
- Construction and Validation of an Immune Infil...⚑ L1 51/100 ⚑
- Identification of a novel 10 immune-related ge...
- Exploration of the shared diagnostic genes and... L1 76/100
- IRSN-23 gene diagnosis enhances breast cancer... L1 71/100
- Molecular Classification Models for Triple Neg... L1 86/100
- Predicting Bone Metastasis Using Gene Expressi... L1 62/100
- Autoencoder Networks Decipher the Association... L1 74/100
- Comprehensive analysis of a novel RNA modifica... L1 71/100
- Discovery and validation of molecular patterns... L1 83/100
- Comparative profiling of skeletal muscle model... L1 64/100
- A curated collection of transcriptome datasets... L1 62/100
- PulmonDB: a curated lung disease gene expressi...⚑ L1 53/100 ⚑
- An NMF-Based Methodology for Selecting Biomark... L1 84/100
- Meta-analysis of gene expression profiles of l... L1 78/100
- Predicting Bone Metastasis Using Gene Expressi... L1 62/100
- Comprehensive analysis of a novel RNA modifica... L1 71/100
- Comparative profiling of skeletal muscle model... L1 64/100
- CoINcIDE: A framework for discovery of patient... L1 87/100
- A curated collection of transcriptome datasets... L1 62/100
- PulmonDB: a curated lung disease gene expressi...⚑ L1 53/100 ⚑
- Meta-analysis of gene expression profiles of l... L1 78/100
- Screening of Diagnostic Biomarkers and Immune...⚑ L1 48/100 ⚑
- IRSN-23 gene diagnosis enhances breast cancer... L1 71/100
- Predicting Bone Metastasis Using Gene Expressi... L1 62/100
- A curated collection of transcriptome datasets... L1 62/100
- PulmonDB: a curated lung disease gene expressi...⚑ L1 53/100 ⚑
- VIGET: A web portal for study of vaccine-induc... L1 64/100
- Exploration of the shared diagnostic genes and... L1 76/100
- Exploring the key genetic association between... L1 91/100
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-33599246 (Gemma database paper)
Paper: Lim et al. 2021, Curation of over 10 000 transcriptomic studies to
enable data reuse. Database (Oxford), 10.1093/database/baab006.
Named dataset (BRIEF): geo:GSE23579. Code: github.com/PavlidisLab/Gemma.
Nature of the paper
This is a database/resource description paper. Most reported numbers are
aggregate descriptive statistics of the live Gemma database as of June 2020
(e.g. "10 811 curated datasets", "811 platforms"). Gemma is a continuously-updated
production service, so those snapshots are not exactly reproducible by design —
the live DB now holds 23 698 public datasets (checked 2026-06-15, REST API
totalElements), i.e. it has more than doubled. Reproducing the June-2020 counts
would require the June-2020 DB dump, which is not shipped. → out of scope for
exact 1:1, recorded as drift context, not graded as mismatch.
In scope — pipeline-derived, per-dataset, reproducible on the named dataset
Gemma's per-dataset processing pipeline (curation → multi-species split → processed-data assembly → QC: sample–sample correlation, outlier flagging, batch detection, GEEQ quality scoring) produces stable, per-dataset outputs that are exposed via the public REST API and can be reproduced on the paper's own named dataset GSE23579.
| # | In-scope result | Pipeline step | How reproduced |
|---|---|---|---|
| A | GSE23579 is split into GSE23579.1 (human) + GSE23579.2 (mouse) |
multi-species split (Methods, "Multi-platform, multi-species and overlapping datasets") | live Gemma REST API: confirm two taxon-specific EEs |
| B | Datasets pass QC when median sample–sample correlation ≥ 0.9 (82% of datasets; Results) | QC / GEEQ correlation-matrix step | download Gemma processed matrices for both split EEs («infra»), recompute Pearson sample–sample correlation, compare to threshold + Gemma's stored geeq.corrMatIssues |
Out of scope (not attempted / not exact-reproducible)
- Aggregate June-2020 counts (10 811 datasets; 4593/4933/894 human/mouse/rat; 811 platforms; 54 316 ontology terms; mean 5.2 topics) — time-varying live-DB snapshots, no June-2020 dump shipped.
- The full Gemma Java/Spring platform build + ETL from raw GEO (massive enterprise app + DB backend). We reproduce pipeline-derived outputs via the deployed Gemma service + an independent re-computation of one QC step, which is a valid reproduction of the pipeline's results on the paper's data (BRIEF P16: running the tool/pipeline on the paper's data is equally valid). We did not rebuild the authors' Java codebase — noted explicitly, no completeness claim.
- Manual curation quality, ontology-mapping accuracy — wet-lab/manual, out of scope.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This resource paper describes the live Gemma database; its reproducible, per-dataset pipeline outputs on the named dataset GSE23579 reproduced 1:1 — the taxon split (GSE23579.1=human EE2670, GSE23579.2=mouse EE2669) is an exact structural match and the QC sample-sample correlation independently recomputed to median 0.985 for both species, comfortably above the paper's >=0.9 gate and consistent with stored geeq.corrMatIssues=0. The only deviation is the headline aggregate count (10811 in June 2020 vs 23698 live today), which is explained entirely by legitimate growth of a continuously-updated service and the absence of a shipped June-2020 dump — a data-availability/time-drift issue, not an authors' defect or fabrication. Net: a clean reproduction of all testable claims, downgraded to overall yellow only because the time-varying headline figure is not exactly reproducible.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.