Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

RummaGEO: Automatic mining of human and mouse gene sets from GEO.

Patterns (N Y) · 2024
L1 63/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • No authors-side cause for any deviation
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
63/100
Reproducibility score
0.6 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 24% of all assessed papers rank 875 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to verify the shipped output, but NOT to re-derive the headline numbers within 80/20. RummaGEO is a DB-building pipeline (limma-voom DE + sentence-transformer/k-means condition detection over ~29k ARCHS4-processed GEO studies); full re-derivation needs tens of GB of ARCHS4 H5 + GPU embedding, which we deliberately did not attempt (the hard ~20%). Reproduction = verifying the shipped GMT gene-set libraries on «our HPC»/«infra» («job», exit 0). RESULT: the documented gene-set size filter (5..2000 genes) reproduces EXACTLY across all 382,402 shipped sets (C4, clean 1:1). The headline counts are PARTIAL due to version drift, not error: our line counts match the live v2.5 (2024-11-04) download page exactly (178,975 human / 203,427 mouse) but the paper reports the 2024-08-22 snapshot (171,441 / 195,265); only v2.5 is still hosted (probed minio: v1.0/1.5/2.0/2.4 -> HTTP 404), so the paper-exact artifact is no longer downloadable. The +~4% growth is consistent with the DB accumulating studies between Aug and Nov 2024. No fabrication signal: counts are internally consistent with shipped data given the known version bump, up/dn sets are near-balanced as the method dictates, and the size filter holds perfectly. NOT attempted: full ARCHS4 re-derivation, Mistral-7B key-term extraction, silhouette QC figure, and the DE adj.P<0.05 threshold (not recoverable from a GMT, which stores only passing gene symbols).

💻 Code ↗ 🗄 Data: 10.5281/zenodo.13358070

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 63
    assessed: 2026-06-14 ⛓ 3c611db564a5
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

GEO metadata can be searched but GEO cannot be searched at the data level; the paper tests whether automatically identifying sample groups and computing differential expression signatures from GEO RNA-seq studies can produce a searchable database of gene sets that recovers known biology and enables discovery.

Core claims
  • RummaGEO is a gene expression signature search engine built from automatically mined human and mouse RNA-seq perturbation studies in GEO resource
  • Groups of samples were automatically identified from GEO metadata and differential expression was computed to extract up/down gene sets for each study method
  • The RummaGEO database contains 171,441 human and 195,265 mouse gene sets extracted from 29,294 GEO studies finding
  • Human and mouse gene sets show significant but incomplete species separation in UMAP space despite gene symbol harmonization finding
  • Gene sets cluster coherently by tissue and, less strongly, by disease when visualized with UMAP finding
  • Most RummaGEO gene sets map closest to the Enrichr 'crowd generated' category, consistent with both being derived from user/automatically extracted GEO signatures finding
  • TF and kinase libraries built from RummaGEO's gene-gene co-occurrence matrix recover known TF targets better than known kinase substrates finding
  • Leiden clustering of the gene similarity space identifies functional gene modules, some of which are also enriched for specific chromosomal locations finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq differential expression (uniformly processed via ARCHS4) human and mouse GEO samples/studies none/other (automatically inferred disease, drug, gene knockout/knockdown/overexpression conditions across studies) up/down differentially expressed gene sets per study ARCHS4
gene set vectorization (IDF + truncated SVD) and UMAP dimensionality reduction 171,441 human and 195,265 mouse gene sets none clustering/separation by species, tissue, disease; silhouette scores
Leiden clustering on gene similarity UMAP human genes (77 clusters) and mouse genes (67 clusters) from RummaGEO co-occurrence matrix none functional gene modules and chromosomal location enrichment
gene set enrichment / overlap comparison with Enrichr library UMAP RummaGEO human/mouse gene sets mapped to human protein-coding genes vs. Enrichr curated gene sets none Euclidean distance between RummaGEO and Enrichr category centroids Enrichr
enrichment analysis benchmarking (Fisher's exact test, ROC/AUC) RummaGEO-derived transcription factor gene set library none recovery rank/AUC of known TF targets ChEA3
enrichment analysis benchmarking (Fisher's exact test, ROC/AUC) RummaGEO-derived kinase gene set library none recovery rank/AUC of known kinase substrates KEA3
Monte Carlo chi-squared test Leiden clusters (n=258) of tissue-labeled human/mouse gene sets none association between cluster identity and tissue type
Key results
  • RummaGEO database contains 171,441 human and 195,265 mouse gene sets from 29,294 GEO studies
  • Species silhouette score for human vs mouse gene sets was significant compared to shuffled labels 0.135 vs -0.0065 ± 0.0038, p=1.456e-17
  • Leiden clusters of tissue-labeled gene sets show significant association with tissue type chi-squared=1.35e6, p<0.0001, n=258 clusters
  • RummaGEO gene sets are closest in UMAP space to the Enrichr 'crowd generated' category, followed by transcription, cell types, then diseases/drugs distances 0.165, 1.518, 2.003, 2.417
  • Cusanovich shRNA TFs library gave the best recovery of known TF targets among benchmarked TF libraries AUC=0.81
  • PTMsigDB drug signatures gave the best recovery of known kinase substrates among benchmarked kinase libraries AUC=0.604
  • RummaGEO TF library recovery was slightly better than Rummagene's, while kinase recovery was slightly worse TF mean AUC 0.757 vs 0.708; kinase mean AUC 0.592 vs 0.640
  • 77 human and 67 mouse gene clusters identified, most with clear functional enrichment and some with chromosome-location bias
Key statistics
  • count 171,441 human gene sets; 195,265 mouse gene sets; 29,294 GEO studies (overall RummaGEO database size)
  • other silhouette score 0.135 (species) vs -0.0065 ± 0.0038 (shuffled) (human vs mouse gene set separation in UMAP)
  • pvalue p = 1.456e-17 (significance of species coherence silhouette score)
  • other chi-squared statistic = 1.35e6, p < 0.0001 (association of Leiden clusters with tissue type)
  • fold_change AUC = 0.81 (Cusanovich shRNA TF library recovery of known TF targets (ChEA3))
  • fold_change AUC = 0.604 (PTMsigDB drug signatures kinase library recovery (KEA3))
  • fold_change mean AUC 0.757 (RummaGEO) vs 0.708 (Rummagene) for TFs; 0.592 vs 0.640 for kinases (comparison of RummaGEO vs Rummagene TF/kinase library performance)
  • other Euclidean distances to Enrichr category centroids: crowd generated 0.165, transcription 1.518, cell types 2.003, diseases/drugs 2.417 (similarity of RummaGEO gene sets to Enrichr categories)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

RummaGEO is a computational database paper that automatically mines gene expression signatures from GEO RNA-seq data, with statistical content focused on descriptive characterization of database contents, cluster quality evaluation (silhouette scores with a permutation-based null), a Monte Carlo chi-squared test for cluster-tissue associations, and Fisher's exact test-based enrichment analysis for benchmarking derived transcription factor and kinase libraries. Library performance was assessed using area under the receiver operating characteristic (ROC) curve (AUC), and results were reported with exact p-values and AUC point estimates. No explicit multiple-testing correction was described despite evaluation across multiple benchmark libraries.

Replicationunclear Sample sizeDatabase scale described (171,441 human and 195,265 mouse gene sets from 29,294 GEO studies); benchmark dataset sizes not stated; no formal power analysis described GroupsHuman vs. mouse gene sets; tissue/disease-labeled Leiden clusters; RummaGEO TF/kinase libraries vs. multiple ChEA3/KEA3 benchmark libraries; RummaGEO vs. Rummagene library AUC Pairingna Randomization/blindingnot stated Dispersionmixed Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Silhouette score with permutation null (shuffled species labels repeated to estimate null mean ± dispersion, implying z-type comparison) Species coherence of human vs. mouse gene sets in UMAP embedding (Figure 2A); observed silhouette = 0.135, null = −0.0065 ± 0.0038, p = 1.456e−17 not stated
Monte Carlo chi-squared test Association between Leiden clusters and tissue type labels on the UMAP (Figure 2B); chi-squared statistic = 1.35e6, p < 0.0001 n = 258 clusters not stated
Fisher's exact test (overlap p-value used to rank enrichment results) Enrichment analysis for benchmarking RummaGEO-derived TF and kinase libraries against ChEA3 and KEA3 benchmark datasets (Figures 5A–5F) not stated
Area under ROC curve (AUC) Benchmarking TF libraries (range >0.70–0.81 AUC) and kinase libraries (up to 0.604 AUC); comparison to Rummagene-derived library performance (mean AUC 0.757 vs. 0.708 for TFs; 0.592 vs. 0.640 for kinases) na
Euclidean distance between UMAP centroids Similarity of RummaGEO gene set space to Enrichr category centroids (Figure 3); crowd-generated closest at 0.165, transcription 1.518, cell types 2.003, diseases and drugs 2.417 na
Approaches that could also have been used
  • ROC AUC was used as the sole performance metric for benchmarking TF and kinase library recovery across all benchmark datasets
    Could also: Precision-recall AUC (AUPRC) could also be reported alongside ROC AUC — When the number of true positives (known TF targets or kinase substrates) is small relative to the total gene space, AUPRC is more sensitive to differences in performance among methods; reporting both metrics together provides a more complete picture of library quality under class imbalance
  • A permutation approach (shuffled species labels, repeated to yield a null mean ± dispersion) was used to assess significance of the observed silhouette score for species coherence
    Could also: A formal bootstrap confidence interval around the observed silhouette score could also quantify uncertainty in the estimate itself — The permutation null characterizes whether the observed score exceeds chance; a bootstrap CI on the observed score additionally conveys sampling variability of the silhouette estimate across the full gene set collection, which may be useful given the very large n
  • Fisher's exact test p-value was used to rank TFs and kinases in enrichment analysis for library benchmarking
    Could also: A hypergeometric test or a rank-based enrichment score (as in GSEA) could also be used to assess overlap significance — Fisher's exact test and the hypergeometric test are mathematically equivalent for this use case and are interchangeable; GSEA-style scoring additionally incorporates the rank ordering of genes, which can be more sensitive when effect sizes vary across the ranked list rather than existing as a binary overlap
  • Leiden community detection was applied to UMAP-projected gene sets and genes to identify clusters, with the number of clusters determined automatically by the algorithm
    Could also: Louvain community detection, or hierarchical clustering with a dendrogram cut, could also partition these spaces — Leiden is known to produce well-connected communities and avoids a known resolution problem of Louvain; hierarchical clustering offers an alternative that does not require a graph construction step and provides a dendrogram that can be cut at multiple resolutions, which may aid interpretability for downstream functional annotation
  • Gene sets were vectorized using inverse document frequency (IDF) weighting followed by truncated SVD as input to UMAP
    Could also: TF-IDF weighting or binary (presence/absence) encoding could also be used as input representations prior to dimensionality reduction — TF-IDF additionally accounts for within-document (within-gene-set) term frequency; binary encoding is simpler and makes no implicit frequency-based assumptions; the choice of representation can affect the structure of the resulting embedding and the clusters it produces
  • A Monte Carlo chi-squared test was used to evaluate the association between 258 Leiden clusters and tissue-type labels
    Could also: A permutation test (randomly shuffling tissue labels across gene sets many times) could also assess the same cluster-label association without distributional assumptions — Monte Carlo chi-squared is well suited when expected cell counts in the contingency table are low (as is likely with 258 clusters and many tissue categories); a label-permutation test makes no distributional assumptions at all and is equally standard for assessing whether clustering tracks a categorical variable
Software: ARCHS4 (uniformly aligned RNA-seq count data source) 2023 release · Leiden algorithm (community detection for clustering) · UMAP (uniform manifold approximation and projection) · ChEA3 (TF enrichment benchmarking framework) · KEA3 (kinase enrichment benchmarking framework)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
15
Impact: medium
Foundation confidence
Built on 3 assessed reference(s) · mean reproducibility 54/100
partly built on non-reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GO:0030301 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0033344 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0033700 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0043062 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0045229 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39569206 (RummaGEO)

Paper: Marino GB, Clarke DJB, Lachmann A, Deng EZ, Ma'ayan A. RummaGEO: Automatic mining of human and mouse gene sets from GEO. Patterns 2024. DOI 10.1016/j.patter.2024.101072. PMCID PMC11573963.

Repo: https://github.com/MaayanLab/rummageo (default branch rummageo, commit 6834cf9940bb28978d3ac8ec10725e79e8a724b4 as of 2026-05-14). Data: zenodo 10.5281/zenodo.13358070 = a code snapshot rummageo-08222024.zip (74 MB), not the gene-set data. Gene-set libraries are served as GMT downloads from https://minio.dev.maayanlab.cloud/rummageo/<version>/.

What RummaGEO is

A pipeline + database + search engine that automatically mines differential- expression gene sets ("signatures") from GEO RNA-seq studies (via the ARCHS4 uniformly-processed expression matrices). For each study it auto-detects sample conditions and computes up/down gene sets between condition pairs.

Pipeline (ETL/, run order from ETL/README.md)

  1. process_ARCHS4.py — pick valid GEO Series from the ARCHS4 H5 (single-cell prob < 0.5; 6 ≤ samples < 50), embed sample metadata (SentenceTransformer all-mpnet-base-v2), k-means (k = n_samples//3), keep clusters ≥ 3 samples, drop studies < 6 samples after clustering.
  2. compute_signatures.py — identify control conditions, run limma-voom pairwise DE per study → per-direction gene lists.
  3. create_gmt.py — write <species>-geo-auto.gmt: one line per signature- direction (... up / ... dn); keep 5 ≤ |genes| ≤ 2000 (if a set > 2000, tighten adj.p from 0.05 → 0.01 → 0.005 → ... until ≤ 2000); DE significance adj.P.Val < 0.05.
  4. create_meta_dict.py, calc_confidence.py (silhouette QC), enrichr_tags.py, extract_key_terms.py (Mistral-7B LLM key terms), make_downloads.py.

In scope (pipeline-derived, attempted)

  • C1/C2 — gene-set counts (human 171,441 / mouse 195,265, abstract/results). Verifiable by counting lines in the shipped GMT libraries.
  • C3 — gene-set size bounds 5–2000 genes (Methods). Verifiable across the full shipped GMT (every set must obey the documented filter).
  • C4 — up/dn pairing & study count (29,294 studies processed). Cross-check against GMT naming + processed-meta JSON.

Out of scope / not attempted (the hard ~20%, with reasons)

  • Full re-derivation of all 171k+195k gene sets from ARCHS4: requires the full ARCHS4 H5 matrices (tens of GB) + GPU embedding + limma-voom over ~29k studies — far beyond an 80/20 reproduction; the method is documented and the output is shipped, so we verify the output instead.
  • LLM key-term extraction (Mistral-7B) — auxiliary annotation, not a core numeric result.
  • Silhouette-score distribution figure — QC visualization, not a headline number.

Key reproducibility limitation (version drift)

The paper's numbers are the 2024-08-22 snapshot. Only v2.5 (2024-11-04) GMTs are still hosted on minio (probed: v1.0/1.5/2.0/2.4 → HTTP 404; v2.5 → 200). The v2.5 download page itself reports 178,975 human / 203,427 mouse gene sets — already larger than the paper. So we reproduce against v2.5 and quantify the drift vs the paper; the exact paper-version artifact is no longer downloadable. This is a genuine reproducibility finding, not a fabrication signal.

C1
Reported
171,441 human gene sets
Reproduced
178,975 (v2.5 GMT line count; matches v2.5 download page exactly)
partial
C2
Reported
195,265 mouse gene sets
Reproduced
203,427 (v2.5 GMT line count; matches v2.5 download page exactly)
partial
C3
Reported
29,294 GEO studies
Reproduced
30,696 distinct GSE in GMT names (14,204 human + 16,492 mouse, not cross-species deduped)
partial
C4
Reported
gene-set size filter 5 <= |genes| <= 2000
Reproduced
min=5, max=2000, 0 violations across all 382,402 shipped gene sets
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 63/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

This is a database-building pipeline where full re-derivation was deliberately out of 80/20 scope; reproduction verified the shipped GMT outputs instead. C4 (size filter 5..2000) is a clean 1:1 match across all 382,402 sets, and the headline counts (C1-C3) differ from the paper by only ~4%, fully explained by version drift: our v2.5 line counts match the live v2.5 download page exactly, but the paper's 2024-08-22 snapshot has been de-hosted and is no longer obtainable. The deviation sits on the data-availability/version side, not the authors' computation, and shows no fabrication signal (counts internally consistent, up/dn sets balanced, filter obeyed). Core conclusion holds; overall yellow because the paper-exact artifact could not be matched 1:1.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

128.9 k
tokens (I/O) · 9.6 M incl. cache
15 min
runtime · 0.01 CPU-h
0.1 GB
peak RAM
1
HPC jobs
hummel
machine