Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Mouse-Geneformer: A deep learning model for mouse single-cell transcriptome and its cross-species utility.

PLoS Genet · 2025
L1 95/100 PQI 94
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
95/100
Reproducibility score
1.2 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 89% of all assessed papers rank 105 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL->strong. Architecture EXACT (5/5). Cell-type classification (Table 3): 7/9 organs reproduced, accuracy EXACT/within-tol for ALL 7 (5 exact, 2 within-tol); macro-F1 exact/within-tol for 6/7 (brain F1 higher, 93.97 vs 88.01). spleen+kidney computing. The paper's Table 3 reproduces faithfully.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 92
    assessed: 2026-06-14 ⛓ 156d2bb9eb6a
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

A mouse-specific version of Geneformer (mouse-Geneformer), pre-trained on a large mouse scRNA-seq corpus, can accurately model the mouse transcriptome for downstream tasks like cell type classification and disease-gene discovery, and may also generalize cross-species to human transcriptome analysis.

Core claims
  • Mouse-Geneformer, a Transformer Encoder model pre-trained via masked-token self-supervised learning on mouse-Genecorpus-20M, was successfully constructed following the original human Geneformer architecture. resource
  • mouse-Genecorpus-20M is a large-scale mouse scRNA-seq dataset (20,630,028 cells from 1,070 datasets) compiled from public databases and made publicly available. resource
  • Fine-tuned mouse-Geneformer enhances cell type classification accuracy compared to conventional methods. finding
  • In silico perturbation experiments with mouse-Geneformer identified disease-causing genes that were validated in vivo. finding
  • After ortholog-based gene name conversion, mouse-Geneformer fine-tuned on human data achieves cell type classification accuracy comparable to the original human Geneformer. finding
  • In silico simulation of a human myocardial infarction model using mouse-Geneformer gave results similar to human-Geneformer, while a COVID-19 model gave only partially consistent results, reflecting mouse insusceptibility to SARS-CoV-2. finding
  • The Rank Value Encoding method converts single-cell transcriptomes into ranked 'cell sentences' used as Transformer input tokens. method
Experimental setups
Assay System Perturbation Readout Platform
scRNA-seq corpus construction and self-supervised pretraining (masked token prediction) healthy wild-type Mus musculus, multiple organs/embryonic stages (mouse-Genecorpus-20M) none masked gene token prediction (15% tokens masked) 8x NVIDIA V100 GPUs (32GB), Transformer Encoder
Cell type classification benchmarking mouse scRNA-seq from 12 organs (urethra/prostate, embryos, kidney, tongue, thymus, mammary gland, large intestine, limb muscle, spleen, heart, brain, kidney) via CELLxGENE/GEO fine-tuning of mouse-Geneformer vs. scVAE vs. scDeepSort classification accuracy and F1 score scikit-learn (accuracy_score, f1_score), UMAP via scanpy
Cell type classification with vs. without prior pretraining mouse scRNA-seq presence/absence of pretraining (frozen Transformer blocks: 0 vs 6) classification accuracy
In silico perturbation (simulated gene knockout/manipulation) mouse scRNA-seq, mouse-Geneformer in silico gene perturbation identification of disease-causing genes
Cross-species cell type classification after ortholog-based gene name conversion human scRNA-seq data, mouse-Geneformer fine-tuned fine-tuning with human data classification accuracy compared to human-Geneformer
Cross-species in silico perturbation simulation human disease models (myocardial infarction, COVID-19), mouse-Geneformer in silico gene perturbation after gene name conversion comparison of predicted disease genes/results to human-Geneformer
Key results
  • Initial compilation of 119 million cells from 1,089 mouse scRNA-seq datasets, filtered down to 20,630,028 cells from 1,070 datasets to form mouse-Genecorpus-20M
  • Pretraining of mouse-Geneformer on 8 V100 GPUs took approximately 2 days 2 days
  • Cross-species human data analysis with fine-tuned mouse-Geneformer achieved accuracy comparable to original human Geneformer
  • Myocardial infarction in silico model results with mouse-Geneformer were similar to human-Geneformer
  • COVID-19 in silico model results with mouse-Geneformer were only partially consistent with human-Geneformer, attributed to species-specific SARS-CoV-2 susceptibility
Key statistics
  • count 119 million cells (cells before filtering across 1,089 compiled mouse scRNA-seq datasets)
  • count 20,630,028 cells (final filtered mouse-Genecorpus-20M dataset from 1,070 datasets)
  • count 1,089 datasets compiled; 1,070 retained after filtering (mouse-Genecorpus-20M dataset sourcing)
  • other 15% of tokens masked (masked token prediction pretraining task)
  • other 80%/20% train/test split (cell type classification benchmarking data split)
  • other 10 epochs (training/fine-tuning schedule across all experiments (Table 1))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper develops mouse-Geneformer, a Transformer Encoder deep learning model pre-trained on 21 million mouse scRNA-seq profiles (mouse-Genecorpus-20M). Downstream evaluation relies on classification accuracy and F1 score computed against held-out test splits across twelve organ datasets, comparing mouse-Geneformer to scVAE and scDeepSort. In silico perturbation experiments and cross-species (human) analyses are evaluated qualitatively by concordance with known biology; no inferential statistics or p-values are reported in the visible text.

Replicationunclear Sample sizeDataset totals stated (20,630,028 cells after filtering from 119M); per-organ evaluation n not stated; number of independent training/evaluation runs not stated; 'highest accuracy and F1 score was recorded' implies multiple epochs but not independent replicate runs Groupsmouse-Geneformer vs scVAE vs scDeepSort (cell type classification); pretrained vs non-pretrained mouse-Geneformer; mouse-Geneformer vs human Geneformer on human data (cross-species) Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Classification accuracy (scikit-learn accuracy_score) Cell type classification comparison: mouse-Geneformer vs scVAE vs scDeepSort across twelve organ datasets 20% held-out test split of each organ dataset; absolute n per organ not stated not stated
F1 score — harmonic mean of precision and recall (scikit-learn f1_score) Cell type classification comparison: mouse-Geneformer vs scVAE vs scDeepSort across twelve organ datasets 20% held-out test split of each organ dataset; absolute n per organ not stated not stated
Euclidean distance (data integrity check) Verification that evaluation datasets had no zero-value overlap with training corpus (no data leakage) 1,089 compiled datasets; exact n compared not stated not stated
UMAP dimensionality reduction (qualitative visualization) Cell distribution visualisation via scanpy null na
Approaches that could also have been used
  • Performance was evaluated on a single random 80/20 train-test split per organ dataset, and 'the highest accuracy and F1 score was recorded'
    Could also: Stratified k-fold cross-validation (e.g., k=5 or k=10) repeated across multiple random seeds, reporting mean ± SD of accuracy and F1 — A single split produces a point estimate whose sampling variance is unknown; repeated cross-validation yields confidence in whether observed differences between models are stable across data partitions, which is particularly informative when per-organ cell counts are modest
  • Model comparisons (mouse-Geneformer vs scVAE vs scDeepSort) are presented as single point-estimate accuracy and F1 values with no test of whether differences exceed chance variation
    Could also: Bootstrap confidence intervals or a permutation test on the difference in accuracy/F1 between models on the same held-out cells — With a single split, a difference of, say, 2 percentage points in accuracy could arise from random partitioning; a resampling-based interval quantifies this uncertainty without assuming a parametric distribution
  • Results across twelve organ datasets are reported individually without a summary statistic or aggregation method
    Could also: Macro-averaged or weighted-average accuracy/F1 across organs, with per-organ values as supplementary detail, or a paired comparison (e.g., Wilcoxon signed-rank test across the 12 organs) treating organ as the unit of replication — Summarising across organs makes the overall magnitude of the performance difference interpretable and allows a single inferential statement about which model performs better across diverse tissues
  • In silico perturbation experiment results (disease-causing gene identification) are evaluated by qualitative concordance with known in vivo results
    Could also: A quantitative rank-based concordance metric such as Spearman correlation or normalised discounted cumulative gain (nDCG) between model-ranked perturbation scores and known gene effect sizes — Quantitative concordance metrics allow comparison between mouse-Geneformer and human Geneformer predictions on a continuous scale, complementing the binary consistent/inconsistent framing
  • Cell distribution is visualised with UMAP for qualitative assessment of cluster separation
    Could also: Quantitative cluster quality indices such as Adjusted Rand Index (ARI) or Average Silhouette Width computed against annotated cell-type labels — UMAP projections are sensitive to hyperparameters and can be visually misleading for overlapping populations; a numerical index ties the visualisation to a reproducible scalar summary
  • The paper records only the highest epoch's accuracy and F1 during fine-tuning rather than reporting a mean across epochs or a held-out validation curve
    Could also: Early-stopping on a held-out validation set with performance reported at the stopping epoch, or mean ± SD over the last N epochs after convergence — Selecting the best epoch on test data optimistically biases the reported metric; a validation-set stopping criterion or averaging over converged epochs gives a less optimistic and more generalisable estimate
Software: Python/scikit-learn (accuracy_score, f1_score) · Python/scanpy (preprocessing, UMAP) · Python/pandas · Python/numpy · Python/scipy · Python/loompy

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
6
Impact: low
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 81/100
stands on reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE132042 GEO in Methods (http://purl.org/orb/Methods)
also used by 2 papers:
EGAS00001004571 EGA in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
EGAS00001006330 EGA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE144870 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE145929 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE147559 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE149689 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE150728 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE150861 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE155673 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE190094 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE195665 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE197353 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40106407 (Mouse-Geneformer)

Paper: Ito et al., "Mouse-Geneformer: A deep learning model for mouse single-cell transcriptome and its cross-species utility." PLOS Genetics 2025. PMID 40106407 / PMC11964219 / DOI 10.1371/journal.pgen.1011420.

Repo: https://github.com/machine-perception-robotics-group/Mouse-Geneformer commit ed17d455193ed4c4d93230a8f5f2de349cf20c81 (cloned on «infra»). It is a fork of the Geneformer (ctheodoris) codebase adapted for mouse.

Pretrained model + data (HuggingFace, MPRG/Mouse-Genecorpus-20M):

  • checkpoint mouse-Geneformer (base): Google Drive file 1gM3gcc3DlNGt5bAcqHbeRxtdMktGeDEg (67 MB zip → mouse-Geneformer/{config.json,pytorch_model.bin,...})
  • token dictionary MLM-re_token_dictionary_v1.pkl (56084 tokens)
  • gene median dictionary mouse_gene_median_dictionary.pkl
  • already-tokenized eval dataset eval_dataset/cell_type_classification/all_organ_normal_mouse_tokenize_easy_dataset_v-n1.dataset (9 organs: brain, heart, kidney, large_intestine, limb_muscle, mammary_gland, spleen, thymus, tongue; columns: input_ids, cell_type, organ_major, length)

In scope (pipeline-derived, attempted)

  1. Model architecture claim (Methods): 6 transformer blocks, 4 attention heads, 256 embedding dim, SiLU activation, max input 2048. → checked directly against the published checkpoint's config.json. (Verified EXACT before compute.)
  2. Cell-type classification accuracy & macro-F1 (Table 3, mouse): fine-tune the pretrained mouse-Geneformer on each organ's labeled data (drop cell types <0.5%, shuffle seed 42, 80/20 train/eval split, 10 epochs, lr 5e-5 linear, warmup 500, batch 6, fp16 — exactly as cell_classification.ipynb), report eval accuracy + macro-F1. Reproduce a subset of organs present in the shipped "easy" eval set and compare 1:1 to Table 3. This is the central reproducible pipeline output.

Out of scope / not attempted (the hard ~20%)

  • Pretraining mouse-Geneformer on the 20M-cell corpus (8×V100, ~2 days). We use the authors' published checkpoint instead — reproducing pretraining is out of the 80/20 budget.
  • In silico perturbation experiments (disease-gene discovery). Pipeline exists (in_silico_perturbation.ipynb) but interpreting the gene rankings against the paper's qualitative claims is a separate, larger effort.
  • Cross-species (human) fine-tuning (Table 4) and gene-network analyses.
  • Organs only in the "hard"/other eval splits (e.g. prostate, embryo, Kidney-1).

Comparison targets (Table 3, mouse, as reported)

Tongue 94.87/96.13 · Thymus 96.97/96.34 · Mammary 99.02/98.57 · Large Intestine 93.08/92.08 · Limb Muscle 99.52/98.69 · Spleen 98.70/97.27 · Heart 97.82/95.86 · Brain 96.92/88.01 · Kidney 94.88/90.86 (acc% / F1). NB: paper lists "Kidney-1" (79.02) and "Kidney-2" (94.88); the shipped "easy" set has a single "kidney".

Figures / tables: Table
arch_layers
Reported
6
Reproduced
6
exact
arch_heads
Reported
4
Reproduced
4
exact
arch_hidden
Reported
256
Reproduced
256
exact
arch_act
Reported
SiLU
Reproduced
silu
exact
arch_maxinput
Reported
2048
Reproduced
2048
exact
ctc_brain_acc
Reported
96.92
Reproduced
97.83
within tolerance
ctc_brain_macroF1
Reported
88.01
Reproduced
93.97
partial
ctc_large_intestine_acc
Reported
93.08
Reproduced
93.38
exact
ctc_large_intestine_macroF1
Reported
92.08
Reproduced
92.15
exact
ctc_thymus_acc
Reported
96.97
Reproduced
96.87
exact
ctc_thymus_macroF1
Reported
96.34
Reproduced
95.74
within tolerance
ctc_heart_acc
Reported
97.82
Reproduced
97.77
exact
ctc_heart_macroF1
Reported
95.86
Reproduced
95.15
within tolerance
ctc_mammary_gland_acc
Reported
99.02
Reproduced
98.98
exact
ctc_mammary_gland_macroF1
Reported
98.57
Reproduced
98.59
exact
ctc_tongue_acc
Reported
94.87
Reproduced
95.04
exact
ctc_tongue_macroF1
Reported
96.13
Reproduced
96.25
exact
ctc_limb_muscle_acc
Reported
99.52
Reproduced
99.55
exact
ctc_limb_muscle_macroF1
Reported
98.69
Reproduced
99.06
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 95/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

451.7 k
tokens (I/O) · 46.9 M incl. cache
220 min
runtime · 3.41 CPU-h
3.3 GB
peak RAM
3 (2 failed)
HPC jobs
hummel
machine