Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Mammary cell gene expression atlas links epithelial cell remodeling events to breast carcinogenesis.

Commun Biol · 2021
L1 96/100 PQI 96
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
96/100
Reproducibility score
1.2 SD above mean
vs. all fields · 1187 studies
🎯 Scores higher than 91% of all assessed papers rank 92 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce, 1:1. The paper ships NO author analysis code (both scaffold pointers were text-mining false positives that are actually integrated SOURCE datasets: czbiohub/tabula-muris -> Study=TabulaMuris 774 cells; GSE111113/Giraddi -> Study=Giraddi 5821 cells). Per P16 we reproduced using a third-party tool (Scanpy 1.11.5) on the authors' own deposited processed data, the UCSC Cell Browser integrated Seurat_v3 object (50,407 cells; downloaded MD5s match the deposited files). RESULT: all headline pipeline-derived numbers reproduce. C1 50407 cells (~50K). C2 exactly 6 clusters with the identical Basal/L-Hor/L-Alv + 3-progenitor label set, sizes summing to 50407. C3 5 studies / 3 strains / 8 developmental stages all match. C4 (genuine recompute): Scanpy Wilcoxon DE recovers 15 of 16 reported cell-type markers in the top-25 of the matching cluster (only Csn2 ranks low at #97, still positively enriched, depressed by the 2000-HVG integrated feature set). No fabrication indicators. NOT attempted (hard ~20%, see scope.md): de-novo Cell Ranger + Seurat integration from raw FASTQ (stochastic, won't match 1:1), STREAM trajectory, CytoTRACE, scGSVA gene-set scoring, the human atlas, and TCGA cell-of-origin inference.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 96
    assessed: 2026-06-15 ⛓ d7a593f5203c
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The authors tested whether integrating single-cell RNA-seq data across key life-stage windows of susceptibility would comprehensively capture mammary gland reorganization, and whether a resulting consensus lineage trajectory could infer cells of origin for breast cancer and link gland reorganization to risk of specific breast cancer subtypes.

Core claims
  • Integration of five mouse scRNAseq datasets reveals a trifurcating lineage trajectory originating from embryonic mammary stem cells (MaSCs) that differentiates into three epithelial lineages (Basal, L-Alv, L-Hor) via unipotent progenitor clusters finding
  • Progenitor clusters (C1, C3, C5) show significantly higher CytoTRACE stemness scores than their corresponding differentiated clusters (C2, C4, C6) finding
  • Curated lineage-specific gene sets outperform existing MSigDB gene sets, CytoTRACE-derived scores, and basic scRNAseq characteristics in correlating with pseudotime for each lineage state method
  • Mouse mammary lineages correspond to human breast epithelial clusters via label transfer (Basal→B/Myo, L-Alv→L1.1/L1.2, L-Hor→L2), supporting conserved mouse-human epithelial biology finding
  • The human adult breast epithelium lacks a bridging MaSC population/junction cluster present in mouse, indicating absence of true MaSCs postnatally in humans finding
  • 17β-estradiol treatment of surgically menopaused mice re-expands the mammary gland, and this can be combined with progesterone and PBDE exposure to model postmenopausal hormone/endocrine-disruptor effects method
  • A publicly accessible mammary cell gene expression atlas and UMAP-based trajectory tool was constructed and deposited on the UCSC Cell Browser resource
  • Lineage-specific gene sets and scGSVA scoring enable de novo, less computationally intensive projection of new scRNAseq/bulk data (including TCGA breast cancer data) onto the mammary lineage trajectory method
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA sequencing mouse mammary gland (embryonic, neonatal, pubertal, pregnant; public datasets from Giraddi et al., Pal et al., Bach et al., Tabula Muris Consortium) none (developmental/life-stage sampling) transcriptome, cell clustering, lineage marker expression
single-cell RNA sequencing surgically menopaused mouse mammary gland (this study) 17β-estradiol, progesterone, PBDEs (environmental endocrine-disrupting chemicals), or combinations transcriptome, cell clustering
whole-mount mammary gland morphological analysis surgically menopaused mouse mammary gland 17β-estradiol, progesterone, PBDEs, or combinations total duct length, branching points, terminal end bud-like structures
data integration (Seurat v3 anchor-based; also Harmony, LIGER, scAlign) mouse mammary epithelial cells (~50K cells, 5 studies) none (computational integration) UMAP clustering, trifurcation structure robustness Seurat v3
CytoTRACE analysis mouse and human mammary epithelial scRNAseq data none (computational) predicted cell differentiation/stemness score CytoTRACE
pseudotime trajectory inference (STREAM pipeline) mouse and human mammary epithelial scRNAseq data none (computational) pseudotime branch structure, differentiation-specific gene expression STREAM (python)
single-cell gene set variation analysis (scGSVA) mouse and human mammary epithelial scRNAseq data; TCGA breast cancer RNAseq; human breast cancer scRNAseq none (computational) lineage gene set scores (Stem, Basal, Alv, Hor) for UMAP/ternary plot placement scGSVA
single-cell RNA sequencing with canonical component analysis-based label transfer human normal breast epithelium (4 individuals, ~24K cells) none cluster annotation (B/Myo, L1.1/L1.2, L2) and agreement with mouse-derived label transfer
Key results
  • 17β-estradiol treatment re-expanded the mammary gland in menopaused mice, increasing total duct length, branching points, and terminal end bud-like structures
  • Progesterone combined with 17β-estradiol further increased gland branching compared to estradiol alone
  • Simultaneous PBDE exposure tended to show weaker gland regrowth
  • Louvain clustering of integrated mouse data identified 6 clusters (C1-C6) matching 3 progenitor and 3 differentiated states forming a trifurcation
  • Progenitor clusters had significantly higher CytoTRACE scores than corresponding differentiated leaf clusters Cliff's delta 0.72, 0.81, 0.53
  • Curated 'Stem' gene set outperformed all other RNA-based features/algorithms tested (including MSigDB gene sets, CytoTRACE, GCS) in correlation with S5 (Stem) pseudotime
  • Best-performing gene sets used top 160 (Stem), 240 (Basal), 500 (Alv), and 200 (Hor) ranked genes
  • Human breast epithelial clusters lacked a bridging cluster between the three major lineages, unlike mouse data, indicating absence of true MaSCs in adult human breast
Key statistics
  • count ~75K total barcodes/cells (combined mouse scRNAseq datasets before quality filtering)
  • count 50K putative single mammary epithelial cells (high-quality mouse cells retained after filtering, used for integration)
  • count 24,377 cells (human breast epithelial cells from 4 individuals used for integration)
  • other Cliff's delta = 0.72 (CI: 0.70-0.74) (CytoTRACE score comparison between progenitor and differentiated clusters)
  • other Cliff's delta = 0.81 (CI: 0.79-0.82) (CytoTRACE score comparison between progenitor and differentiated clusters)
  • other Cliff's delta = 0.53 (CI: 0.49-0.57) (CytoTRACE score comparison between progenitor and differentiated clusters)
  • count n = 2404, 2393, 2659, 2429, 828, 2249 (cell numbers in clusters C1-C6 respectively used for CytoTRACE comparisons)
  • count MSigDB n = 22,540 gene sets (as of 3-20-2020) (reference gene sets compared against curated lineage gene sets for pseudotime correlation performance)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a single-cell RNA-sequencing atlas study that is largely descriptive and computational rather than hypothesis-test driven. Five mouse (and four human) scRNAseq data sets were integrated (Seurat v3 anchor-based integration, with Harmony, LIGER, and scAlign as cross-checks), clustered (Louvain), and ordered along inferred lineage trajectories (UMAP, STREAM pseudotime, CytoTRACE, scGSVA/GSVA). Where group differences were quantified, effect sizes (Cliff's delta with confidence intervals) and correlation coefficients were reported, and distributions were shown with box plots and LOESS fits with confidence bands.

Replicationmixed Sample sizesample sizes reported as numbers of cells/barcodes per cluster and per integrated data set (e.g., ~75K barcodes, 50K mouse and 24K human epithelial cells); no formal power analysis described Groupsepithelial clusters/lineages (progenitor vs differentiated; stages; mouse vs human); treated menopaused mice (estradiol, progesterone, PBDEs) Pairingunclear Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Cliff's delta effect size (with confidence intervals) for comparing CytoTRACE scores between clusters Fig. 1e, comparisons of progenitor vs. leaf clusters (e.g., C1 vs C2, C3 vs C4, C5 vs C6) per-cell counts stated (e.g., C1 n=2404, C2 n=2393, C3 n=2659, C4 n=2429, C5 n=828, C6 n=2249) na
correlation coefficient (scGSVA scores vs. pseudotime); LOESS regression with confidence intervals Fig. 2c, evaluation of curated gene-set performance against pseudotime and other RNA-based features five studies aggregated; integrated mouse data N=50,407 not stated
correlation between GSVA stem-gene-set scores and CytoTRACE scores (described as significant) Supplementary Fig. 14d, human breast epithelium not stated
difference in CytoTRACE scores described as 'significantly higher' (specific test not stated) progenitor vs. corresponding leaf clusters per-cell counts as above not stated
Approaches that could also have been used
  • Cluster differences in CytoTRACE scores were quantified with Cliff's delta and confidence intervals, with some comparisons described as 'significant.'
    Could also: Reporting the accompanying nonparametric test statistic and p-value (e.g., Mann-Whitney U / Wilcoxon rank-sum) alongside the effect size. — Pairing the effect size with an explicit test statistic and p-value gives readers both the magnitude and the formal inferential basis in one place; the two conventions are complementary.
  • Each pair of clusters was compared individually across several pairings.
    Could also: A single omnibus comparison across clusters (e.g., Kruskal-Wallis) followed by post-hoc pairwise comparisons with a correction such as Benjamini-Hochberg or Dunn's test. — An omnibus-plus-post-hoc framework controls the error rate across the family of cluster comparisons and documents the multiplicity scope explicitly.
  • Because clusters contain large numbers of cells, comparisons are based on per-cell n.
    Could also: Treating the biological replicate (animal/individual) as the unit, e.g., pseudobulk or mixed-effects models that nest cells within samples. — Sample-level modeling distinguishes biological from technical replication and is often preferred to avoid pseudoreplication when many cells come from few individuals; the authors themselves note some stages derive from a single strain.
  • Lineage relationships were inferred primarily from one integration plus three additional algorithms and STREAM/CytoTRACE.
    Could also: Adding quantitative trajectory-uncertainty metrics (e.g., RNA velocity, partition-based graph abstraction, or bootstrap stability of branch assignments). — Explicit uncertainty or stability measures would complement the qualitative agreement across methods and convey confidence in the inferred branch points.
  • Distribution spread was shown with box plots (median, quartiles, 1.5×IQR) and LOESS confidence bands.
    Could also: Supplementing with per-point or violin/jittered displays of the full distribution. — Showing the underlying distribution can convey modality and density for very large cell counts in addition to the summary quartiles.
  • Gene-set performance was summarized by correlation coefficients between scGSVA scores and pseudotime.
    Could also: Reporting correlation confidence intervals and a held-out or cross-validated evaluation of the curated gene sets. — Interval estimates and out-of-sample validation help characterize how stable the gene-set rankings are when the same data inform both selection and evaluation.
Software: Seurat v3 (anchor-based integration, label transfer) v3 · Harmony (integration) · LIGER (integration) · scAlign (integration) · STREAM (Python, pseudotime trajectory) · CytoTRACE / scGSVA (GSVA-based scoring); MSigDB feature reference

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
54
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE103275 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
also used by 1 paper:
GSE19446 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
also used by 1 paper:
GSE75688 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
also used by 1 paper:
10.5281/zenodo.4674274 DOI in References (http://purl.org/orb/References)
no other assessed paper uses this yet
GSE106273 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE111113 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE113197 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE149949 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE16997 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34079055

Paper: Saeki et al. 2021, Mammary cell gene expression atlas links epithelial cell remodeling events to breast carcinogenesis. Commun Biol 4:660. PMID 34079055 · PMCID PMC8172904 · DOI 10.1038/s42003-021-02201-2.

Scaffold pointers were text-mining false positives (corrected)

  • code_url = github.com/czbiohub/tabula-muriswrong as "the paper's code". Tabula Muris is one of the source datasets the atlas integrates (774 cells in the final object, field Study=TabulaMuris), not this paper's analysis code.
  • data_accession = GSE111113wrong as "the paper's own data". GSE111113 is an earlier paper (Giraddi 2018, PMID 30089273); it is another integrated source dataset (Study=Giraddi, 5821 cells).
  • The authors ship NO analysis-code repository (no GitHub/Zenodo in Methods/ Data-availability). Per brief rule P16 this does not force a drop: we reproduce the pipeline-derived results using a standard third-party tool (Scanpy) on the paper's own deposited processed data.

The reproducible asset: the deposited integrated object (UCSC Cell Browser)

https://mouse-mammary-epithelium-integrated.cells.ucsc.edu Four integration methods shipped (Harmony, scAlign, LIGER, Seurat_v3 = primary), each 50,407 cells. Download base (used here): https://cells.ucsc.edu/mouse-mammary-epithelium-integrated/seurat-v3/meta.tsv (per-cell annotation, 4.5 MB) + exprMatrix.tsv.gz (log-norm integrated expression, 437 MB). This is the authors' final pipeline output.

In scope (pipeline-derived, attempted)

The headline computational outputs of the integration + clustering pipeline (Cell Ranger v2 → Seurat v3 integration → Louvain clustering → DE markers):

id reported result how reproduced
C1 ~50 K integrated mammary epithelial cells wc of deposited meta.tsv
C2 6 epithelial clusters: Basal, L-Hor, L-Alv + 3 progenitors (MaSC/B-pro, LH-pro, LA-pro) unique Cluster_annotation
C3 composition: 5 integrated studies, 3 strains, 8 developmental stages value counts of meta fields
C4 cluster marker genes (Basal: Krt14/Acta2/Krt17/Myl9; L-Hor: Areg/Cited1/Ly6d/Prlr; L-Alv: Csn3/Lalba/Csn2/Spp1) Scanpy rank_genes_groups (Wilcoxon) on the deposited matrix; check reported markers rank as top DE genes per matching cluster

Out of scope (the hard ~20%, NOT attempted — and why)

  • Re-running Cell Ranger v2 from raw FASTQ (5 source datasets) and re-deriving the integration de novo: integration is stochastic/parameter-sensitive; exact cell counts will not match 1:1; enormous compute for little added evidence.
  • STREAM trajectory, CytoTRACE potency, scGSVA gene-set scoring, ggtern ternary plots, the human atlas, TCGA cell-of-origin inference: each a separate pipeline; beyond the low-hanging "is the deposited atlas internally consistent with the reported headline numbers, and do its clusters carry the reported markers" goal.

Honest framing

C1–C3 verify that the deposited object is internally consistent with the numbers printed in the paper (a fabrication check: text numbers must be backed by the shipped data). C4 is a genuine recomputation — we run a DE pipeline on the paper's matrix and test whether the reported cell-type markers actually emerge.

Figures / tables: Fig 1
C1
Reported
~50,000 integrated mammary epithelial cells
Reproduced
50407
within tolerance
C2
Reported
6 epithelial clusters: Basal, L-Hor, L-Alv, MaSC/B-pro, LA-pro, LH-pro
Reproduced
6 clusters, identical label set; sizes sum to 50407
exact
C3a
Reported
5 integrated source datasets
Reproduced
5 (Bach, Pal, Giraddi, Saeki, TabulaMuris)
exact
C3b
Reported
3 mouse strains (C57BL/6, FVB, Balb/c)
Reproduced
3 (C57BL/6, FVB, Balb/c)
exact
C3c
Reported
8 developmental stages
Reproduced
8 stages
exact
C4a
Reported
Basal markers Krt14, Acta2, Krt17, Myl9
Reproduced
all top-10 DE (Krt14#1, Acta2#6, Myl9#7, Krt17#8)
exact
C4b
Reported
L-Hor markers Areg, Cited1, Ly6d, Prlr
Reproduced
all top-17 DE (Prlr#2, Ly6d#6, Cited1#7, Areg#17)
exact
C4c
Reported
L-Alv markers Csn3, Lalba, Csn2, Spp1
Reproduced
Csn3#5, Lalba#10, Spp1#12 (top-12); Csn2#97
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 96/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

All reproducible headline results match the authors' deposited integrated object: C1 50407 (~50K), C2 the exact 6-cluster label set, C3 5 studies/3 strains/8 stages, and a genuine Scanpy Wilcoxon recompute recovers 15/16 reported markers in the top-25 (only Csn2 at #97, still positively enriched, depressed by the 2000-HVG export). Deviations are negligible and on the technical/expected side (rounding + feature selection), with no fabrication indicators and provably matching input MD5s. The main caveat is coverage, not correctness: the paper ships no author code, C1–C3 are internal-consistency checks, and the carcinogenesis-link analyses (trajectory/CytoTRACE/scGSVA/TCGA) were out of scope — so the title claim is unrefuted but untested.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

121.8 k
tokens (I/O) · 7.1 M incl. cache
15 min
runtime · 0.02 CPU-h
2 GB
peak RAM
1
HPC jobs
hummel
machine