Mammary cell gene expression atlas links epithelial cell remodeling events to breast carcinogenesis.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce, 1:1. The paper ships NO author analysis code (both scaffold pointers were text-mining false positives that are actually integrated SOURCE datasets: czbiohub/tabula-muris -> Study=TabulaMuris 774 cells; GSE111113/Giraddi -> Study=Giraddi 5821 cells). Per P16 we reproduced using a third-party tool (Scanpy 1.11.5) on the authors' own deposited processed data, the UCSC Cell Browser integrated Seurat_v3 object (50,407 cells; downloaded MD5s match the deposited files). RESULT: all headline pipeline-derived numbers reproduce. C1 50407 cells (~50K). C2 exactly 6 clusters with the identical Basal/L-Hor/L-Alv + 3-progenitor label set, sizes summing to 50407. C3 5 studies / 3 strains / 8 developmental stages all match. C4 (genuine recompute): Scanpy Wilcoxon DE recovers 15 of 16 reported cell-type markers in the top-25 of the matching cluster (only Csn2 ranks low at #97, still positively enriched, depressed by the 2000-HVG integrated feature set). No fabrication indicators. NOT attempted (hard ~20%, see scope.md): de-novo Cell Ranger + Seurat integration from raw FASTQ (stochastic, won't match 1:1), STREAM trajectory, CytoTRACE, scGSVA gene-set scoring, the human atlas, and TCGA cell-of-origin inference.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 96assessed: 2026-06-15 ⛓ d7a593f5203c
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe authors tested whether integrating single-cell RNA-seq data across key life-stage windows of susceptibility would comprehensively capture mammary gland reorganization, and whether a resulting consensus lineage trajectory could infer cells of origin for breast cancer and link gland reorganization to risk of specific breast cancer subtypes.
- ★ Integration of five mouse scRNAseq datasets reveals a trifurcating lineage trajectory originating from embryonic mammary stem cells (MaSCs) that differentiates into three epithelial lineages (Basal, L-Alv, L-Hor) via unipotent progenitor clusters finding
- ★ Progenitor clusters (C1, C3, C5) show significantly higher CytoTRACE stemness scores than their corresponding differentiated clusters (C2, C4, C6) finding
- ★ Curated lineage-specific gene sets outperform existing MSigDB gene sets, CytoTRACE-derived scores, and basic scRNAseq characteristics in correlating with pseudotime for each lineage state method
- ★ Mouse mammary lineages correspond to human breast epithelial clusters via label transfer (Basal→B/Myo, L-Alv→L1.1/L1.2, L-Hor→L2), supporting conserved mouse-human epithelial biology finding
- ★ The human adult breast epithelium lacks a bridging MaSC population/junction cluster present in mouse, indicating absence of true MaSCs postnatally in humans finding
- 17β-estradiol treatment of surgically menopaused mice re-expands the mammary gland, and this can be combined with progesterone and PBDE exposure to model postmenopausal hormone/endocrine-disruptor effects method
- ★ A publicly accessible mammary cell gene expression atlas and UMAP-based trajectory tool was constructed and deposited on the UCSC Cell Browser resource
- ★ Lineage-specific gene sets and scGSVA scoring enable de novo, less computationally intensive projection of new scRNAseq/bulk data (including TCGA breast cancer data) onto the mammary lineage trajectory method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA sequencing | mouse mammary gland (embryonic, neonatal, pubertal, pregnant; public datasets from Giraddi et al., Pal et al., Bach et al., Tabula Muris Consortium) | none (developmental/life-stage sampling) | transcriptome, cell clustering, lineage marker expression | — |
| single-cell RNA sequencing | surgically menopaused mouse mammary gland (this study) | 17β-estradiol, progesterone, PBDEs (environmental endocrine-disrupting chemicals), or combinations | transcriptome, cell clustering | — |
| whole-mount mammary gland morphological analysis | surgically menopaused mouse mammary gland | 17β-estradiol, progesterone, PBDEs, or combinations | total duct length, branching points, terminal end bud-like structures | — |
| data integration (Seurat v3 anchor-based; also Harmony, LIGER, scAlign) | mouse mammary epithelial cells (~50K cells, 5 studies) | none (computational integration) | UMAP clustering, trifurcation structure robustness | Seurat v3 |
| CytoTRACE analysis | mouse and human mammary epithelial scRNAseq data | none (computational) | predicted cell differentiation/stemness score | CytoTRACE |
| pseudotime trajectory inference (STREAM pipeline) | mouse and human mammary epithelial scRNAseq data | none (computational) | pseudotime branch structure, differentiation-specific gene expression | STREAM (python) |
| single-cell gene set variation analysis (scGSVA) | mouse and human mammary epithelial scRNAseq data; TCGA breast cancer RNAseq; human breast cancer scRNAseq | none (computational) | lineage gene set scores (Stem, Basal, Alv, Hor) for UMAP/ternary plot placement | scGSVA |
| single-cell RNA sequencing with canonical component analysis-based label transfer | human normal breast epithelium (4 individuals, ~24K cells) | none | cluster annotation (B/Myo, L1.1/L1.2, L2) and agreement with mouse-derived label transfer | — |
- ▲ 17β-estradiol treatment re-expanded the mammary gland in menopaused mice, increasing total duct length, branching points, and terminal end bud-like structures
- ▲ Progesterone combined with 17β-estradiol further increased gland branching compared to estradiol alone
- ▼ Simultaneous PBDE exposure tended to show weaker gland regrowth
- – Louvain clustering of integrated mouse data identified 6 clusters (C1-C6) matching 3 progenitor and 3 differentiated states forming a trifurcation
- ▲ Progenitor clusters had significantly higher CytoTRACE scores than corresponding differentiated leaf clusters Cliff's delta 0.72, 0.81, 0.53
- ▲ Curated 'Stem' gene set outperformed all other RNA-based features/algorithms tested (including MSigDB gene sets, CytoTRACE, GCS) in correlation with S5 (Stem) pseudotime
- – Best-performing gene sets used top 160 (Stem), 240 (Basal), 500 (Alv), and 200 (Hor) ranked genes
- – Human breast epithelial clusters lacked a bridging cluster between the three major lineages, unlike mouse data, indicating absence of true MaSCs in adult human breast
- count ~75K total barcodes/cells (combined mouse scRNAseq datasets before quality filtering)
- count 50K putative single mammary epithelial cells (high-quality mouse cells retained after filtering, used for integration)
- count 24,377 cells (human breast epithelial cells from 4 individuals used for integration)
- other Cliff's delta = 0.72 (CI: 0.70-0.74) (CytoTRACE score comparison between progenitor and differentiated clusters)
- other Cliff's delta = 0.81 (CI: 0.79-0.82) (CytoTRACE score comparison between progenitor and differentiated clusters)
- other Cliff's delta = 0.53 (CI: 0.49-0.57) (CytoTRACE score comparison between progenitor and differentiated clusters)
- count n = 2404, 2393, 2659, 2429, 828, 2249 (cell numbers in clusters C1-C6 respectively used for CytoTRACE comparisons)
- count MSigDB n = 22,540 gene sets (as of 3-20-2020) (reference gene sets compared against curated lineage gene sets for pseudotime correlation performance)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a single-cell RNA-sequencing atlas study that is largely descriptive and computational rather than hypothesis-test driven. Five mouse (and four human) scRNAseq data sets were integrated (Seurat v3 anchor-based integration, with Harmony, LIGER, and scAlign as cross-checks), clustered (Louvain), and ordered along inferred lineage trajectories (UMAP, STREAM pseudotime, CytoTRACE, scGSVA/GSVA). Where group differences were quantified, effect sizes (Cliff's delta with confidence intervals) and correlation coefficients were reported, and distributions were shown with box plots and LOESS fits with confidence bands.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Cliff's delta effect size (with confidence intervals) for comparing CytoTRACE scores between clusters | Fig. 1e, comparisons of progenitor vs. leaf clusters (e.g., C1 vs C2, C3 vs C4, C5 vs C6) | per-cell counts stated (e.g., C1 n=2404, C2 n=2393, C3 n=2659, C4 n=2429, C5 n=828, C6 n=2249) | na |
| correlation coefficient (scGSVA scores vs. pseudotime); LOESS regression with confidence intervals | Fig. 2c, evaluation of curated gene-set performance against pseudotime and other RNA-based features | five studies aggregated; integrated mouse data N=50,407 | not stated |
| correlation between GSVA stem-gene-set scores and CytoTRACE scores (described as significant) | Supplementary Fig. 14d, human breast epithelium | — | not stated |
| difference in CytoTRACE scores described as 'significantly higher' (specific test not stated) | progenitor vs. corresponding leaf clusters | per-cell counts as above | not stated |
-
Cluster differences in CytoTRACE scores were quantified with Cliff's delta and confidence intervals, with some comparisons described as 'significant.'↳ Could also: Reporting the accompanying nonparametric test statistic and p-value (e.g., Mann-Whitney U / Wilcoxon rank-sum) alongside the effect size. — Pairing the effect size with an explicit test statistic and p-value gives readers both the magnitude and the formal inferential basis in one place; the two conventions are complementary.
-
Each pair of clusters was compared individually across several pairings.↳ Could also: A single omnibus comparison across clusters (e.g., Kruskal-Wallis) followed by post-hoc pairwise comparisons with a correction such as Benjamini-Hochberg or Dunn's test. — An omnibus-plus-post-hoc framework controls the error rate across the family of cluster comparisons and documents the multiplicity scope explicitly.
-
Because clusters contain large numbers of cells, comparisons are based on per-cell n.↳ Could also: Treating the biological replicate (animal/individual) as the unit, e.g., pseudobulk or mixed-effects models that nest cells within samples. — Sample-level modeling distinguishes biological from technical replication and is often preferred to avoid pseudoreplication when many cells come from few individuals; the authors themselves note some stages derive from a single strain.
-
Lineage relationships were inferred primarily from one integration plus three additional algorithms and STREAM/CytoTRACE.↳ Could also: Adding quantitative trajectory-uncertainty metrics (e.g., RNA velocity, partition-based graph abstraction, or bootstrap stability of branch assignments). — Explicit uncertainty or stability measures would complement the qualitative agreement across methods and convey confidence in the inferred branch points.
-
Distribution spread was shown with box plots (median, quartiles, 1.5×IQR) and LOESS confidence bands.↳ Could also: Supplementing with per-point or violin/jittered displays of the full distribution. — Showing the underlying distribution can convey modality and density for very large cell counts in addition to the summary quartiles.
-
Gene-set performance was summarized by correlation coefficients between scGSVA scores and pseudotime.↳ Could also: Reporting correlation confidence intervals and a held-out or cross-validated evaluation of the curated gene sets. — Interval estimates and out-of-sample validation help characterize how stable the gene-set rankings are when the same data inform both selection and evaluation.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
17β-estradiol increases mammary duct length, branching points, and terminal end bud formation in menopaused mice, progesterone increases branching, and PBDE exposure attenuates hormone-driven regrowth.imaging mouse mammary gland mixed 2021×1papers★ This paper is the founder (earliest)
-
Cross-species label transfer maps mouse basal, luminal-alveolar, and luminal-hormone-sensing lineages to human B/Myo, L1.1/L1.2, and L2 clusters respectively.scRNA-seq human breast epithelium 2021×1papers★ This paper is the founder (earliest)
-
Human adult breast epithelium lacks a bridging progenitor cluster between the three major lineage clusters, indicating absence of true multipotent MaSCs in adult human breast.scRNA-seq human breast epithelium none 2021×1papers★ This paper is the founder (earliest)
-
Human mammary stem gene set GSVA scores significantly correlate with unbiased CytoTRACE stemness scores in human breast epithelium, validating cross-species gene set transfer.scRNA-seq human breast epithelium up 2021×1papers★ This paper is the founder (earliest)
-
Integrated mouse mammary epithelial scRNA-seq reveals a trifurcating trajectory from embryonic MaSCs to basal, luminal-alveolar, and luminal-hormone-sensing lineages connected by a bridging progenitor population.scRNA-seq mouse mammary epithelium 2021×1papers★ This paper is the founder (earliest)
-
A curated mammary stem gene set score shows the highest correlation with stem-lineage pseudotime among all MSigDB gene sets and competing algorithms tested.scRNA-seq mouse mammary epithelium up 2021×1papers★ This paper is the founder (earliest)
-
Putative progenitor clusters have significantly higher CytoTRACE stemness scores than their differentiated leaf-cluster counterparts in mouse mammary epithelium.scRNA-seq mouse mammary epithelium up 2021×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34079055
Paper: Saeki et al. 2021, Mammary cell gene expression atlas links epithelial cell remodeling events to breast carcinogenesis. Commun Biol 4:660. PMID 34079055 · PMCID PMC8172904 · DOI 10.1038/s42003-021-02201-2.
Scaffold pointers were text-mining false positives (corrected)
code_url=github.com/czbiohub/tabula-muris→ wrong as "the paper's code". Tabula Muris is one of the source datasets the atlas integrates (774 cells in the final object, fieldStudy=TabulaMuris), not this paper's analysis code.data_accession=GSE111113→ wrong as "the paper's own data". GSE111113 is an earlier paper (Giraddi 2018, PMID 30089273); it is another integrated source dataset (Study=Giraddi, 5821 cells).- The authors ship NO analysis-code repository (no GitHub/Zenodo in Methods/ Data-availability). Per brief rule P16 this does not force a drop: we reproduce the pipeline-derived results using a standard third-party tool (Scanpy) on the paper's own deposited processed data.
The reproducible asset: the deposited integrated object (UCSC Cell Browser)
https://mouse-mammary-epithelium-integrated.cells.ucsc.edu
Four integration methods shipped (Harmony, scAlign, LIGER, Seurat_v3 = primary),
each 50,407 cells. Download base (used here):
https://cells.ucsc.edu/mouse-mammary-epithelium-integrated/seurat-v3/
→ meta.tsv (per-cell annotation, 4.5 MB) + exprMatrix.tsv.gz (log-norm
integrated expression, 437 MB). This is the authors' final pipeline output.
In scope (pipeline-derived, attempted)
The headline computational outputs of the integration + clustering pipeline (Cell Ranger v2 → Seurat v3 integration → Louvain clustering → DE markers):
| id | reported result | how reproduced |
|---|---|---|
| C1 | ~50 K integrated mammary epithelial cells | wc of deposited meta.tsv |
| C2 | 6 epithelial clusters: Basal, L-Hor, L-Alv + 3 progenitors (MaSC/B-pro, LH-pro, LA-pro) | unique Cluster_annotation |
| C3 | composition: 5 integrated studies, 3 strains, 8 developmental stages | value counts of meta fields |
| C4 | cluster marker genes (Basal: Krt14/Acta2/Krt17/Myl9; L-Hor: Areg/Cited1/Ly6d/Prlr; L-Alv: Csn3/Lalba/Csn2/Spp1) | Scanpy rank_genes_groups (Wilcoxon) on the deposited matrix; check reported markers rank as top DE genes per matching cluster |
Out of scope (the hard ~20%, NOT attempted — and why)
- Re-running Cell Ranger v2 from raw FASTQ (5 source datasets) and re-deriving the integration de novo: integration is stochastic/parameter-sensitive; exact cell counts will not match 1:1; enormous compute for little added evidence.
- STREAM trajectory, CytoTRACE potency, scGSVA gene-set scoring, ggtern ternary plots, the human atlas, TCGA cell-of-origin inference: each a separate pipeline; beyond the low-hanging "is the deposited atlas internally consistent with the reported headline numbers, and do its clusters carry the reported markers" goal.
Honest framing
C1–C3 verify that the deposited object is internally consistent with the numbers printed in the paper (a fabrication check: text numbers must be backed by the shipped data). C4 is a genuine recomputation — we run a DE pipeline on the paper's matrix and test whether the reported cell-type markers actually emerge.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
All reproducible headline results match the authors' deposited integrated object: C1 50407 (~50K), C2 the exact 6-cluster label set, C3 5 studies/3 strains/8 stages, and a genuine Scanpy Wilcoxon recompute recovers 15/16 reported markers in the top-25 (only Csn2 at #97, still positively enriched, depressed by the 2000-HVG export). Deviations are negligible and on the technical/expected side (rounding + feature selection), with no fabrication indicators and provably matching input MD5s. The main caveat is coverage, not correctness: the paper ships no author code, C1–C3 are internal-consistency checks, and the carcinogenesis-link analyses (trajectory/CytoTRACE/scGSVA/TCGA) were out of scope — so the title claim is unrefuted but untested.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.