Mammary cell gene expression atlas links epithelial cell remodeling events to breast carcinogenesis.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce, 1:1. The paper ships NO author analysis code (both scaffold pointers were text-mining false positives that are actually integrated SOURCE datasets: czbiohub/tabula-muris -> Study=TabulaMuris 774 cells; GSE111113/Giraddi -> Study=Giraddi 5821 cells). Per P16 we reproduced using a third-party tool (Scanpy 1.11.5) on the authors' own deposited processed data, the UCSC Cell Browser integrated Seurat_v3 object (50,407 cells; downloaded MD5s match the deposited files). RESULT: all headline pipeline-derived numbers reproduce. C1 50407 cells (~50K). C2 exactly 6 clusters with the identical Basal/L-Hor/L-Alv + 3-progenitor label set, sizes summing to 50407. C3 5 studies / 3 strains / 8 developmental stages all match. C4 (genuine recompute): Scanpy Wilcoxon DE recovers 15 of 16 reported cell-type markers in the top-25 of the matching cluster (only Csn2 ranks low at #97, still positively enriched, depressed by the 2000-HVG integrated feature set). No fabrication indicators. NOT attempted (hard ~20%, see scope.md): de-novo Cell Ranger + Seurat integration from raw FASTQ (stochastic, won't match 1:1), STREAM trajectory, CytoTRACE, scGSVA gene-set scoring, the human atlas, and TCGA cell-of-origin inference.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 96assessed: 2026-06-15 ⛓ d7a593f5203c
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe authors tested the hypothesis that constructing an integrated single-cell RNA-seq atlas covering key windows of susceptibility would comprehensively capture mammary gland reorganization throughout life, and that projecting a consensus lineage trajectory could infer cells of origin for breast carcinogenesis and link gland reorganization to risk of different breast cancer subtypes.
- ★ An integrated 50K mouse and 24K human mammary epithelial cell atlas (scRNA-seq) captures mammary epithelium reorganization across most lifetime stages. resource
- ★ A putative lineage trajectory originates from embryonic mammary stem cells (MaSCs) and differentiates into three epithelial lineages: basal, luminal hormone-sensing, and luminal alveolar. finding
- ★ The three differentiated lineages arise from corresponding unipotent progenitor clusters (B-Pro, LA-Pro, LH-Pro) in postnatal glands. mechanism
- ★ Lineage-specific gene sets infer cells of origin of breast cancer using TCGA and human breast cancer scRNA-seq data and associate gland reorganization with different breast cancer subtypes. finding
- ★ Curated lineage-specific gene sets scored via scGSVA enable de novo reconstruction of the mammary trajectory and outperform existing RNA-based features and algorithms. method
- ★ Mouse and human mammary epithelial lineages are largely conserved: Basal, L-Alv, L-Hor correspond to human B/Myo, L1.1/L1.2, and L2 clusters. finding
- 17β-estradiol re-expands the gland in menopaused mice, progesterone increases branching, and PBDE co-exposure tends to weaken regrowth. finding
- Human adult breast lacks true MaSCs, shown by absence of a bridging cluster between the three major epithelial clusters. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA sequencing (scRNA-seq) | surgically menopaused mouse mammary gland (C57BL/6/FVB/Balb/c strains) | 17β-estradiol, progesterone, PBDEs, or combinations | single-cell gene expression / epithelial cell transcriptomes | — |
| scRNA-seq data integration (Seurat v3 anchor-based) | mouse mammary epithelium, 5 datasets, 8 life stages, 3 strains (~50K cells) | none (integration of public + new data) | UMAP clusters / lineage trajectory | Seurat v3 |
| mammary gland morphology/whole-mount analysis | menopaused mouse mammary gland | 17β-estradiol, progesterone, PBDE treatment | total duct length, branching points, terminal end bud-like structures | — |
| trajectory inference (pseudotime) | integrated mouse mammary epithelium scRNA-seq | none | pseudotime lineage trajectory, branch-specific gene expression | STREAM python pipeline |
| differentiation-state inference (CytoTRACE) | integrated mouse mammary epithelial clusters | none | CytoTRACE stemness scores per cell/cluster | CytoTRACE |
| single-cell gene set variation analysis (scGSVA) | mouse and human mammary epithelial scRNA-seq | none | lineage gene set scores correlated to pseudotime | scGSVA / MSigDB (n=22,540 gene sets) |
| cross-species label transfer (CCA anchor-based) | human normal breast epithelium scRNA-seq, 4 individuals (~24K cells) | none | transferred lineage cluster annotations | Seurat v3 |
| bulk RNA-seq gene set scoring | TCGA breast cancer RNA-seq and human breast cancer scRNA-seq | none | lineage gene set scores / ternary cell-of-origin placement | — |
- – Integrated mouse data formed three major clusters connected by a bridging population in a trifurcation shape, supporting a trajectory from embryonic MaSCs to basal, L-Alv, and L-Hor lineages.
- ▲ Putative progenitor clusters had significantly higher CytoTRACE (stemness) scores than corresponding leaf clusters; Cliff's delta C1 vs C3 = 0.72 (CI 0.70–0.74), C3 vs C4 = 0.81 (CI 0.79–0.82), C5 vs C6 = 0.53 (CI 0.49–0.57). Cliff's delta 0.72, 0.81, 0.53
- ▲ Curated 'Stem' gene set outperformed all other MSigDB features and algorithms in correlation with S5 (Stem) pseudotime.
- – Best-performing gene sets comprised top 160 (Stem), 240 (Basal), 500 (Alv), and 200 (Hor) genes. 160/240/500/200 genes
- – Mouse lineages Basal, L-Alv, L-Hor largely corresponded to human B/Myo, L1.1/L1.2, and L2 clusters via label transfer.
- – 17β-estradiol re-expanded the gland with increased duct length, branching points, and terminal end bud-like structures; progesterone increased branching; PBDEs tended to weaken regrowth.
- – Human adult breast showed no bridging cluster between three major clusters, indicating absence of true MaSCs.
- ▲ Human stem gene set GSVA scores showed significant correlation with unbiased CytoTRACE scores.
- correlation Cliff's delta 0.81 (CI 0.79–0.82) (CytoTRACE score difference C3 vs C4 (LA-Pro vs L-Alv))
- correlation Cliff's delta 0.72 (CI 0.70–0.74) (CytoTRACE score difference C1 vs C3)
- correlation Cliff's delta 0.53 (CI 0.49–0.57) (CytoTRACE score difference C5 vs C6 (LH-Pro vs L-Hor))
- count ~50,407 mouse epithelial cells integrated (50K) (high-quality single mammary epithelial cells from 5 studies)
- count 24,377 human breast epithelial cells (24K) (integrated from 4 individuals)
- count ~75K total barcodes (raw cells across 5 datasets before QC filtering)
- count MSigDB n = 22,540 gene sets (as of 3-20-2020) (RNA-based features compared against curated gene sets)
- count CytoTRACE cluster sizes: C1 n=2404, C2 n=2393, C3 n=2659, C4 n=2429, C5 n=828, C6 n=2249 (biologically independent cells per cluster compared)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a single-cell RNA-sequencing atlas study that is largely descriptive and computational rather than hypothesis-test driven. Five mouse (and four human) scRNAseq data sets were integrated (Seurat v3 anchor-based integration, with Harmony, LIGER, and scAlign as cross-checks), clustered (Louvain), and ordered along inferred lineage trajectories (UMAP, STREAM pseudotime, CytoTRACE, scGSVA/GSVA). Where group differences were quantified, effect sizes (Cliff's delta with confidence intervals) and correlation coefficients were reported, and distributions were shown with box plots and LOESS fits with confidence bands.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Cliff's delta effect size (with confidence intervals) for comparing CytoTRACE scores between clusters | Fig. 1e, comparisons of progenitor vs. leaf clusters (e.g., C1 vs C2, C3 vs C4, C5 vs C6) | per-cell counts stated (e.g., C1 n=2404, C2 n=2393, C3 n=2659, C4 n=2429, C5 n=828, C6 n=2249) | na |
| correlation coefficient (scGSVA scores vs. pseudotime); LOESS regression with confidence intervals | Fig. 2c, evaluation of curated gene-set performance against pseudotime and other RNA-based features | five studies aggregated; integrated mouse data N=50,407 | not stated |
| correlation between GSVA stem-gene-set scores and CytoTRACE scores (described as significant) | Supplementary Fig. 14d, human breast epithelium | — | not stated |
| difference in CytoTRACE scores described as 'significantly higher' (specific test not stated) | progenitor vs. corresponding leaf clusters | per-cell counts as above | not stated |
-
Cluster differences in CytoTRACE scores were quantified with Cliff's delta and confidence intervals, with some comparisons described as 'significant.'↳ Could also: Reporting the accompanying nonparametric test statistic and p-value (e.g., Mann-Whitney U / Wilcoxon rank-sum) alongside the effect size. — Pairing the effect size with an explicit test statistic and p-value gives readers both the magnitude and the formal inferential basis in one place; the two conventions are complementary.
-
Each pair of clusters was compared individually across several pairings.↳ Could also: A single omnibus comparison across clusters (e.g., Kruskal-Wallis) followed by post-hoc pairwise comparisons with a correction such as Benjamini-Hochberg or Dunn's test. — An omnibus-plus-post-hoc framework controls the error rate across the family of cluster comparisons and documents the multiplicity scope explicitly.
-
Because clusters contain large numbers of cells, comparisons are based on per-cell n.↳ Could also: Treating the biological replicate (animal/individual) as the unit, e.g., pseudobulk or mixed-effects models that nest cells within samples. — Sample-level modeling distinguishes biological from technical replication and is often preferred to avoid pseudoreplication when many cells come from few individuals; the authors themselves note some stages derive from a single strain.
-
Lineage relationships were inferred primarily from one integration plus three additional algorithms and STREAM/CytoTRACE.↳ Could also: Adding quantitative trajectory-uncertainty metrics (e.g., RNA velocity, partition-based graph abstraction, or bootstrap stability of branch assignments). — Explicit uncertainty or stability measures would complement the qualitative agreement across methods and convey confidence in the inferred branch points.
-
Distribution spread was shown with box plots (median, quartiles, 1.5×IQR) and LOESS confidence bands.↳ Could also: Supplementing with per-point or violin/jittered displays of the full distribution. — Showing the underlying distribution can convey modality and density for very large cell counts in addition to the summary quartiles.
-
Gene-set performance was summarized by correlation coefficients between scGSVA scores and pseudotime.↳ Could also: Reporting correlation confidence intervals and a held-out or cross-validated evaluation of the curated gene sets. — Interval estimates and out-of-sample validation help characterize how stable the gene-set rankings are when the same data inform both selection and evaluation.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
17β-estradiol increases mammary duct length, branching points, and terminal end bud formation in menopaused mice, progesterone increases branching, and PBDE exposure attenuates hormone-driven regrowth.imaging mouse mammary gland mixed 2021×1papers★ This paper is the founder (earliest)
-
Cross-species label transfer maps mouse basal, luminal-alveolar, and luminal-hormone-sensing lineages to human B/Myo, L1.1/L1.2, and L2 clusters respectively.scRNA-seq human breast epithelium 2021×1papers★ This paper is the founder (earliest)
-
Human adult breast epithelium lacks a bridging progenitor cluster between the three major lineage clusters, indicating absence of true multipotent MaSCs in adult human breast.scRNA-seq human breast epithelium none 2021×1papers★ This paper is the founder (earliest)
-
Human mammary stem gene set GSVA scores significantly correlate with unbiased CytoTRACE stemness scores in human breast epithelium, validating cross-species gene set transfer.scRNA-seq human breast epithelium up 2021×1papers★ This paper is the founder (earliest)
-
Integrated mouse mammary epithelial scRNA-seq reveals a trifurcating trajectory from embryonic MaSCs to basal, luminal-alveolar, and luminal-hormone-sensing lineages connected by a bridging progenitor population.scRNA-seq mouse mammary epithelium 2021×1papers★ This paper is the founder (earliest)
-
A curated mammary stem gene set score shows the highest correlation with stem-lineage pseudotime among all MSigDB gene sets and competing algorithms tested.scRNA-seq mouse mammary epithelium up 2021×1papers★ This paper is the founder (earliest)
-
Putative progenitor clusters have significantly higher CytoTRACE stemness scores than their differentiated leaf-cluster counterparts in mouse mammary epithelium.scRNA-seq mouse mammary epithelium up 2021×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34079055
Paper: Saeki et al. 2021, Mammary cell gene expression atlas links epithelial cell remodeling events to breast carcinogenesis. Commun Biol 4:660. PMID 34079055 · PMCID PMC8172904 · DOI 10.1038/s42003-021-02201-2.
Scaffold pointers were text-mining false positives (corrected)
code_url=github.com/czbiohub/tabula-muris→ wrong as "the paper's code". Tabula Muris is one of the source datasets the atlas integrates (774 cells in the final object, fieldStudy=TabulaMuris), not this paper's analysis code.data_accession=GSE111113→ wrong as "the paper's own data". GSE111113 is an earlier paper (Giraddi 2018, PMID 30089273); it is another integrated source dataset (Study=Giraddi, 5821 cells).- The authors ship NO analysis-code repository (no GitHub/Zenodo in Methods/ Data-availability). Per brief rule P16 this does not force a drop: we reproduce the pipeline-derived results using a standard third-party tool (Scanpy) on the paper's own deposited processed data.
The reproducible asset: the deposited integrated object (UCSC Cell Browser)
https://mouse-mammary-epithelium-integrated.cells.ucsc.edu
Four integration methods shipped (Harmony, scAlign, LIGER, Seurat_v3 = primary),
each 50,407 cells. Download base (used here):
https://cells.ucsc.edu/mouse-mammary-epithelium-integrated/seurat-v3/
→ meta.tsv (per-cell annotation, 4.5 MB) + exprMatrix.tsv.gz (log-norm
integrated expression, 437 MB). This is the authors' final pipeline output.
In scope (pipeline-derived, attempted)
The headline computational outputs of the integration + clustering pipeline (Cell Ranger v2 → Seurat v3 integration → Louvain clustering → DE markers):
| id | reported result | how reproduced |
|---|---|---|
| C1 | ~50 K integrated mammary epithelial cells | wc of deposited meta.tsv |
| C2 | 6 epithelial clusters: Basal, L-Hor, L-Alv + 3 progenitors (MaSC/B-pro, LH-pro, LA-pro) | unique Cluster_annotation |
| C3 | composition: 5 integrated studies, 3 strains, 8 developmental stages | value counts of meta fields |
| C4 | cluster marker genes (Basal: Krt14/Acta2/Krt17/Myl9; L-Hor: Areg/Cited1/Ly6d/Prlr; L-Alv: Csn3/Lalba/Csn2/Spp1) | Scanpy rank_genes_groups (Wilcoxon) on the deposited matrix; check reported markers rank as top DE genes per matching cluster |
Out of scope (the hard ~20%, NOT attempted — and why)
- Re-running Cell Ranger v2 from raw FASTQ (5 source datasets) and re-deriving the integration de novo: integration is stochastic/parameter-sensitive; exact cell counts will not match 1:1; enormous compute for little added evidence.
- STREAM trajectory, CytoTRACE potency, scGSVA gene-set scoring, ggtern ternary plots, the human atlas, TCGA cell-of-origin inference: each a separate pipeline; beyond the low-hanging "is the deposited atlas internally consistent with the reported headline numbers, and do its clusters carry the reported markers" goal.
Honest framing
C1–C3 verify that the deposited object is internally consistent with the numbers printed in the paper (a fabrication check: text numbers must be backed by the shipped data). C4 is a genuine recomputation — we run a DE pipeline on the paper's matrix and test whether the reported cell-type markers actually emerge.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
All reproducible headline results match the authors' deposited integrated object: C1 50407 (~50K), C2 the exact 6-cluster label set, C3 5 studies/3 strains/8 stages, and a genuine Scanpy Wilcoxon recompute recovers 15/16 reported markers in the top-25 (only Csn2 at #97, still positively enriched, depressed by the 2000-HVG export). Deviations are negligible and on the technical/expected side (rounding + feature selection), with no fabrication indicators and provably matching input MD5s. The main caveat is coverage, not correctness: the paper ships no author code, C1–C3 are internal-consistency checks, and the carcinogenesis-link analyses (trajectory/CytoTRACE/scGSVA/TCGA) were out of scope — so the title claim is unrefuted but untested.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.