Distinct senotypes in p16- and p21-positive cells across human and mouse aging tissues.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough only PARTIALLY. The authors' repo (SaulLab_code @ 0ba9634) is NOT a runnable pipeline: just figure-plotting snippets (P16_vs_P21) over un-shipped intermediate Seurat objects, with zero QC/normalization/integration/clustering/labelling code. So the headline figures (Fig.1 tSNE/DotPlot/velocity, Fig.4 SenMayo radars, Fig.5 14-dataset compendium) are NOT independently re-derivable from shipped artifacts. However the repo DOES ship the self-contained per-cell p16/p21 classification logic (Cd45/Mki67/S-phase gate -> Cdkn2a/Cdkn1a positivity). Using a standard third-party tool (scanpy) on the paper's OWN Fig.1 data (GSE161340 single-cell OLD brain), this 1:1 reproduces the paper's stated claim that p21+ cells dominate the senescent population in aged brain (p21+ 8.9-12x p16+), with p16+ and dPo cells rare. Graded 'partial' because the claims are qualitative in the paper (no printed proportions to compare numerically) and only Fig.1/GSE161340 of 14 datasets was attempted (80/20). NOT attempted: exact embeddings/labelling, RNA velocity, the other 13 tissue datasets, SASP/secretome overlap counts, SenMayo radars, SCENIC -- the hard ~20% requiring un-shipped intermediates.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-14 ⛓ d61d4f528df4
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether p16Ink4a-positive and p21Cip1-positive senescent cells represent distinct or overlapping populations with tissue-specific SASP profiles across mouse and human aging tissues.
- ★ p16+ and p21+ senescent cells represent largely distinct, non-overlapping populations across murine and human aging tissues finding
- ★ p16+ and p21+ cells follow independent RNA velocity and pseudotime trajectories with no evidence of a common ancestor or direct transition between states finding
- ★ p16+ cells display heterogeneous, tissue-specific secretomes, whereas p21+ cells exhibit broader but more conserved SASP profiles finding
- ★ Only a small set of 'core' SASP factors (e.g., ICAM1, IGFBP4/6, CXCL16, PLAUR) are shared across tissues and species finding
- ★ The separation of p16+ and p21+ cells was confirmed at the protein level using CyTOF in mouse bone finding
- The SenMayo gene set was used as a curated SASP-related gene resource for cross-dataset analysis resource
- p21+ cell populations are consistently more abundant than p16+ populations across skeletal muscle, bone, and liver finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| scRNA-seq | murine hippocampus (young 4m vs old 24m) | aging | p16/p21 expression, cell composition, SASP gene expression | GSE161340 |
| scRNA-seq | whole mouse brain, excluding hindbrain (8 young, 8 old) | aging | p16/p21 overlap and secretory profile validation | — |
| scRNA-seq | murine skeletal muscle (young vs old) | aging | cell population identity, p16/p21 expression, SASP profile | GSE172410 |
| scRNA-seq | murine bone (young vs old) | aging | cell population identity, p16/p21 expression, SASP profile | GSE128423 |
| scRNA-seq | murine liver (young vs old) | aging | cell population identity, p16/p21 expression, SASP profile | GSE166504 |
| CyTOF (mass cytometry) | murine bone (young vs aged) | aging | protein-level p16/p21 co-expression and secretory proteins (e.g., Serpine1, IL-6) | CyTOF |
| scRNA-seq | murine spleen and kidney | aging | p16/p21 population overlap and secretome | Calico murine aging cell atlas |
| RNA velocity / pseudotime analysis (scVelo, PhyloVelo, Monocle3) | murine hippocampus and other tissues | none | developmental trajectory relationships between p16+ and p21+ cells | — |
- – Only Cxcl16 and Plaur SASP genes (SenMayo set) were expressed in both p16+ and p21+ hippocampal subpopulations
- – Using an expanded 1989-gene secreted protein list plus SenMayo, only Col8a2, Plaur, Cxcl16 and Fbln5 were shared between p21+ and p16+ cells
- – No common ancestor or direct lineage progression was found between p16+ and p21+ cells across brain, muscle, bone, and liver via RNA velocity/pseudotime
- – Standard error of SASP factor expression was substantially higher in p16+ than p21+ cells in the hippocampus
- ▲ p21+ cells expressed a broader range of SenMayo SASP genes than p16+ cells across brain, muscle, bone, and liver
- – At the protein level, Serpine1 was exclusive to p16+ cells and IL-6 exclusive to p21+ cells in bone CyTOF data
- – p16 expression increased with age across all hippocampal clusters, whereas p21 upregulation was restricted to specific clusters
- count 9, 14, and 13 distinct cell populations identified in skeletal muscle, bone, and liver respectively (cell type diversity per tissue)
- count 4 shared genes (Col8a2, Plaur, Cxcl16, Fbln5) (overlap between p21+ and p16+ secretomes using expanded 1989-gene secreted protein list)
- count 2 shared genes (Cxcl16, Plaur) (overlap between p16+ and p21+ SenMayo SASP genes in hippocampus)
- count 1989 secreted protein genes (Human Protein Atlas secreted gene list used for expanded secretome analysis)
- count 8 young and 8 old whole mouse brains (independent brain validation dataset (Ximerakis et al, 2019))
- count young (4 months) vs old (24 months) mice (ages compared in hippocampal scRNA-seq dataset)
- count 17 cell types distinguished in murine bone (Baryawno et al 2019 bone dataset, GSE128423)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This computational resource paper performed comparative single-cell RNA sequencing (scRNA-seq) analyses of p16+ and p21+ cell populations across multiple murine (brain, skeletal muscle, bone, liver, spleen, kidney) and human (skin, lung) aging tissues using publicly available datasets. Cells were filtered to exclude proliferating (Ki67+, S-phase) and high-CD45 immune cells before defining p16+, p21+, and double-positive subpopulations. Trajectory dynamics were assessed with RNA velocity (scVelo, PhyloVelo) and pseudotime analysis (Monocle3), and SASP profiles were characterized using the SenMayo gene set and Human Protein Atlas secreted-protein lists, with overlap and expression reported as dot plots scaled by fold change relative to other populations.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| RNA velocity (scVelo, PhyloVelo) | Unbiased trajectory/lineage relationship analysis between p16+ and p21+ cells across all tissues | — | not stated |
| Pseudotime analysis (Monocle3) | Biased developmental trajectory tracing as follow-up to PhyloVelo in brain dataset and other tissues | — | not stated |
| t-SNE dimensionality reduction | Visualization of cell population structure across all tissues (Figs. 1A, 1C, 2A, 2E, 2I, and extended figures) | — | na |
| Fold change comparison (relative to all other populations) | SASP factor expression in p16+, p21+, and double-positive cells (Fig. 1G dot plots and analogues across tissues) | — | not stated |
| Standard error comparison of SASP factor expression | Informal comparison of SASP variability between p16+ and p21+ populations (Fig. 1F) | — | not stated |
| Cell-type proportion analysis (percentage of G0/G1 cells across filtering steps) | Characterization of p16+ and p21+ subpopulation composition and age-related changes (Fig. 1B, EV1A-B) | — | not stated |
-
Differential SASP gene expression between p16+ and p21+ cells was characterized using fold change relative to all other populations, with overlap assessed by counting shared genes from predefined lists↳ Could also: A pseudobulk negative-binomial differential expression framework (e.g., DESeq2 or edgeR aggregating cells per biological sample) with Benjamini-Hochberg FDR correction could also identify differentially expressed SASP genes between the two populations — Pseudobulk methods account for within-sample cell-to-cell correlation that inflates degrees of freedom in naive single-cell tests, and FDR correction would control the false-positive rate across the many simultaneous gene comparisons
-
The degree of gene-list overlap between p16+ and p21+ SASP profiles was reported descriptively by counting shared genes (e.g., 2 shared SenMayo genes in brain, 4 shared in expanded secretome list)↳ Could also: A hypergeometric test or Fisher's exact test could also formally assess whether the observed overlap is greater or less than expected by chance given the sizes of the two gene lists and the total tested gene universe — A formal overlap test would provide a p-value and effect size (odds ratio) for the degree of separation, making the claim of distinctness statistically quantifiable rather than relying on the raw count alone
-
SASP variability between p16+ and p21+ populations was compared informally by comparing their standard errors (Fig. 1F)↳ Could also: Levene's test or Brown-Forsythe test for equality of variances, or the coefficient of variation (CV), could also formally compare dispersion between the two groups — SEM is sensitive to group size and conflates measurement precision with biological variability; CV or a formal variance-equality test would more directly and size-independently characterize heterogeneity differences between the populations
-
Trajectory independence of p16+ and p21+ cells was concluded from visual inspection of RNA velocity and Monocle3 pseudotime plots↳ Could also: A quantitative metric such as the transition probability matrix from scVelo, or a permutation-based test of trajectory correlation, could also provide a numerical measure of the independence between the two cell states — Quantitative trajectory divergence scores would allow readers to gauge statistical confidence in the independence conclusion beyond visual assessment of arrow-field plots
-
Age-related changes in p16+ and p21+ cell proportions were described visually (Fig. EV1A-B and analogues) without a reported formal test↳ Could also: Compositional analysis methods (e.g., propeller, Dirichlet regression, or ANCOM-BC) could also formally test whether cell-type proportions differ between young and old animals while respecting the sum-to-one constraint of compositional data — Compositional methods account for the dependency structure inherent in proportions and yield p-values and confidence intervals for age-related proportion shifts, enabling comparison across tissues
-
Cross-dataset and cross-tissue consistency of findings was assessed by repeating the same analysis in independent datasets and describing qualitative agreement narratively↳ Could also: A fixed- or random-effects meta-analysis of a common effect size (e.g., log fold change of a core SASP gene or Jaccard similarity of p16+/p21+ gene lists) across datasets could also quantitatively pool evidence — Formal meta-analysis would yield a summary effect estimate with confidence interval and an I² statistic for between-dataset heterogeneity, making the cross-dataset generalizability directly testable rather than narrative
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
CyTOF mass cytometry in mouse bone identifies SERPINE1 protein exclusively in p16+ cells and IL-6 exclusively in p21+ cells, providing protein-level evidence for distinct senotypesflow-cytometry mouse bone 2025×1papers★ This paper is the founder (earliest)
-
RNA velocity and PhyloVelo trajectory analysis finds no common ancestor between p16+ (CDKN2A+) and p21+ (CDKN1A+) cells across brain and multiple aging tissues, indicating independent cellular originsscRNA-seq mouse brain 2025×1papers★ This paper is the founder (earliest)
-
All p16+ (CDKN2A+) cell clusters expand with age in mouse brain, whereas age-related expansion of p21+ (CDKN1A+) clusters is restricted to specific cell typesscRNA-seq mouse brain up 2025×1papers★ This paper is the founder (earliest)
-
p16+ cells display substantially higher inter-cell variability in expressed SASP factors than p21+ cells, indicating a more heterogeneous secretory senotypescRNA-seq mouse brain mixed 2025×1papers★ This paper is the founder (earliest)
-
p16+ and p21+ cells in mouse hippocampus share only CXCL16 and PLAUR from the SenMayo SASP gene set, indicating largely non-overlapping secretory profilesscRNA-seq mouse hippocampus none 2025×1papers★ This paper is the founder (earliest)
-
Mouse liver contains the lowest proportion of p21+ (CDKN1A+) cells among all aging tissues analyzedscRNA-seq mouse liver down 2025×1papers★ This paper is the founder (earliest)
-
p21+ (CDKN1A+) cells are considerably more abundant than p16+ cells in mouse skeletal muscle, bone, and liver with rare double-positive cellsscRNA-seq mouse skeletal-muscle up 2025×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41162753
Paper: Saul et al. 2025, EMBO J. Distinct senotypes in p16- and p21-positive cells across human and mouse aging tissues. PMID 41162753 / PMC12669595 / doi:10.1038/s44318-025-00601-2.
Repo: https://github.com/donshiva88/SaulLab_code @ commit
0ba9634402cc65f2d5dcf3b85d7d23d6b9078e96 (default branch main, pushed
2025-09-03). The repo contains exactly two files: README.md and P16_vs_P21.
Nature of the repo (critical for scope)
P16_vs_P21 is not a runnable pipeline. It is a 9.4 KB collection of
figure-plotting snippets (R: Seurat/ggplot2/ggalluvial/fmsb; one Python
reticulate block for RNA-velocity). The snippets operate on pre-computed,
un-shipped intermediate objects, e.g.:
old_sc_brain_adjusted_labelled_perfect_works, seurat_brain1.2,
Fig5_compendium_genes_complete, summary_both_datasets.xlsx,
naked_frame_SenMayo_clean_p16_p21_both.
There is no code for: QC/filtering thresholds, normalization parameters, integration method, PCA/neighbors, clustering resolution, tSNE params, or the manual cell-type labelling that produced those objects. Several blocks are also syntactically broken (the reticulate Python block has all statements on one line). Figs 1f, 4 ("manual calculation"), 1g (Cytoscape) are explicitly manual.
This is a compendium / meta-analysis spanning 14 public datasets (one per tissue; see Data availability). GSE161340 is the Fig. 1 dataset only.
In scope (attempted)
The repo does ship the per-cell p16/p21 classification logic (Fig. 1b/c), which is self-contained and tool-agnostic:
Cd45 = Ptprc > 0(immune gate; keep Cd45-neg)Mki67 = Mki67 > 0(proliferation gate; keep Mki67-neg)G10M = Phase=="S"(cell-cycle gate; keep non-S)- among the gated (Cd45-neg, Mki67-neg, non-S) cells:
p16+= Cdkn2a>0,p21+= Cdkn1a>0,dPo= both>0,none= neither.
Reproduction target (clear, falsifiable, pipeline-derivable): Using the
paper's own Fig. 1 data (GSE161340, single-cell Old brain samples
GSM4905055 sc_OBrain1 + GSM4905056 sc_OBrain3) and the classification logic
above, regenerate the senescent-cell composition and test the paper's stated
claim that "p21+ cells constituted the dominant senescent population" in the
aged brain (Fig. 1, Results). Pipeline: standard scanpy load+QC+normalize+
cell-cycle-score → gate → classify → count. A standard third-party tool
(scanpy) applied to the paper's data is an explicitly valid reproduction mode
per the brief.
Out of scope (NOT attempted — the hard ~20%, stated honestly)
- Exact reproduction of Figs 1a/c/d tSNE/DimPlot/DotPlot — requires the un-shipped integrated, clustered, hand-labelled Seurat object; no parameters given. Not reproducible from shipped artifacts.
- Fig 1e RNA velocity — depends on un-shipped velocity object; snippet is broken.
- Fig 5 compendium / SenMayo radar charts (Fig 4) — depend on un-shipped
summary_both_datasets.xlsx/ compendium frames aggregated across 14 datasets; "manual calculation" per the repo. - The other 13 tissue datasets — same classification could be applied, but out of the 80/20 budget for this unit; GSE161340 (Fig.1) is the representative.
- SASP/secretome overlap counts (e.g. "only 2 genes Cxcl16/Plaur shared") — needs an un-shipped secretome gene list + DE; deferred as the hard 20%.
Possible-fabrication watch
The intermediate object names (..._perfect_works, ..._correct3) and the
absence of any pipeline from raw GEO data to those objects mean the headline
figures are not independently re-derivable from the shipped code+data. The
in-scope classification result is the part that can be checked 1:1 against the
paper's qualitative claim; any divergence is recorded provisionally for a human.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The one self-contained, pipeline-derived claim — that p21+ cells dominate the senescent population in aged brain, with p16+ and dPo cells rare — reproduces 1:1 in direction (p21+ 8.85x gated / 12.1x all cells; p16+ 2.9%, dPo 1.4%) using the repo's own positivity gate on the paper's own GSE161340 data. The deviation is negligible and the core claim holds, so this is not a fabrication or discrepancy case. The limitations are reproducibility-side: the paper's claims are qualitative (no printed numbers to match), the authors' repo is figure-plotting snippets over un-shipped intermediates rather than a runnable pipeline, and only 1 of 14 datasets was attempted — hence overall yellow, a solid but bounded reproduction with explainable, mostly deposit-driven gaps.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.