Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Distinct senotypes in p16- and p21-positive cells across human and mouse aging tissues.

EMBO J · 2025
L1 50/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough only PARTIALLY. The authors' repo (SaulLab_code @ 0ba9634) is NOT a runnable pipeline: just figure-plotting snippets (P16_vs_P21) over un-shipped intermediate Seurat objects, with zero QC/normalization/integration/clustering/labelling code. So the headline figures (Fig.1 tSNE/DotPlot/velocity, Fig.4 SenMayo radars, Fig.5 14-dataset compendium) are NOT independently re-derivable from shipped artifacts. However the repo DOES ship the self-contained per-cell p16/p21 classification logic (Cd45/Mki67/S-phase gate -> Cdkn2a/Cdkn1a positivity). Using a standard third-party tool (scanpy) on the paper's OWN Fig.1 data (GSE161340 single-cell OLD brain), this 1:1 reproduces the paper's stated claim that p21+ cells dominate the senescent population in aged brain (p21+ 8.9-12x p16+), with p16+ and dPo cells rare. Graded 'partial' because the claims are qualitative in the paper (no printed proportions to compare numerically) and only Fig.1/GSE161340 of 14 datasets was attempted (80/20). NOT attempted: exact embeddings/labelling, RNA velocity, the other 13 tissue datasets, SASP/secretome overlap counts, SenMayo radars, SCENIC -- the hard ~20% requiring un-shipped intermediates.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-14 ⛓ d61d4f528df4
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Because p21^Cip1 and p16^Ink4a are both used as senescence markers but their individual in vivo roles remain unclear, the paper tests whether p16+ and p21+ senescent cells represent overlapping or distinct populations with shared or divergent SASP profiles across mouse and human aging tissues.

Core claims
  • p16+ and p21+ senescent cells are largely distinct, non-overlapping populations across murine and human aging tissues, with only rare double-positive cells finding
  • RNA velocity and pseudotime analyses show p16+ and p21+ cells follow independent trajectories with no common ancestor or direct transition between states finding
  • p16+ cells display heterogeneous, tissue-specific secretomes whereas p21+ cells exhibit broader but more conserved SASP profiles finding
  • Only a small set of 'core' SASP factors (e.g., ICAM1, IGFBP4/6, CXCL16, PLAUR) are shared across tissues and species finding
  • p16 expression increases broadly with age in brain while p21 upregulation is restricted to specific cell clusters finding
  • The separation of p16+ and p21+ cells was confirmed at the protein level by CyTOF in bone (Serpine1 in p16+, IL-6 in p21+ cells) finding
  • p21+ cells are more abundant than p16+ cells across muscle, bone, and liver finding
  • A meta-analysis pipeline integrating multiple single-cell RNA-seq atlases with SenMayo gene set is used to map senescent cell heterogeneity in vivo method
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA-seq analysis young (4 m) and old (24 m) mouse hippocampus aging (none) p16 (Cdkn2a)/p21 (Cdkn1a) expression, cell-type composition, SASP/secretome profiles GSE161340 (Ogrodnik et al)
single-cell RNA-seq analysis young and old whole mouse brain (excluding hindbrain) aging (none) p16/p21 positivity, secretory profiles Ximerakis et al dataset
single-cell RNA-seq analysis young and old murine skeletal muscle aging (none) cell-type identity, p16/p21 expression, SASP factors GSE172410 (Zhang et al); Tabula Muris Senis
single-cell RNA-seq analysis young and old murine bone aging (none) cell types, p16/p21 expression, secretome GSE128423 (Baryawno et al); Tabula Muris Senis
single-cell RNA-seq analysis young and old murine liver aging (none) cell types, p16/p21 expression, secretome GSE166504 (Su et al); Tabula Muris Senis
single-cell RNA-seq analysis murine spleen and kidney aging (none) p16/p21 distinct populations and secretome Calico murine aging cell atlas (Kimmel et al)
mass cytometry (CyTOF) young and aged mouse bone aging (none) p21 and p16 protein expression, Serpine1, IL-6 CyTOF (Doolittle et al)
RNA velocity / pseudotime trajectory analysis murine brain, muscle, bone, liver scRNA-seq populations none developmental trajectory / common ancestor between p16+ and p21+ cells scVelo, PhyloVelo, Monocle3
Key results
  • In hippocampus, only Cxcl16 and Plaur were expressed in both p16+ and p21+ subpopulations using the SenMayo gene set 2 shared genes
  • Expanded analysis of 1989 secreted protein genes plus SenMayo shared only four genes (Col8a2, Plaur, Cxcl16, Fbln5) between p21+ and p16+ cells 4 of 1989
  • Variability (standard error) of expressed SASP factors among p16+ cells was substantially higher than among p21+ cells
  • RNA velocity/PhyloVelo showed no common ancestor for p16+ and p21+ cells across brain and all tissues
  • All p16+ cells increased with age in brain whereas p21+ cell increase was restricted to specific clusters
  • Across muscle, bone, and liver p21+ cells were considerably more abundant than p16+ cells, with rare double-positive cells
  • CyTOF in bone identified Serpine1 exclusively in p16+ cells and IL-6 exclusively in p21+ cells
  • Liver had the lowest proportion of p21+ cells among tissues analyzed, though still the majority of senescent cells
Key statistics
  • count 2 shared SASP genes (Cxcl16, Plaur) (overlap between p16+ and p21+ hippocampal cells via SenMayo)
  • count 4 of 1989 secreted protein genes shared (Col8a2, Plaur, Cxcl16, Fbln5) (Human Protein Atlas + SenMayo overlap in brain)
  • count 5 main cell populations (identified in mouse hippocampus)
  • count 9, 14/17, and 13 cell populations (skeletal muscle, bone, liver respectively)
  • count young 4 m vs old 24 m (mouse ages compared in hippocampus dataset)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This computational resource paper performed comparative single-cell RNA sequencing (scRNA-seq) analyses of p16+ and p21+ cell populations across multiple murine (brain, skeletal muscle, bone, liver, spleen, kidney) and human (skin, lung) aging tissues using publicly available datasets. Cells were filtered to exclude proliferating (Ki67+, S-phase) and high-CD45 immune cells before defining p16+, p21+, and double-positive subpopulations. Trajectory dynamics were assessed with RNA velocity (scVelo, PhyloVelo) and pseudotime analysis (Monocle3), and SASP profiles were characterized using the SenMayo gene set and Human Protein Atlas secreted-protein lists, with overlap and expression reported as dot plots scaled by fold change relative to other populations.

Replicationbiological Sample sizeSample sizes described per dataset: e.g., 8 young and 8 old whole mouse brains (Ximerakis et al. 2019 validation cohort); cell counts per cluster not stated in provided text; findings validated in at least one independent dataset per tissue Groupsp16+ vs p21+ vs double-positive cells; young vs old animals across murine and human tissues; multiple independent scRNA-seq datasets per tissue for validation Pairingunpaired Randomization/blindingnot stated DispersionSEM Effect sizesyes Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
RNA velocity (scVelo, PhyloVelo) Unbiased trajectory/lineage relationship analysis between p16+ and p21+ cells across all tissues not stated
Pseudotime analysis (Monocle3) Biased developmental trajectory tracing as follow-up to PhyloVelo in brain dataset and other tissues not stated
t-SNE dimensionality reduction Visualization of cell population structure across all tissues (Figs. 1A, 1C, 2A, 2E, 2I, and extended figures) na
Fold change comparison (relative to all other populations) SASP factor expression in p16+, p21+, and double-positive cells (Fig. 1G dot plots and analogues across tissues) not stated
Standard error comparison of SASP factor expression Informal comparison of SASP variability between p16+ and p21+ populations (Fig. 1F) not stated
Cell-type proportion analysis (percentage of G0/G1 cells across filtering steps) Characterization of p16+ and p21+ subpopulation composition and age-related changes (Fig. 1B, EV1A-B) not stated
Approaches that could also have been used
  • Differential SASP gene expression between p16+ and p21+ cells was characterized using fold change relative to all other populations, with overlap assessed by counting shared genes from predefined lists
    Could also: A pseudobulk negative-binomial differential expression framework (e.g., DESeq2 or edgeR aggregating cells per biological sample) with Benjamini-Hochberg FDR correction could also identify differentially expressed SASP genes between the two populations — Pseudobulk methods account for within-sample cell-to-cell correlation that inflates degrees of freedom in naive single-cell tests, and FDR correction would control the false-positive rate across the many simultaneous gene comparisons
  • The degree of gene-list overlap between p16+ and p21+ SASP profiles was reported descriptively by counting shared genes (e.g., 2 shared SenMayo genes in brain, 4 shared in expanded secretome list)
    Could also: A hypergeometric test or Fisher's exact test could also formally assess whether the observed overlap is greater or less than expected by chance given the sizes of the two gene lists and the total tested gene universe — A formal overlap test would provide a p-value and effect size (odds ratio) for the degree of separation, making the claim of distinctness statistically quantifiable rather than relying on the raw count alone
  • SASP variability between p16+ and p21+ populations was compared informally by comparing their standard errors (Fig. 1F)
    Could also: Levene's test or Brown-Forsythe test for equality of variances, or the coefficient of variation (CV), could also formally compare dispersion between the two groups — SEM is sensitive to group size and conflates measurement precision with biological variability; CV or a formal variance-equality test would more directly and size-independently characterize heterogeneity differences between the populations
  • Trajectory independence of p16+ and p21+ cells was concluded from visual inspection of RNA velocity and Monocle3 pseudotime plots
    Could also: A quantitative metric such as the transition probability matrix from scVelo, or a permutation-based test of trajectory correlation, could also provide a numerical measure of the independence between the two cell states — Quantitative trajectory divergence scores would allow readers to gauge statistical confidence in the independence conclusion beyond visual assessment of arrow-field plots
  • Age-related changes in p16+ and p21+ cell proportions were described visually (Fig. EV1A-B and analogues) without a reported formal test
    Could also: Compositional analysis methods (e.g., propeller, Dirichlet regression, or ANCOM-BC) could also formally test whether cell-type proportions differ between young and old animals while respecting the sum-to-one constraint of compositional data — Compositional methods account for the dependency structure inherent in proportions and yield p-values and confidence intervals for age-related proportion shifts, enabling comparison across tissues
  • Cross-dataset and cross-tissue consistency of findings was assessed by repeating the same analysis in independent datasets and describing qualitative agreement narratively
    Could also: A fixed- or random-effects meta-analysis of a common effect size (e.g., log fold change of a core SASP gene or Jaccard similarity of p16+/p21+ gene lists) across datasets could also quantitatively pool evidence — Formal meta-analysis would yield a summary effect estimate with confidence interval and an I² statistic for between-dataset heterogeneity, making the cross-dataset generalizability directly testable rather than narrative
Software: scVelo · PhyloVelo · Monocle3 · SenMayo gene set

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
21
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE132042 GEO in Methods (http://purl.org/orb/Methods)
also used by 2 papers:
GSE122960 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE128423 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
ENST00000579755 Ensembl in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE117278 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE128033 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE129788 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE130148 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE130973 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE132901 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE161340 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE166504 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE172410 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE212109 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE218300 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE237301 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE269660 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
P10144 UniProt in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41162753

Paper: Saul et al. 2025, EMBO J. Distinct senotypes in p16- and p21-positive cells across human and mouse aging tissues. PMID 41162753 / PMC12669595 / doi:10.1038/s44318-025-00601-2.

Repo: https://github.com/donshiva88/SaulLab_code @ commit 0ba9634402cc65f2d5dcf3b85d7d23d6b9078e96 (default branch main, pushed 2025-09-03). The repo contains exactly two files: README.md and P16_vs_P21.

Nature of the repo (critical for scope)

P16_vs_P21 is not a runnable pipeline. It is a 9.4 KB collection of figure-plotting snippets (R: Seurat/ggplot2/ggalluvial/fmsb; one Python reticulate block for RNA-velocity). The snippets operate on pre-computed, un-shipped intermediate objects, e.g.: old_sc_brain_adjusted_labelled_perfect_works, seurat_brain1.2, Fig5_compendium_genes_complete, summary_both_datasets.xlsx, naked_frame_SenMayo_clean_p16_p21_both.

There is no code for: QC/filtering thresholds, normalization parameters, integration method, PCA/neighbors, clustering resolution, tSNE params, or the manual cell-type labelling that produced those objects. Several blocks are also syntactically broken (the reticulate Python block has all statements on one line). Figs 1f, 4 ("manual calculation"), 1g (Cytoscape) are explicitly manual.

This is a compendium / meta-analysis spanning 14 public datasets (one per tissue; see Data availability). GSE161340 is the Fig. 1 dataset only.

In scope (attempted)

The repo does ship the per-cell p16/p21 classification logic (Fig. 1b/c), which is self-contained and tool-agnostic:

  • Cd45 = Ptprc > 0 (immune gate; keep Cd45-neg)
  • Mki67 = Mki67 > 0 (proliferation gate; keep Mki67-neg)
  • G10M = Phase=="S" (cell-cycle gate; keep non-S)
  • among the gated (Cd45-neg, Mki67-neg, non-S) cells: p16+ = Cdkn2a>0, p21+ = Cdkn1a>0, dPo = both>0, none = neither.

Reproduction target (clear, falsifiable, pipeline-derivable): Using the paper's own Fig. 1 data (GSE161340, single-cell Old brain samples GSM4905055 sc_OBrain1 + GSM4905056 sc_OBrain3) and the classification logic above, regenerate the senescent-cell composition and test the paper's stated claim that "p21+ cells constituted the dominant senescent population" in the aged brain (Fig. 1, Results). Pipeline: standard scanpy load+QC+normalize+ cell-cycle-score → gate → classify → count. A standard third-party tool (scanpy) applied to the paper's data is an explicitly valid reproduction mode per the brief.

Out of scope (NOT attempted — the hard ~20%, stated honestly)

  • Exact reproduction of Figs 1a/c/d tSNE/DimPlot/DotPlot — requires the un-shipped integrated, clustered, hand-labelled Seurat object; no parameters given. Not reproducible from shipped artifacts.
  • Fig 1e RNA velocity — depends on un-shipped velocity object; snippet is broken.
  • Fig 5 compendium / SenMayo radar charts (Fig 4) — depend on un-shipped summary_both_datasets.xlsx / compendium frames aggregated across 14 datasets; "manual calculation" per the repo.
  • The other 13 tissue datasets — same classification could be applied, but out of the 80/20 budget for this unit; GSE161340 (Fig.1) is the representative.
  • SASP/secretome overlap counts (e.g. "only 2 genes Cxcl16/Plaur shared") — needs an un-shipped secretome gene list + DE; deferred as the hard 20%.

Possible-fabrication watch

The intermediate object names (..._perfect_works, ..._correct3) and the absence of any pipeline from raw GEO data to those objects mean the headline figures are not independently re-derivable from the shipped code+data. The in-scope classification result is the part that can be checked 1:1 against the paper's qualitative claim; any divergence is recorded provisionally for a human.

Figures / tables: Fig. 1
C1
Reported
p21+ cells constituted the dominant senescent population in aged brain (qualitative; Fig.1, GSE161340)
Reproduced
p21+ dominant: gated Cd45-/Mki67-/non-S cells p21=363 vs p16=41 vs dPo=19 (8.85x); all cells p21=942 vs p16=78 (12.1x)
partial
C2
Reported
p16+ cells a minority of senescent cells in aged brain (qualitative; Fig.1)
Reproduced
p16+ = 41/1395 gated cells (2.9%)
partial
C3
Reported
double-positive (dPo) cells rare (qualitative; Fig.1)
Reproduced
dPo = 19/1395 gated (1.4%)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

The one self-contained, pipeline-derived claim — that p21+ cells dominate the senescent population in aged brain, with p16+ and dPo cells rare — reproduces 1:1 in direction (p21+ 8.85x gated / 12.1x all cells; p16+ 2.9%, dPo 1.4%) using the repo's own positivity gate on the paper's own GSE161340 data. The deviation is negligible and the core claim holds, so this is not a fabrication or discrepancy case. The limitations are reproducibility-side: the paper's claims are qualitative (no printed numbers to match), the authors' repo is figure-plotting snippets over un-shipped intermediates rather than a runnable pipeline, and only 1 of 14 datasets was attempted — hence overall yellow, a solid but bounded reproduction with explainable, mostly deposit-driven gaps.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

112.8 k
tokens (I/O) · 6 M incl. cache
12 min
runtime · 0.01 CPU-h
1.3 GB
peak RAM
1
HPC jobs
hummel
machine