Why an integrated view of gene expression studies on hematopoiesis in mouse aging is better than the sum of their parts.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for the INTEGRATION step, which reproduced 1:1 for the pinnable claim. This is a FEBS Lett perspective/meta-analysis; the author's GitHub repo (commit b8597e3) ships the derived result tables (CSV/XLSX), not raw GEO data or an end-to-end pipeline. We reproduced the citation-frequency filtering on the shipped 'aging list' (valid per P16). C1 (human-mouse overlap) is an EXACT match: filtering Subset_AL>1_ovl_Human>1.csv to Counts>3 yields precisely the five genes the paper names (Gadd45g, Tsc22d1, Thbd, Zfp36, Cd55) - fully derivable from shipped data, no fabrication concern. C2 (aging-signature size) is a soft inequality: '>3 -> over 200 genes' holds (we get 524 from 5279 genes across 22 studies), but the paper's ~200 figure described the older 16-study analysis, so it is not a tight 1:1 number for this update (graded partial; Counts>=7 -> 199, close to ~200). NOT ATTEMPTED (hard ~20%): (a) re-deriving each study's DE list from raw GEO by re-running edgeR/DESeq2/Seurat on the 22 source datasets - the paper ships no per-study pipeline/unified params; (b) the GO-lineage counts (myeloid 16, B-cell 12, T-cell 11, platelet 4, inflammation ~12) which require an external GO database plus the author's manual curation choices, not reproducible from the shipped tables alone. Compute ran on «our HPC» («infra», «job»); repo cloned inside the compute job; pure-stdlib Python, no env needed.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 75assessed: 2026-06-14 ⛓ b3c0e1ae8ee2
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetAn integrated, citation-weighted analysis across many independent hematopoietic-aging gene expression datasets in mice yields more reliable and informative conclusions about aging-related gene expression changes than examining each publication's DE gene list separately.
- ★ Combining differentially expressed (DE) gene lists from multiple publications into a unified 'aging list' (AL) with citation counts, and deriving a shorter high-confidence 'aging signature' (AS, genes cited in >3 publications, ~200 genes), produces a more reliable and consistent picture of hematopoietic aging than any single study. method
- ★ Variation between DE gene lists across studies arises from inconsistent definitions of 'young' and 'old', technical differences in RNA-seq/analysis protocols, and inconsistent gene category (GO) definitions, rather than primarily from the sequencing technology used. finding
- ★ The most frequently detected DE genes in the AS list are highly consistent across independent publications, indicating they are reliable markers of hematopoietic aging. finding
- ★ Cross-referencing the AS gene list against specific Gene Ontology categories (e.g., inflammatory response, myeloid differentiation, platelet activation) recovers additional aging-associated genes not highlighted in individual original publications. finding
- ★ Hematopoietic stem cell aging in C57BL/6 mice is associated with increased self-renewal/cycling, inflammation, myeloid lineage skewing, and platelet activation gene signatures. finding
- Type of sequencing platform (expression microarray, bulk RNA-seq, single-cell RNA-seq) has little effect on the DE gene lists obtained, compared to experimental design factors. finding
- A public GitHub repository (aging_update) provides an updated compilation and recount of aging-related genes in mice across 34 identified datasets. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| gene expression microarray (ExpArray) | hematopoietic stem cells (HSCs), primarily C57BL/6 mice | aging (young vs old cohorts) | differentially expressed genes between young and old HSCs | — |
| bulk RNA sequencing (bulkRNA) | hematopoietic stem cells (HSCs), C57BL/6 mice | aging (young vs old cohorts) | differentially expressed genes between young and old HSCs | — |
| single-cell RNA sequencing (SC_RNA) | hematopoietic stem cells (HSCs), C57BL/6 mice (also one DBA/2 and one BALB/c dataset) | aging (young vs old cohorts) | differentially expressed genes between young and old HSCs | — |
| Gene Ontology (GO) classification / cross-referencing | compiled aging signature (AS) gene list derived from HSC aging studies | none | membership and up/down-regulation direction of AS genes within specific GO categories (e.g., GO:0006954, GO:0030099, GO:0030168) | — |
- – Threshold of >3 citations across publications defines the aging signature (AS), a list of over 200 highly reproducible DE genes. >200 genes
- ▼ Dnmt3b is a frequently reported DE gene in the AS list, consistently downregulated upon aging.
- – Scanning the AS list against GO:0006954 (inflammatory response) recovered additional inflammation-related genes (e.g., Cyp26b1, Ptger4, Vwf, Camk1d, Plscr1, Bcl6, Sema7a, Nupr1, Jun, Slamf1, Syk, C4b, Prcp); all but two (Syk, Camk1d) were upregulated with aging. 11 of 13 genes upregulated
- – Scanning the AS list against GO:0030099 (myeloid cell differentiation) recovered 16 genes; 3 (Anxa2, Il15, Gata1) were downregulated and the remaining 13 were upregulated, confirming myeloid skewing. 13 of 16 genes upregulated
- – Scanning the AS list against GO:0030168 (platelet activation) recovered 4 upregulated genes (F2rl3, Plek, Vwf, Itgb3), while Gata1 in the same GO group was downregulated. 4 genes up, 1 down
- – Only 2 genes (Vwf, Itgb3) intersected between genes suggested in prior individual publications and the current GO:0030168 platelet activation group, out of 10 originally suggested. 2 of 10 genes overlapped
- – The number of phenotypically identified HSCs increases with age in C57BL/6 mice while individual cell functionality decreases.
- count 34 papers/datasets identified (updated from original 16, of which 12 passed quality check) (number of hematopoietic aging gene expression datasets compiled)
- count >200 genes (size of the aging signature (AS) list using citation threshold >3)
- count 16 genes total in GO:0030099 (myeloid cell differentiation) overlap with AS list (cross-check of AS list against myeloid differentiation GO category)
- count 3 of 16 genes downregulated (Anxa2, Il15, Gata1) (direction breakdown within GO:0030099 overlap)
- count 4 upregulated genes (F2rl3, Plek, Vwf, Itgb3) in GO:0030168 overlap (platelet activation GO category cross-check with AS list)
- other P-value cutoff typically set at 5%, adjusted for multiple testing (standard significance threshold used in DE gene analysis protocols across studies)
- other male mouse lifespan varies from 450 to approximately 1000 days across laboratory strains (cited lifespan variation across mouse strains relevant to defining 'old age')
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a perspective/meta-analysis article that introduces no new primary experimental data. Its core analytical contribution is a citation-frequency vote-counting approach: differentially expressed (DE) gene lists from 12 published datasets (of 16 passing quality review; updated to 34 total) are pooled into an 'aging list' (AL), and a reproducible subset termed the 'aging signature' (AS) is derived by retaining genes cited as DE in more than three independent publications. Functional characterization of the AS is then performed by manual cross-referencing against specific Gene Ontology (GO) term member lists. The underlying primary datasets are described as having used standard R-based differential expression pipelines (edgeR, DESeq2, Seurat, Scater/Scran) with a 5%-level p-value adjusted for multiple testing.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Citation-frequency vote-counting aggregation (not a formal statistical test; tallies the number of independent publications reporting each gene as DE to form the aging list and aging signature) | Construction of the aging list (AL) and aging signature (AS) across 12 pooled DE gene lists | 12 publications (out of 16 passing quality check) for the main analysis; 34 datasets mentioned as updated count | not stated |
| Standard differential expression analysis via edgeR, DESeq2, Seurat, or Scater/Scran (described as methods of the underlying primary datasets, not performed by this paper) | Primary datasets feeding into the meta-analytic aggregation | — | not stated |
| Gene Ontology term membership scan (manual overlap of AS gene list against specific GO term member sets: GO:0006954, GO:0030099, GO:0030168, and others) | Functional characterization of the AS for inflammation, myeloid differentiation, platelet activation, and lymphoid categories | — | na |
-
Cross-study gene aggregation was performed by counting how many publications reported each gene as DE (citation-frequency vote-counting), with a threshold of >3 citations to define the aging signature↳ Could also: A random-effects meta-analysis of log fold changes (e.g., using the metafor R package or a dedicated RNA-seq meta-analysis framework such as MetaVolcano or RankProd) could be applied across the standardized effect estimates from each dataset — Effect-size meta-analysis weights each study by its precision and yields a pooled fold-change estimate with a confidence interval, allowing genes with large but noisy effects to be distinguished from genes with small but consistently detected effects — information that binary presence/absence vote-counting does not capture
-
The >3-citation threshold that defines the aging signature is described by the authors themselves as arbitrary↳ Could also: A permutation- or bootstrap-based threshold could be derived by randomly reshuffling gene labels across studies to estimate the expected citation frequency under the null hypothesis of no true overlap, and setting the cutoff at a desired false-discovery rate — A data-driven threshold would quantify how many citations are required to exceed chance-level overlap given the number, size, and diversity of the contributing studies, replacing the arbitrary numerical choice with a statistically grounded criterion
-
Gene Ontology enrichment was assessed by manually scanning the AS list against specific pre-selected GO term member lists↳ Could also: A formal overrepresentation analysis (e.g., Fisher's exact test across all GO terms with Benjamini-Hochberg FDR correction, as implemented in clusterProfiler, g:Profiler, or WebGestalt) applied to the full AS against an appropriate background gene set could be used — Formal enrichment testing provides a statistical measure of whether each GO category is more represented in the AS than expected by chance, and correction across all tested GO terms would help prioritize the most robustly enriched pathways rather than relying on pre-selected categories
-
DE analysis in the underlying primary studies used binary young/old group contrasts; the paper notes that multi-age-point datasets were forced into this two-group framework, which the authors describe as not optimal↳ Could also: Regression-based DE analysis treating age as a continuous or ordinal covariate (supported natively in DESeq2 and edgeR via GLM design matrices) could model gene expression as a function of age across the full time-series — Continuous-age modelling would capture monotonic or dose-response-like trends, make fuller use of multi-timepoint datasets already present in the literature, and potentially identify genes whose regulation is graded rather than threshold-like
-
Cross-dataset batch effects are identified as a concern contributing to variation in DE gene lists, but no cross-study batch correction is applied before aggregating DE calls↳ Could also: A harmonized joint re-analysis applying cross-study normalization and batch correction (e.g., ComBat-seq, limma::removeBatchEffect, or RUVSeq) to raw count matrices from the GEO datasets could precede a pooled DE analysis — A single harmonized analysis across all raw datasets would increase statistical power, provide a unified effect estimate for each gene, and reduce the influence of study-specific technical variation that the citation-counting approach cannot disentangle from biological signal
-
The aging signature is reported as a list of gene names with citation counts, without any quantification of the direction consistency or magnitude of change across studies↳ Could also: Reporting for each AS gene the proportion of studies showing upregulation versus downregulation, together with the median and range of log fold changes across studies, would also be informative — Summarizing effect direction consistency alongside citation frequency would allow readers to distinguish robustly directional genes from those with mixed directionality across studies, and would make the AS list more actionable for functional interpretation and follow-up experiments
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-38627103
Paper: Bystrykh LV (2024) Why an integrated view of gene expression studies on hematopoiesis in mouse aging is better than the sum of their parts. FEBS Lett 598:2765. PMID 38627103 / PMC11586588 / DOI 10.1002/1873-3468.14869.
Nature of paper: Perspective / meta-analysis (review article). It integrates differential-expression (DE) gene lists from many published mouse-HSC aging studies into one "aging list" (AL), then filters by citation frequency (number of studies in which a gene is DE) to define an "aging signature" (AS).
Code/data artifact: Author's own GitHub repo
github.com/LeonidBystrykh/aging_update (default branch main, not archived, no
license, last push 2024-04-25, 758 KB). It ships the derived result tables as CSV/
XLSX — not raw GEO data and not an end-to-end pipeline script. The repo IS the
supplementary data. (Per P16, applying analysis to the paper's own shipped data is a
valid reproduction.)
Key files:
AL_Table_22.csv— the integrated aging list: 576 genes × 22 studies. Columns:Gene,Counts(= citation frequency = #studies in which the gene is DE), andLogFC_*+Adj.P.Val_*per study.Subset_AL>1_ovl_Human>1.csv— genes overlapping between the mouse AL and a human aging dataset (23 genes), with their mouseCounts.Table 1.csv,Data_sources_*.xlsx— the data-source inventory (Table 1).
In scope (pipeline-derived, clearly specified → attempt)
| id | reported claim | paper location | how to reproduce |
|---|---|---|---|
| C1 | "only five genes (Gadd45g, Tsc22d1, Thbd, Zfp36, Cd55) are in the mouse AS list" (human↔mouse overlap that survives the AS threshold) | Results, human-mouse comparison | filter Subset_AL>1_ovl_Human>1.csv to Counts > 3 → expect exactly those 5 symbols |
| C2 | aging signature = genes with citation threshold ">3", "a list of over 200 genes" | Methods/Results | count genes in AL_Table_22.csv with Counts > 3 |
Out of scope / not attempted (the hard ~20%)
- GO-lineage counts (myeloid 16 genes [13 up/3 down], B-cell 12, T-cell 11, platelet 4, inflammation ~12). These require mapping AS genes onto specific GO terms (GO:0030099, GO:0030168, GO:0006954, …) using an external GO database and the author's manual curation choices. Not reproducible from the shipped tables alone → not attempted, flagged as GO/manual-dependent.
- Re-deriving each study's DE list from raw GEO (re-running edgeR/DESeq2/Seurat on every one of the 22 source datasets). This is the upstream that produced AL; the paper does not ship per-study pipelines or unified parameters, and GSE4332 is only one of dozens of source accessions. Out of scope — we reproduce the integration step on the shipped DE values, not the per-study DE calling.
Compute plan
Trivial compute (pure-stdlib Python over a 576-row CSV). Per HARD RULES, executed on
«our HPC»: clone repo into «infra» inside a SLURM job (compute nodes have internet), run
analysis, pull back the small result JSON. No conda env needed (stdlib csv only).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
C1 reproduces 1:1 — filtering the shipped Subset_AL>1_ovl_Human>1.csv to Counts>3 yields precisely the five genes the paper names (Gadd45g, Tsc22d1, Thbd, Zfp36, Cd55), fully derivable from author-shipped data with no fabrication concern. C2 is a soft inequality that holds (524 > 200), but the discrepancy from the paper's ~200 is on our/version side: AL_Table_22 integrates 22 studies whereas the ~200 figure described the older 16-study analysis (Counts>=7 -> 199 ≈ original). The deviation sits at the input/sample-definition level and is fully explained — not an authors' defect. Overall solid and reproduced for the pinnable claim, but partial: the upstream per-study DE re-derivation and GO-lineage counts were not attempted, so yellow rather than green.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.