Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Under the Shadow: Old-biased Genes Are Subject to Weak Purifying Selection at Both the Tissue- and Cell Type-Specific Levels.

Genome Biol Evol · 2025
L1 93/100 PQI 97
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
93/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 85% of all assessed papers rank 154 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough and reproduced 1:1. Paper: 'Under the Shadow' (GBE 2025), multi-species test of the ADICT hypothesis (old-biased genes under weaker purifying selection). Reproduced the killifish dataset (GSE66712, the brief's named accession) end-to-end from the authors' own repo (NisanYildiz/selection-shadow @ 91b7948): shipped raw counts + shipped dN/dS xlsx + GEO metadata -> DESeq2 size-factor normalization -> per-gene Spearman(expr,age) -> BH adjust -> age-class assignment (|rho|>0.5, padj<0.1) -> expression-conservation-vs-age rho (Fig1C). Compared against the authors' OWN shipped derived tables (supplements gene_class CSVs), a file-level check. RESULT: skin table BYTE-IDENTICAL (same SHA256); liver table identical except 2 of 9851 BH-adjusted p-values differing only in the 15th significant digit (IEEE-754 last bit, BLAS/R minor-version drift) with all gene classes identical; all 6 class counts exact (liver 537/618/8696, skin 1468/1405/5996); Fig1C rho values reproduced (liver -0.76/-0.32, skin -0.84/-0.71), matching the paper's negative-ADICT direction. One reproducibility-critical env gotcha: the authors' scripts only run under R<4.0 (stringsAsFactors=TRUE) -- under R 4.3 they silently produce EMPTY results; faithful re-run needs that default restored. No fabrication concern: every value is fully derivable from shipped data+code. NOT attempted (optional tail): the other species (chicken GSE114129, fly Pacifico, mouse astrocyte GSE99791, naked mole-rat GSE30337), the Tabula Muris Senis single-cell cell-type figures (Fig3-7; h5ad inputs .gitignored), and the Tajima's D population-genomics tracks -- same code idiom on different shipped data, skipped to keep to a few clear data points on the named dataset.

💻 Code ↗ 🗄 Data: GSE66712

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 93
    assessed: 2026-06-14 ⛓ 623cc2b65a36
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Does the age-related decrease in transcriptome conservation (ADICT)—weaker purifying selection on old-biased genes, predicted by Medawar's mutation accumulation theory and the selection shadow—extend across diverse nonmammalian metazoans, and does it arise cell-autonomously within specific cell types rather than merely from age-related shifts in tissue/cell type composition?

Core claims
  • Age-related decrease in transcriptome conservation (ADICT) is commonly found in ageing tissues of nonmammalian species (chicken brain, killifish liver and skin, fruit fly brain). finding
  • Old-biased genes show consistently lower average sequence conservation than young-biased genes across the analyzed datasets. finding
  • The ADICT trend is detectable at the single-cell-type level in adult mouse tissues, supporting a cell-autonomous component rather than purely composition-driven changes. finding
  • Cell type-specific transcriptomes vary dramatically in average conservation levels, with neural cells highest and immune cells lowest. finding
  • Gene class (young- vs old-biased) is a significant predictor of conservation independent of gene age and mean expression level (linear models). finding
  • An expression-conservation-age correlation approach using age-series data provides a more powerful metric than dividing genes into young/old-biased classes. method
  • Results support Medawar's mutation accumulation process / selection shadow as a shaper of metazoan tissue ageing. mechanism
  • In long-lived naked-mole rat, old-biased genes showed weaker conservation across brain, kidney, liver, but lack of replicates precludes generalization. finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (expression-conservation analysis) Gallus gallus (chicken) brain none (ageing, 100-1825 days old) expression-conservation correlation per individual vs age; gene-wise dN/dS conservation
bulk RNA-seq Nothobranchius furzeri (turquoise killifish) liver and skin none (ageing, 5-39 weeks old) expression-conservation correlation; old- vs young-biased gene conservation
bulk RNA-seq Drosophila melanogaster (fruit fly) brain none (ageing, 5-40 days old) expression-conservation correlation; old- vs young-biased gene conservation
bulk RNA-seq Heterocephalus glaber (naked-mole rat) brain, kidney, liver none (ageing, 4 vs 20 years old) old- vs young-biased gene conservation
single-cell RNA-seq (44 cell types) Mus musculus lung, skeletal muscle, brain, skin, kidney, liver (Tabula Muris Senis) none (ageing, 3/18/24 months old) per-cell-type transcriptome conservation (correlation of conservation scores with mean expression); cell-type-specific ADICT
bulk RNA-seq with astrocyte ribotagging Mus musculus astrocyte-enriched cerebellum, hypothalamus, motor cortex, visual cortex none (ageing, 4 vs 24 months old) age-related expression changes; cell-type-specific transcriptome conservation
Key results
  • Whole-transcriptome analysis revealed moderate ADICT (negative expression-conservation-age correlations) in all four nonmammalian datasets, becoming more conspicuous when limited to differentially expressed genes.
  • Old-biased gene sets showed consistently lower average gene-wise conservation than young-biased genes in all four datasets. Welch's t-test P < 0.003 across all four tests
  • Significant cell type effect and tissue effect on transcriptome conservation in two-way ANOVA. F_tissue=89.1 (d.f.=4), F_celltype=17.8 (d.f.=43), P < 1e-16
  • Immune status had a significant negative effect on transcriptome conservation (mixed model ANOVA). F_immune status=95.79, d.f.=1, P < 0.0001
  • Gene class remained a statistically significant explanatory variable of conservation when gene age and expression were accounted for.
  • In all three naked-mole rat tissues, old-biased genes showed weaker conservation than young-biased genes, but without biological replicates.
  • Neural cells, oligodendrocytes and astrocytes showed highest transcriptome conservation; immune cells (B-cells, T-cells, neutrophils, myeloid dendritic cells) on the lowest end.
  • ADICT signature (cell-type-specific decrease in transcriptome conservation with age) detected within astrocyte transcriptomes of cerebellum and hypothalamus.
Key statistics
  • pvalue P < 0.003 (Welch's t-test, lower conservation of old- vs young-biased genes across all four nonmammalian datasets)
  • other F_tissue=89.1, d.f.=4; F_celltype=17.8, d.f.=43; P < 1e-16 (two-way ANOVA of transcriptome conservation across cell types and tissues)
  • other F_immune status=95.79, d.f.=1, P < 0.0001 (mixed model ANOVA, negative effect of immune status on transcriptome conservation)
  • count 13,989 expressed genes (combined set of expressed genes across cell types for mouse single-cell conservation analysis)
  • count 10,986 whole transcriptome; 1,920 differentially expressed (G. gallus brain dataset gene counts; n=13 individuals)
  • count old-b=1,066, young-b=864 (G. gallus brain old- vs young-biased gene counts)
  • count old-b=792, young-b=831 (D. melanogaster brain old- vs young-biased gene counts)
  • count ADICT found in 76% of 66 datasets (prior work, Turan et al. 2019) (previously reported prevalence of ADICT in mammalian bulk-tissue datasets)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study investigated whether genes with higher expression in old versus young adults (old-biased genes) show weaker evolutionary sequence conservation than young-biased genes, using bulk-tissue and single-cell RNA-seq datasets from five species. The primary analytical framework paired per-sample Spearman correlations between gene expression levels and dN/dS-based conservation scores, then correlated those per-sample scores with individual age to detect age-related decrease in transcriptome conservation (ADICT). Comparisons between old-biased and young-biased gene sets used Welch's t-tests, and cell-type variation in conservation was assessed with two-way and mixed-model ANOVAs. Results were reported as correlation coefficients, F-statistics, and group mean differences with 95% CIs derived from 1,000 bootstraps.

Replicationbiological Sample sizePer-age-group sample sizes given in Table 1 for each species; no formal power analysis described. Naked mole rat dataset noted explicitly as n = 1 per age group, precluding generalization. GroupsYoung vs. old adult age groups (within species/tissue); old-biased vs. young-biased genes; 44 cell types across 6 mouse tissues Pairingunpaired Randomization/blindingnot stated DispersionCI Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Spearman's rank correlation (within-sample: expression level vs. dN/dS conservation score across genes) Calculation of per-individual transcriptome conservation metric in all bulk-tissue and cell-type datasets Varies by dataset; e.g., n = 10,986 genes (whole transcriptome) or n = 1,920 differentially expressed genes for G. gallus brain not stated
Spearman's rank correlation (expression-conservation score vs. individual age) Detection of ADICT across age series in each species/tissue dataset (Fig. 1b, c) Varies by dataset; e.g., n = 13 individuals for G. gallus brain not stated
Welch's two-sided t-test Comparison of gene-wise conservation score distributions between old-biased and young-biased gene sets in each of four non-mammalian datasets (Fig. 2a) Varies by dataset; e.g., D. melanogaster brain: n_old-biased = 792 genes, n_young-biased = 831 genes not stated
Multiple linear regression (OLS) Testing whether gene class (old- vs. young-biased) predicts dN/dS independently of gene age and mean/maximum expression level (Table S8) null not stated
Two-way ANOVA Testing tissue and cell type effects on transcriptome conservation levels across 44 cell types in young-adult mice (Fig. 3, Table S6) 44 cell types across 5 tissues; F_tissue = 89.1 df = 4, F_celltype = 17.8 df = 43 not stated
Mixed-model ANOVA (immune status as fixed effect, tissue as random effect) Testing effect of immune cell status on transcriptome conservation (Fig. 3) null not stated
Bootstrap resampling (1,000 iterations) for 95% confidence intervals Mean conservation scores of old-biased and young-biased genes relative to constantly expressed genes (Fig. 2b) Gene-level; bootstrapped within each gene class per dataset na
Approaches that could also have been used
  • Conservation score distributions between old-biased and young-biased gene sets were compared with Welch's t-test
    Could also: Mann-Whitney U (Wilcoxon rank-sum) test — dN/dS ratios are typically right-skewed and bounded at zero; a rank-based nonparametric test makes no distributional assumptions and would also be a standard choice for comparing two independent groups of gene conservation scores
  • Separate Welch's t-tests were run for each of four species/tissue datasets independently
    Could also: A linear mixed-effects model with dataset as a random effect, testing the overall old- vs. young-biased gene class effect across all datasets simultaneously — A mixed-effects framework would test the global effect in a single model, provide an estimate of between-dataset variability, and implicitly address the multiple-testing issue that arises from running parallel tests
  • Genes were dichotomized into old-biased (rho > 0.5) and young-biased (rho < −0.5) classes using a fixed threshold, with the remaining genes treated as constantly expressed
    Could also: Treating the continuous expression-age correlation (rho) as a predictor of gene conservation in a regression model — A continuous approach retains information lost by thresholding and avoids sensitivity to the chosen cutoff value; the paper acknowledged this by also showing robustness across alternative cutoffs (Fig. S3)
  • A two-way ANOVA was applied to test tissue and cell type effects on mean transcriptome conservation in young-adult mice
    Could also: A linear mixed-effects model with individual mouse as a random effect nested within tissue — Because multiple cell types are measured from the same individual, a mixed model would explicitly account for within-individual correlation and the unbalanced cell-type-per-individual design, which standard ANOVA assumes away
  • Per-sample expression-conservation Spearman correlations were computed and then correlated with age as the primary ADICT metric
    Could also: A partial correlation or regression approach that simultaneously controls for known confounders (e.g., mean expression level, gene age) within the same model — The two-step correlation approach estimates the expression-conservation-age relationship indirectly; a single model controlling for confounders at the gene level would provide a more direct test of the age effect on the conservation-expression relationship while adjusting for confounders in one step
  • 95% confidence intervals for mean conservation scores were estimated via 1,000 bootstrap resamples
    Could also: Permutation-based confidence intervals or parametric CIs (e.g., t-distribution-based) where normality holds — Both approaches are standard alternatives; permutation tests make no distributional assumptions and directly address the null hypothesis of no group difference, while parametric CIs are computationally simpler when the central limit theorem applies to the large gene sets used here
Software: null

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE132042 GEO in Results (http://purl.org/orb/Results)
also used by 2 papers:
GO:0002376 Gene Ontology (GO) in Article (http://semanticscience.org/resource/SIO_001029)
also used by 1 paper:
GSE114129 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE30337 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE66712 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE99791 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: tablestableFig1Fig2
C1
Reported
GSE66712 liver age-class counts: Old-Biased=537, Young-Biased=618, Not-significant=8696
Reproduced
Old-Biased=537, Young-Biased=618, Not-significant=8696
exact
C2
Reported
GSE66712 skin age-class counts: Old-Biased=1468, Young-Biased=1405, Not-significant=5996
Reproduced
Old-Biased=1468, Young-Biased=1405, Not-significant=5996
exact
C3
Reported
GSE66712 liver full per-gene table (rho,p,BH-padj,class), 9851 genes, SHA256 017c6282...
Reproduced
9849/9851 rows identical; 2 rows differ only in 15th significant digit of BH-adjusted p (IEEE-754 last bit); identical gene memberships/classes
within tolerance
C4
Reported
GSE66712 skin full per-gene table, 8869 genes, SHA256 67266f48dc5d34b6163bddb851968c682032f88847f7fb128a6369adea896373
Reproduced
byte-identical, same SHA256 67266f48dc5d34b6163bddb851968c682032f88847f7fb128a6369adea896373
exact
C5
Reported
Fig1C killifish expression-conservation~age rho for DE genes: negative, 'more conspicuous for DE genes' (exact label not in paper text)
Reproduced
liver rho_DE=-0.76, skin rho_DE=-0.84
within tolerance
C6
Reported
Fig1C killifish expression-conservation~age rho for all genes: negative moderate ADICT (exact label not in paper text)
Reproduced
liver rho_All=-0.32 (p=0.12), skin rho_All=-0.71 (p=1e-4)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 93/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

Reproduction of the brief's named killifish dataset (GSE66712) is essentially perfect: the skin per-gene table is byte-identical (same SHA256), all six age-class counts match exactly (liver 537/618/8696, skin 1468/1405/5996), and the liver table matches 9849/9851 rows with the remaining two differing only in the 15th significant digit of a BH-adjusted p-value (IEEE-754 last-bit, BLAS/R version drift). The central ADICT conclusion — negative expression-conservation~age correlations — is confirmed (liver -0.76/-0.32, skin -0.84/-0.71). Every value is fully derivable from shipped data+code; no fabrication concern. Caveat (not a defect): only the killifish dataset was reproduced, not the optional other-species/single-cell tail, but that is the accession named in the brief.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

174.1 k
tokens (I/O) · 13.8 M incl. cache
20 min
runtime · 0.02 CPU-h
2.8 GB
peak RAM
4 (3 failed)
HPC jobs
hummel
machine