Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Age Deceleration and Reversal Gene Patterns in Dauer Diapause.

Aging Cell · 2025
L1 88/100 PQI 96
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
88/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 74% of all assessed papers rank 276 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (1:1 on the core results), re-run with fresh genuine compute on «our HPC» («job», node n094) after the «infra» workdir was reclaimed. The paper applies the BiT age transcriptomic aging clock (github.com/Meyer-DH/AgingClock @a047583 - a third-party tool on the authors' own data, P16-valid) to GEO GSE288723 C. elegans daf-2 dauer RNA-seq (19909 genes x 54 samples; raw-counts file sha256 236b863f..., byte-identical across two independent fetches). Pipeline inside the SLURM job: tarball-fetch repo -> download GSE288723 -> build conda env -> raw counts -> CPM -> the repo's own make_binary (per-sample non-zero-median binarization, called directly from shipped code) -> PBA = sum(coef*binary)+intercept. The v2 coefficient set (7198 clock genes, matching the paper's stated count) reproduces the reported predicted biological ages essentially exactly: D1 122.5h vs 123, D4 179.8h vs 180, D30 258.7h vs 256; aging-rate ratios 0.795 vs 0.79 (D1->D4) and 0.127 vs 0.12 (D4->D30); L3 controls 26.8h far younger than dauer (age deceleration/reversal confirmed); PC1<->BiT-age Pearson 0.89-0.90 vs reported 0.91. The v1 (576-gene) clock does NOT match, independently confirming the version used. Every per-sample PBA is bit-identical to the prior genuine run (diff=0), and no reported value was found non-derivable from the shipped data+code -> no fabrication concern. NOT attempted: DEG/pathway counts and gene clustering (Fig 3; enrichment pipeline not fully shipped) and UV/CPD repair kinetics + GLTD (Fig 5-6; wet-lab + bespoke). Provisional grades; a human reviewer signs AUDIT.md.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 88
    assessed: 2026-06-16 ⛓ 5df9b716182e
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether aging is slowed during C. elegans dauer diapause and reversed upon dauer exit, and seeks to characterize the transcriptomic and molecular mechanisms underlying this deceleration and rejuvenation.

Core claims
  • Dauer arrest, regardless of duration, does not cause lasting functional decline in lifespan, brood size, or developmental resumption finding
  • The BiT Age transcriptome clock reveals a decelerated increase in biological age during dauer and an age reversal upon dauer exit finding
  • Dauer aging is characterized by downregulation of metabolic pathways and upregulation of longevity, proteasome, and autophagy pathways finding
  • Transcription-blocking DNA lesions induce lasting transcription stress in dauers that is rapidly resolved by transcription-coupled nucleotide excision repair (TC-NER) during dauer exit finding
  • BiT Age is a binarized, lifespan-trained transcriptomic aging clock enabling accurate biological age prediction in C. elegans resource
  • PC1 of transcriptomic variation strongly correlates with BiT Age predicted biological age, indicating it mirrors the biological aging trajectory finding
  • Duration of prior dauer arrest influences the rate of transcriptomic recovery upon exit, with longer arrest delaying molecular (but not developmental) recovery finding
  • An independent stochastic data-based aging clock and a published wild-type dauer exit dataset both corroborate the rejuvenation pattern upon dauer exit finding
Experimental setups
Assay System Perturbation Readout Platform
developmental resumption assay daf-2(e1370) C. elegans dauer larvae dauer arrest duration (1, 10, 20, 30 days) percentage of animals in dauer, L4, or young adult stage over time
lifespan assay wild-type and daf-2(e1370) C. elegans dauer arrest duration (1, 10, 20 days) vs non-dauer adult survival
brood size assay wild-type and daf-2(e1370) C. elegans dauer arrest duration (1, 10, 20 days) vs non-dauer number of viable progeny per animal
bulk RNA-sequencing daf-2(e1370) C. elegans: dauer (D1, D4, D15, D30), dauer exit (6h, 24h), and L3 larvae dauer arrest duration / dauer exit timepoint genome-wide gene expression / transcriptome
principal component analysis same bulk RNA-seq dataset (dauer, exit, L3 samples) none (computational analysis) major sources of transcriptomic variation and sample clustering
BiT Age transcriptomic aging clock prediction same bulk RNA-seq dataset (dauer, exit, L3 samples) none (computational analysis) predicted biological age (PBA, in hours)
stochastic data-based aging clock prediction dauer larvae RNA-seq data none (computational analysis) predicted biological age
KEGG pathway enrichment analysis differentially expressed genes from dauer aging and dauer exit RNA-seq time courses none (computational analysis) significantly up- or downregulated biological pathways
Key results
  • No difference in lifespan or brood size between daf-2 animals arrested in dauer for different durations (1, 10, 20 days) and non-dauer daf-2 worms
  • daf-2 mutants showed extended lifespan compared to wild-type
  • daf-2 mutants showed reduced brood size compared to wild-type
  • Predicted biological age (PBA) increased from D1 (123 h) to D30 (256 h), but the biological aging rate slowed sharply from ~0.79 (D1–D4) to ~0.12 (D4–D30) 123h to 256h; rate 0.79 to 0.12
  • During dauer aging, 24 KEGG pathways were significantly downregulated versus only 5 significantly upregulated 24 down vs 5 up
  • PC1 of the transcriptomic PCA strongly correlates with BiT Age predictions Pearson r=0.91
  • L3 control samples showed significantly lower predicted biological age than dauer or 24h post-dauer samples
Key statistics
  • pvalue p < 0.001 (One-way ANOVA with Dunnett's test comparing lifespan/brood size across dauer durations vs non-dauer daf-2)
  • mean PBA: D1=123h, D4=180h, D30=256h (BiT Age predicted biological age across dauer aging time course)
  • other biological aging rate ≈0.79 (57h/72h) (rate of biological aging between D1 and D4)
  • other biological aging rate ≈0.12 (76h/624h) (rate of biological aging between D4 and D30)
  • correlation Pearson r=0.91, p=4.4e-16 (correlation between PC1 and BiT Age predictions)
  • correlation Pearson r=-0.15, p=0.36 (all samples); r=0.39, p=0.02 (excluding L3) (correlation between PC2 and BiT Age predictions)
  • count 24 pathways downregulated, 5 upregulated (KEGG pathway enrichment during dauer aging)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study performs bulk RNA-seq across C. elegans dauer arrest timepoints (D1–D30) and post-exit timepoints (6 h, 24 h) to characterize transcriptomic aging and rejuvenation, applying the BiT Age binarized transcriptome clock to predict biological age. Between-group phenotypic comparisons (brood size, biological age predictions) used one-way ANOVA with post-hoc tests, while gene expression trajectories were modeled per gene using linear slopes across ordered timepoints, followed by KEGG pathway enrichment. Results are reported as mean ± SD with asterisk-coded significance thresholds, with exact p-values provided for Pearson correlations.

Replicationbiological Sample sizeThree independent experiments stated for brood size; RNA-seq sample counts per timepoint not specified in main text; lifespan and developmental assay statistics deferred to supplementary Table S1 GroupsDauer D1, D4, D15, D30; dauer exit 6 h and 24 h; L3 non-dauer controls; WT and daf-2 non-dauer for lifespan and brood size Pairingunpaired Randomization/blindingnot stated DispersionSD Effect sizesno Confidence intervalsno Multiplicity correctionDunnett's test (brood size ANOVA); Tukey HSD (BiT Age ANOVA); adjusted p-values for KEGG enrichment (adjustment method not specified in main text)
Statistical tests used
Test Applied to n Assumptions
one-way ANOVA with Dunnett's multiple comparison test Brood size assay comparing post-dauer duration groups to daf-2 non-dauer reference (Figure 1E) three independent experiments; exact per-group n not stated in main text not stated
one-way ANOVA with post hoc Tukey test Predicted biological age (BiT Age) across dauer arrest and L3 timepoints (Figure 2C) number of RNA-seq samples per timepoint not specified in main text; supplementary Table S2 referenced not stated
linear model (per-gene slope estimation over ordered timepoints) Gene expression trajectory analysis across dauer aging (D1–D30) and dauer exit (6–24 h post-exit) not stated
Pearson correlation Correlation of PC1 and PC2 with BiT Age predictions across all samples and excluding L3 PC1 r=0.91, p=4.4e-16; PC2 r=-0.15, p=0.36 (all samples); PC2 r=0.39, p=0.02 (excluding L3) not stated
principal component analysis (PCA) Global visualization of transcriptomic variation across dauer and exit samples (Figure 2B) na
KEGG pathway enrichment analysis with adjusted p-values Pathway-level analysis of DEGs identified by significant linear slopes during dauer aging and exit (Figure 3A) not stated
Approaches that could also have been used
  • Gene expression trajectories over the dauer time course were summarized by fitting a per-gene linear slope across ordered timepoints to classify genes as consistently up- or down-regulated
    Could also: Time-series-aware methods such as maSigPro, ImpulseDE2, or spline-based generalized additive models could also be used to characterize temporal expression patterns — These approaches explicitly model non-linear trajectories and can capture the plateau phase noted between D4 and D15, potentially identifying genes with non-monotonic dynamics that a linear slope would underweight
  • Pathway-level analysis used over-representation testing of DEGs (significant slopes) against KEGG gene sets with adjusted p-values
    Could also: Gene Set Enrichment Analysis (GSEA) using the full ranked list of per-gene slope statistics could also be applied — GSEA uses the continuous effect-size ranking of all genes rather than a binary DEG threshold, utilizing the full quantitative gradient in the trajectory data and avoiding sensitivity to the significance cutoff
  • Predicted biological age differences across timepoints were assessed with one-way ANOVA followed by Tukey HSD
    Could also: A linear mixed-effects model with timepoint as a fixed effect and experimental batch as a random effect could also be applied, or Kruskal-Wallis with Dunn's post-hoc if normality is not supported by the small per-timepoint sample counts — Mixed-effects models accommodate unequal group sizes and potential batch structure in RNA-seq experiments; rank-based alternatives do not require the normality assumption that ANOVA relies on
  • Post-dauer brood size across dauer-duration groups was compared using one-way ANOVA with Dunnett's test
    Could also: A negative binomial or Poisson regression model, or a non-parametric Kruskal-Wallis test with Dunn's correction, could also be used for count-based offspring data — Offspring counts are discrete and potentially right-skewed; regression models for count data or rank-based tests may better match the distributional properties of brood size measurements
  • The association between PCA coordinates and BiT Age predictions was quantified with Pearson correlation
    Could also: Spearman rank correlation could also measure this association — Spearman correlation is robust to outliers and monotonic but non-linear relationships, which is relevant when aging trajectories may not be perfectly linear in PC space
  • Dispersion for the brood size summary was reported as mean ± SD across three independent experiments
    Could also: Reporting 95% confidence intervals or overlaying individual replicate data points could also convey the group estimates — With three biological replicates, a 95% CI directly communicates the precision of the mean estimate and supports inference about the parameter of interest, complementing SD which describes spread of observations
Software: BiT Age (binarized transcriptomic aging clock, Meyer and Schumacher 2021) · stochastic data-based aging clock (Meyer and Schumacher 2024)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
2
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE130811 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE141514 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE288723 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE52861 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Fig 2C
C1
Reported
PBA D1 dauer = 123 h
Reproduced
122.5 h (v2, n=3)
exact
C2
Reported
PBA D4 dauer = 180 h
Reproduced
179.8 h (v2, n=3)
exact
C3
Reported
PBA D30 dauer = 256 h
Reproduced
258.7 h (v2, n=3)
within tolerance
C4
Reported
biological aging rate D1->D4 = 0.79
Reproduced
0.795
exact
C5
Reported
biological aging rate D4->D30 = 0.12
Reproduced
0.127
within tolerance
C6
Reported
L3 PBA significantly < dauer
Reproduced
L3 26.8h vs dauer 122-259h (clear separation)
partial
C7
Reported
PC1 vs BiT age Pearson 0.91 (p=4.4e-16)
Reproduced
0.89 (log2CPM) / 0.90 (binarized)
within tolerance
C0
Reported
BiT age clock = 7,198 genes
Reproduced
v2 (7198) reproduces PBAs; v1 (576) does not
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 88/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

Clean 1:1 reproduction. Run on the authors' own publicly deposited raw counts (GSE288723) plus the shipped AgingClock repo, the BiT-age v2 (7198-gene) clock reproduces every attempted core value within rounding/averaging tolerance: PBAs 122.5/179.8/258.7h vs 123/180/256, aging rates 0.795/0.127 vs 0.79/0.12, and the age-deceleration/reversal conclusion (L3 26.8h << dauer 122-259h; PC1 corr ~0.90 vs 0.91). All residual deviations are on the technical/expected side (paper's integer rounding, n=3 averaging, unspecified PCA input), none on the authors' side, and nothing was found non-derivable from the shared data — no fabrication concern. Out-of-scope items (DEG/pathway counts, clustering, UV/CPD wet-lab) are honestly disclosed rather than graded.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

202.6 k
tokens (I/O) · 13.7 M incl. cache
108 min
runtime · 0.01 CPU-h
2.5 GB
peak RAM
1
HPC jobs
hummel
machine