Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A blood-based DNA damage signature in patients with Parkinson's disease is associated with disease progression.

Nat Aging · 2025
L1 69/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
69/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 35% of all assessed papers rank 745 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH FOR THE CORE CLAIM, NOT FOR THE EXACT P-VALUES. The paper's central computational validation (Fig 8a,b ALBATRO gene-length bias on two PUBLIC datasets) was reproduced on «our HPC». METHOD (per Methods + repo GeneLenght_Script1.txt): differential expression -> split DEGs into up/down -> gene length = gene_end-gene_start (Ensembl GRCh38.110 GTF, == BioMart genomic span) -> two-sided Mann-Whitney Wilcoxon on log10(length). RESULT: in BOTH datasets downregulated DEGs are significantly longer than upregulated ones (biased suppression of longer transcripts) -- the paper's qualitative claim (Fig 8b, P<0.0001) is REPRODUCED 1:1, and the reproduced Fig-8a density plot matches the figure visually (fig8a_repro.png). C1 GSE99039 (limma iPD-vs-CONTROL on GPL570; iPD=205/ctrl=233 selected EXACTLY): at padj<0.05 (the only DEG cut yielding the method-required >=100 genes/side) Wilcoxon P=2.23e-5 vs reported 5.5e-5 -- same order of magnitude, same direction (within-tol). C2 GSE68719 (shipped DESeq2 diffexp): direction + high significance reproduced robustly (P=5.98e-27 at |log2FC|>0.322); reported 2e-7 is bracketed by the LFC-threshold sweep, so it is consistent with the documented method at an unspecified threshold (partial). The exact panel-a p-values are NOT uniquely reproducible because the authors do not state the validation-set DE method/thresholds; a threshold sweep is reported for transparency. NOT ATTEMPTED (out of scope): all PPMI Figs 1-7 (controlled-access LONI data + repo withholds metadata) and wet-lab gamma-H2AX/IHC (non-computational). No value was fabricated. NB: Fig 8 panel labels the blood set 'GSE99309', a typo for GSE99039.

💻 Code ↗ 🗄 Data: GSE99039

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 69
    assessed: 2026-06-20 ⛓ a992b0140bea
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-20
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Aging is the principal risk factor for Parkinson's disease, and since progressive nuclear DNA damage is a causative mechanism of aging, the study tests whether accumulated nuclear DNA damage contributes to PD pathophysiology and can serve as a blood-based biomarker of disease progression.

Core claims
  • A blood-based DNA damage signature is present in PD patients and is associated with disease progression. finding
  • PD patients show disrupted DNA repair pathways and biased suppression of longer transcripts, indicating age-related, transcription-stalling DNA damage. finding
  • At the intake visit, the DNA damage signature was detected only in patients who went on to have more severe motor symptom progression over 3 years, suggesting predictive value for disease severity. finding
  • The DNA damage signature was validated in independent PD cohorts. finding
  • Increased DNA damage was confirmed in peripheral blood cells and in dopamine neurons of the substantia nigra pars compacta in postmortem PD brains. finding
  • GSEA shows downregulation of DNA repair pathways, including nucleotide excision repair, in iPD, LRRK2 G2019S and prodromal groups but not in GBA mutation carriers at visit 1. finding
  • PD blood transcriptomes show shared downregulation of pathways related to RNA processing/transcription/translation (e.g., MYC_TARGETS_V2, ribosome, RNA processing) and mitochondrial function/oxidative phosphorylation across iPD, GBA, LRRK2 and prodromal groups. finding
  • The NER gene ERCC8 shows a PD association (rs11744756 polymorphism) in the PDGene meta-analysis database, supporting a role for DNA repair genes in PD risk. resource
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (blood transcriptomics) human whole blood, PPMI cohort (iPD, GBA mutation carriers, LRRK2 G2019S carriers, prodromal individuals, healthy controls) disease state (PD/prodromal) vs healthy control, longitudinal (visit 1 baseline vs visit 8 at 36 months) differential gene expression (DEGs), gene set enrichment (GSEA/ORA) across Hallmark, Reactome, KEGG, WikiPathways, GO-BP collections Salmon (transcript quantification), DESeq2 (Wald test)
genetic association meta-analysis (database lookup) human PD GWAS meta-analysis data (PDGene database) none (observational polymorphism association) odds ratio and p-value for association of ERCC8 rs11744756 with PD PDGene database (pdgene.org)
Key results
  • 69 DEGs were common between visit 1 and visit 8 in PD vs HC comparisons, a 26-fold enrichment over chance 26-fold
  • HALLMARK_MYC_TARGETS_V2 and HALLMARK_OXIDATIVE_PHOSPHORYLATION were the only two Hallmark pathways downregulated across all four PD/prodromal groups at visit 1
  • Leading-edge gene analysis identified 28 shared genes in the oxidative phosphorylation set related to mitochondrial respiration/complex I 28 genes
  • Leading-edge gene analysis identified 9 shared genes in the MYC_TARGETS_V2 set related to macromolecular synthesis (transcription/translation) 9 genes
  • Reactome GSEA identified 466 deregulated pathways, of which 13 were shared across groups, all downregulated (2 mitochondrial, remainder transcription/translation-related) 13 of 466 shared
  • HALLMARK_DNA_REPAIR was downregulated in iPD, LRRK2 and prodromal groups at visit 1, but not in GBA carriers; Reactome confirmed downregulation of nucleotide excision repair specifically
  • ERCC8 rs11744756 polymorphism associated with PD in PDGene meta-analysis OR 1.15
Key statistics
  • fold_change log2(FC) > 0.322 (FC = 1.25) (DEG significance threshold applied throughout study)
  • pvalue P-adj < 0.05 (significance threshold for DEGs and GSEA pathways throughout study)
  • fold_change 26-fold (enrichment of common DEGs between visit 1 and visit 8 over chance)
  • count 484 PD / 187 HC at visit 1; 268 PD / 157 HC at visit 8 (PPMI cohort sizes analyzed)
  • pvalue P = 9.27 × 10−07 (association of ERCC8 rs11744756 polymorphism with PD in PDGene meta-analysis)
  • other odds ratio 1.15, 95% CI 1.08–1.21 (ERCC8 rs11744756 PD association effect size)
  • other PC1 = 44% of variance at visit 1; 37% at visit 8 (PCA of PD/genetic-PD transcriptomic data)
  • count 53 of 58 prodromal individuals (prodromal cases included after excluding those who developed PD symptoms within 2 years or had missing data)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper is an observational, longitudinal cohort study using bioinformatic analysis of blood RNA-seq data from the PPMI cohort. Differential gene expression between PD subgroups (idiopathic PD, GBA carriers, LRRK2 G2019S carriers, prodromal individuals) and healthy controls at two visits was assessed with DESeq2's Wald test, followed by gene set enrichment analysis (GSEA via fgseaMultilevel) and overrepresentation analysis (ORA) across five pathway collections. Group age differences were evaluated with the Kruskal-Wallis test, age was used as a covariate for the prodromal group, and a retrospective post hoc power analysis was performed to support the adequacy of sample sizes.

Replicationbiological Sample sizeGroup sizes given as patient/control counts per visit (e.g., 484 PD vs 187 HC at visit 1; 268 PD vs 157 HC at visit 8); a retrospective post hoc power analysis was performed and led to exclusion of the LRRK2 G2019S group at visit 8 for insufficient power GroupsiPD, GBA carriers, LRRK2 G2019S carriers, and prodromal individuals vs healthy controls, at baseline (visit 1) and 36-month follow-up (visit 8) Pairingmixed Randomization/blindingna Dispersionnot stated Exact p-valuesyes Effect sizesyes Confidence intervalsyes Multiplicity correctionNot explicitly named beyond reporting of adjusted p-values ('P-adj'); GSEA enrichment via fgseaMultilevel is described as including multiple-comparison correction internally
Statistical tests used
Test Applied to n Assumptions
DESeq2 Wald test Differential gene expression, PD subgroups vs healthy controls, visit 1 and visit 8 (Fig. 1b) 484 PD vs 187 HC at visit 1; 268 PD vs 157 HC at visit 8 (subgroup ns not fully specified in this excerpt) not stated
Kruskal-Wallis test Comparison of age between patients with PD and healthy controls, and prodromal vs controls not stated
GSEA, one-sided permutation-based test on normalized enrichment score (fgseaMultilevel, adaptive multilevel Monte Carlo) Pathway enrichment against Hallmark, Reactome, KEGG, WikiPathways and GO-BP collections, visit 1 (and visit 8) Same cohort ns as DESeq2 analyses; based on full ranked gene lists not stated
Overrepresentation analysis (ORA) Pathway analysis based on significant DEG lists, five gene set collections DEG lists per comparison (counts not specified in this excerpt) not stated
Enrichment factor / fold-enrichment analysis Overlap of 69 common DEGs between visit 1 and visit 8 PD-vs-HC comparisons (Extended Data Fig. 1f) 69 shared DEGs; 26-fold enrichment reported not stated
Retrospective post hoc power analysis Confirming adequate statistical power for transcriptome differences per group/visit (Extended Data Fig. 2a,b) Group sample sizes per visit not stated
Approaches that could also have been used
  • Differential expression was defined using a fixed log2(FC) > 0.322 and P-adj < 0.05 threshold via DESeq2's Wald test.
    Could also: limma-voom or edgeR's quasi-likelihood F-test — These use different mean-variance modeling approaches for RNA-seq counts and can serve as a complementary cross-check that the DEG list is not sensitive to the specific dispersion model chosen.
  • Pathway-level analysis combined ORA (based on significant DEG lists) with GSEA on ranked gene lists across five collections.
    Could also: Threshold-free, competitive gene set tests such as CAMERA or ROAST — These account for inter-gene correlation within a pathway while, like GSEA, avoiding a hard significance cutoff, which can be useful when effect sizes are small and distributed across many genes.
  • Age differences for the prodromal group (which was significantly older than controls) were addressed by including age as a covariate specifically in analyses involving that group.
    Could also: A single linear mixed-effects model with age as a covariate applied uniformly across all group comparisons — Applying the same covariate-adjustment framework to every comparison, rather than group-specific corrections, can make the handling of age consistent and easier to compare across subgroups.
  • Overlap between visit 1 and visit 8 DEG sets was quantified with an enrichment-factor (fold-enrichment) calculation.
    Could also: A formal hypergeometric or Fisher's exact test for gene-list overlap — This would yield an explicit p-value alongside the fold-enrichment estimate, giving a direct statistical significance statement for the overlap in addition to its magnitude.
  • Age was compared between patients with PD and controls using the Kruskal-Wallis test.
    Could also: A parametric ANOVA or linear model on age, with normality diagnostics reported — If age distributions are approximately normal, a parametric test can offer greater power; reporting the distributional check alongside the nonparametric choice would let readers evaluate the assumptions either way.
  • The two visits (baseline and 36-month follow-up) appear to have been analyzed as separate cross-sectional comparisons against controls.
    Could also: A longitudinal mixed-effects model treating visit as a within-subject repeated factor — For participants sampled at both visits, modeling the repeated measurement structure directly can increase statistical power and explicitly separate within-subject change from between-group differences.
Software: Salmon (transcript quantification) · DESeq2 (Wald test for differential expression) · fgsea / fgseaMultilevel (GSEA)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40913219

Title: A blood-based DNA damage signature in patients with Parkinson's disease is associated with disease progression. Venue: Nat Aging 2025 · DOI 10.1038/s43587-025-00926-x · PMCID PMC12443628 Code: https://github.com/DNAdamageinPD/Natureageing @ 8c035512bf4a6c758fbfa603fd25e65cc7cdda9c (main, pushed 2025-08-02) also Zenodo 10.5281/zenodo.16728586 ; Code Ocean 10.24433/CO.1403558.v2

What the paper does

Longitudinal blood transcriptomics from the PPMI cohort. Core narrative: PD shows disrupted DNA-repair pathways and biased suppression of longer transcripts ("ALBATRO" = Analysis for Length-Biased Alterations in TRanscription Output), an indirect read-out of transcription-stalling DNA damage. The signature appears at intake only in patients with more severe 3-yr motor progression. Validated in two independent public GEO datasets, plus wet-lab (γH2AX foci, postmortem brain IHC).

Pipeline used

  • RNA-seq quantification: Salmon (lengthScaledTPM via tximport) → DESeq2 DE (Wald test). DEG criterion applied throughout: |log2FC| > 0.322 and padj < 0.05.
  • GSEA: fgsea fgseaMultilevel against Hallmark/KEGG/Reactome/WikiPathways/GO-BP (MSigDB v7.2 / GO v2023.2).
  • ALBATRO / gene-length analysis: BioMart (v3.17) gene length = gene_end − gene_start (genomic span, exons+introns). Split DEGs into up/down, test log10(length) for normality (Shapiro–Wilk → non-normal), then two-sided Mann–Whitney/Wilcoxon on up vs down gene lengths; ≥100 genes per distribution required. (Repo file GeneLenght_Script1.txt implements exactly this with thresholds i=log2FC, k=FDR.)

IN SCOPE (pipeline-derived, reproducible from PUBLIC data) — Fig. 8a,b

The validation-cohort ALBATRO gene-length-bias results run on public GEO data:

  • C1 — GSE99039 (whole blood, GPL570 Affymetrix; 205 iPD + 233 controls): biased suppression of longer genes, two-sided Wilcoxon P = 0.000055.
  • C2 — GSE68719 (substantia nigra RNA-seq, GPL11154; 29 PD vs 44 controls; ships DESeq2 diffexp + normalized PCG counts as supplementary): Wilcoxon P = 0.0000002. Both datasets are fully public; both have an identifiable expected value. The method (DE → split up/down DEGs → BioMart length → Wilcoxon on log10 length) is fully specified. This is the intended reproduction target.

Secondary (also public, supportive): Fig 8c GO/GSEA of LEGs on the same two sets (Suppl. Table 9) — coarser, lower priority.

OUT OF SCOPE

  • All PPMI analyses (Figs 1–7; Extended Data 1–8): data_restricted. PPMI Salmon files + clinical metadata are controlled-access (LONI IDA, https://ida.loni.usc.edu, data use agreement). The repo explicitly withholds the metadata file "due to patient privacy concerns." The repo also ships no data, no DE result tables (PDSeverityGroups/), and no GenelengthAll.xlsx — only 4 analysis scripts. So PPMI DESeq2/GSEA/ALBATRO (Figs 1–7) cannot be reproduced from scratch.
  • Wet-lab / manual: γH2AX foci in PBMCs (Fig 6c,d), postmortem SNpc IHC, semiautomated image analysis — non-computational, not attempted.

Status of execution

Scope fully established and the public reproduction targets pinned with exact reported values. «our HPC» compute was NOT executed: the «infra» VPN required operator 2FA, which did not complete within the run window (tunnel never came up; connect timed out), and the room was finalized before the tunnel was available. No values were reproduced; none are fabricated. See ROOM_RESULT.json (status: error) and AUDIT.md.

Figures / tables: Fig 8bFig 8aFig 8
C0_qual
Reported
ALBATRO central claim: downregulated DEGs significantly LONGER than upregulated (biased suppression of longer transcripts), both validation datasets, two-sided Wilcoxon P<0.0001 (Fig 8b)
Reproduced
down_longer in BOTH; GSE99039 P=2.23e-5, GSE68719 P=5.98e-27 (both <1e-4)
exact
C1_p
Reported
GSE99039 (blood microarray) ALBATRO Wilcoxon P=0.000055 (Fig 8a)
Reproduced
P=2.23e-5 at padj<0.05 (936 up/140 down genes); limma iPD-vs-CONTROL DE
within tolerance
C1_n
Reported
GSE99039: 205 iPD + 233 controls
Reproduced
IPD=205 CONTROL=233 (exact, from 558-sample series)
exact
C2_p
Reported
GSE68719 (substantia nigra RNA-seq) ALBATRO Wilcoxon P=0.0000002 (Fig 8a)
Reproduced
down_longer, P=5.98e-27 at |log2FC|>0.322&padj<0.05; reported 2e-7 bracketed by LFC-cut sweep (2.4e-13@cut0.5 .. 1.4e-3@cut1.0)
partial
C2_n
Reported
GSE68719: 29 PD + 44 controls (=73)
Reproduced
74 sample columns in deposited DESeq2 norm-counts (+1 unexplained)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 69/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

The paper's central computational claim — ALBATRO gene-length bias, with downregulated DEGs significantly longer than upregulated DEGs in two independent public datasets (Fig 8b, Wilcoxon P<0.0001) — reproduces 1:1 in direction and significance (GSE99039 P=2.23e-5; GSE68719 P=5.98e-27). The deviations are on our methodology side, downstream of a mild authors' documentation gap: the paper never specifies the validation-set DE method/thresholds, so the exact Fig-8a p-values are only reproducible within an order of magnitude (GSE99039) or as a point on a threshold sweep (GSE68719), plus a minor +1 sample column in the deposited GSE68719 counts. No fabrication: all values are derivable from the shared public data; the restricted PPMI core (Figs 1-7) is simply out of scope, which is a data-availability limit, not an authors' defect.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

308 k
tokens (I/O) · 22 M incl. cache
61 min
runtime · 0.02 CPU-h
2 GB
peak RAM
3
HPC jobs
hummel
machine