Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Investigating epigenetic biomarkers of age, sex, and disease in captive South African cheetahs (Acinonyx jubatus jubatus).

PLoS One · 2026
L1 58/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
58/100
Reproducibility score
0.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 18% of all assessed papers rank 950 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the PIPELINE, and we did: decoded all 57 GEO GSE310779 Mammal40 IDAT pairs to betas and ran the repo's documented elastic-net age + sex clocks (alpha=0.5, log-linear age transform ASM=2/k=0.2, LOOCV, lambda.min) verbatim on the SAME 52 liver+blood training samples. Result is a PARTIAL reproduction: the clocks clearly work (age LOOCV r=0.71, p<1e-4; sex 90.2% accurate; predictions track chronological age and sex), corroborating the paper's qualitative claim, but the exact reported metrics (r=0.97 / MAE=0.86 / 52 CpGs; sex 100% / 67 CpGs) were NOT matched. The gap is NOT a fabrication signal: it is fully explained by declared, reproducible deviations. (1) We could not run SeSaMe noob/pOOBAH normalization because Bioconductor's experiment-data CDN mghp.osn.xsede.org (serving sesameData/ExperimentHub incl. the Mammal40 idatSignature/address) was globally unreachable from BOTH «our HPC» and «host»; we worked around it by decoding IDATs directly with illuminaio + the zhou-lab GitHub Mammal40 manifest (raw beta M/(M+U+100)) and installed the sesameData package from the TU-Dortmund Bioconductor mirror. (2) The authors' submission-1 'shifted' correction, exact ComBat, and hand removal of 7 stillborn/outlier samples are not fully shipped (the SID->GSM metadata CSV cheetah_metadata_vod2.csv is absent), and several neonates are mispredicted in our raw run, which alone depresses r. NOT ATTEMPTED: C6 FelidClock multi-species clock (needs external lion/tiger methylation not in this accession) and C7 SOS differential-methylation analysis (hard-20% per-CpG DMA) - both out of the 80/20 low-hanging scope.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 58
    assessed: 2026-06-16 ⛓ 78d7f2896968
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

DNA methylation-based models can accurately predict age and sex in captive cheetahs despite limited sample sizes, with multi-tissue and multi-species data enhancing predictive power, and epigenetic differences can distinguish cheetahs affected by hepatic sinusoidal obstruction syndrome (SOS) from unaffected individuals.

Core claims
  • A cheetah-specific epigenetic age clock using 52 CpG sites predicts chronological age across blood and liver with r=0.97 and MAE=0.86 years finding
  • A multi-species (cheetah, lion, tiger) age clock using 46 CpG sites predicts age across these felids with r=0.94 and MAE=1.16 years finding
  • A sex clock using 67 CpG sites accurately predicts sex in all test samples finding
  • Differential methylation analysis identified 4,377 CpG sites differing significantly between SOS-positive and SOS-negative cheetahs finding
  • Elastic net regression on HorvathMammalMethylChip40 methylation profiles can build accurate epigenetic clocks from small wildlife datasets (n=52) method
  • The age clock is accurate for adult cheetahs (>3 years) but less precise around age of sexual maturity finding
  • Cheetah-specific clocks outperform existing pan-mammalian (UniversalClock) and domestic-cat (CatClock) clocks on cheetah samples finding
  • DNA methylation establishes a foundation for biomarkers of disease (SOS) in wildlife conservation resource
Experimental setups
Assay System Perturbation Readout Platform
DNA methylation array (Infinium, bisulfite-converted) cheetah (Acinonyx jubatus) liver tissue none CpG methylation beta values for age clock training HorvathMammalMethylChip40 (Illumina), GPL28271
DNA methylation array cheetah whole blood (live animals) none CpG methylation beta values for age/sex clock testing HorvathMammalMethylChip40 Illumina Array
DNA methylation array cheetah skin tissue (necropsy) none CpG methylation beta values HorvathMammalMethylChip40 Illumina Array
DNA methylation array (public MCDB profiles) lion and tiger blood none CpG methylation beta values for multi-species clock testing mammalian methylation array (HorvathMammalMethylChip40)
Epigenome-wide association study (EWAS) cheetah liver (n=38), blood (n=7), skin (n=4) none CpG methylation correlation with age/sex (Pearson; Student's t-test for sex)
Differential methylation analysis (limma lmFit) cheetah liver (>3 years, n=30) SOS disease status (positive vs negative) differentially methylated CpG sites between disease groups
KEGG pathway enrichment analysis genes associated with differentially methylated CpGs none enriched biological pathways clusterProfiler (R)
Key results
  • Cheetah age clock (52 CpGs) predicted age across blood and liver r=0.97, MAE=0.86 years
  • Multi-species feline age clock (46 CpGs) predicted age across cheetah, lion, tiger r=0.94, MAE=1.16 years
  • Sex clock (67 CpGs) correctly predicted sex in all test samples
  • Differential methylation between SOS-positive and SOS-negative cheetahs 4,377 CpG sites (adj p<0.05)
  • CatClock performed well on cheetah blood but poorly on combined tissues blood r=0.79 MAE=1.38 yr; combined r=0.64 MAE=3.5 yr
  • UniversalClock2/3 showed high error on older cheetahs MAE 3–4.42 years
Key statistics
  • correlation r = 0.97 (cheetah age clock across blood and liver)
  • other MAE = 0.86 (cheetah age clock median absolute error (years))
  • correlation r = 0.94 (multi-species feline age clock)
  • other MAE = 1.16 (multi-species feline clock error (years))
  • count 4,377 CpG sites (adjusted p-value < 0.05) (differentially methylated sites SOS+ vs SOS-)
  • correlation correlation 0.79, MAE 1.38 years (CatClock on cheetah blood samples)
  • count n = 11, 25% (cheetahs with SOS diagnosis in cohort)
  • other MAE 3–4.42 years (UniversalClock performance on cheetah samples)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This retrospective observational study used elastic net regression (glmnet) with leave-one-out cross-validation (LOOCV) on DNA methylation beta values from the HorvathMammalMethylChip40 array to build cheetah-specific epigenetic age and sex prediction clocks, reporting Pearson r and median absolute error (MAE) as accuracy metrics. Epigenome-wide association studies (EWAS) screened CpG sites using Pearson correlation (age, continuous trait) and Student's t-test (sex, binary trait) against tissue-specific raw p-value thresholds. Differential methylation between SOS-positive and SOS-negative cheetahs was assessed with limma linear models and Benjamini-Hochberg FDR correction (adjusted p < 0.05), followed by KEGG pathway enrichment analysis; unsupervised hierarchical clustering was used both for outlier exclusion and for visualisation of the top 100 differentially methylated CpGs.

Replicationbiological Sample sizeSample sizes described narratively (44 cheetahs; after hierarchical-clustering exclusions: 38 liver, 4 skin, 7 blood from SDZWA; supplemented with MCDB profiles: 14 cheetah blood, 7 lion blood, 8 tiger blood); authors explicitly acknowledge the dataset is smaller than the recommended minimum of 70–134 samples for epigenetic clocks; no formal a priori power calculation reported GroupsAge (continuous, 0 days–16 years); sex (female [n=26] vs male [n=18]); SOS disease status (positive vs negative, liver samples restricted to age >3 years, n=30) Pairingmixed Randomization/blindingstated Dispersionmixed Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionBenjamini-Hochberg false discovery rate (FDR) for DMA; raw p-value thresholds (10^-3 for liver/blood EWAS; 10^-2 for skin EWAS) with no stated correction for EWAS
Statistical tests used
Test Applied to n Assumptions
Elastic net regression (glmnet) with Leave-One-Out Cross-Validation (LOOCV) Age clock (CheetahClock, feline multi-species clock) and sex clock construction from methylation beta values CheetahClock trained on n=38 liver + n=14 MCDB cheetah blood profiles (52 samples total); sex clock training n not separately stated not stated
Pearson correlation EWAS for age: CpG-level screening in liver (n=38), blood (n=7), and skin (n=4); also reported as clock accuracy metric (r) between inverse-transformed predicted and known age n=38 liver, n=7 blood, n=4 skin for EWAS; test-set n varies by tissue for clock evaluation not stated
Student's t-test (via WGCNA standardScreeningBinaryTrait) Sex EWAS: comparing methylation levels at each CpG between females and males null not stated
limma lmFit (linear model for array data) Differential methylation analysis (DMA): SOS-positive vs SOS-negative liver samples n=30 (liver samples from cheetahs >3 years of age) not stated
Unsupervised hierarchical clustering Outlier detection per tissue type (pre-analysis exclusion); visualisation of top 100 DMA CpG sites across liver and blood samples null na
KEGG pathway enrichment analysis (enrichKEGG, clusterProfiler) Genes associated with all significant DMA CpGs (p.adj<0.05); separate analyses for hypermethylated and hypomethylated subsets 4377 significant CpG sites na
Approaches that could also have been used
  • EWAS for age and sex used raw p-value thresholds (10^-3 or 10^-2) across approximately 37,000 probes without a formal genome-wide multiple-testing correction
    Could also: Benjamini-Hochberg FDR correction (as applied in the DMA) or a Bonferroni threshold (p < 0.05/37,492 ≈ 1.3×10^-6) could also be applied genome-wide to the EWAS — Applying the same FDR framework used in the DMA to the EWAS would quantify the expected false-discovery proportion among selected CpGs; this is particularly relevant when characterising biologically meaningful age-associated sites beyond their use as clock inputs
  • Clock performance was evaluated with LOOCV and reported as single point estimates (r and MAE) without uncertainty intervals
    Could also: Bootstrap resampling (e.g., 1,000 iterations) or repeated k-fold cross-validation could also be used to generate confidence intervals around r and MAE — With a small training set (n=52 samples), bootstrap CIs would communicate estimation uncertainty in performance metrics, helping readers assess whether differences between the CheetahClock and pan-mammalian or CatClock benchmarks are within noise
  • Student's t-test was used for the sex EWAS to compare methylation beta values between females and males
    Could also: Mann-Whitney U (Wilcoxon rank-sum) test or logistic regression could also be used; the former makes no distributional assumption, the latter directly models a binary outcome — DNA methylation beta values are bounded [0,1] and can be bimodally distributed; a non-parametric test does not rely on the normality assumption, which may be difficult to verify at this sample size
  • Outlier samples were identified and excluded using unsupervised hierarchical clustering with visual inspection of dendrograms
    Could also: Principal component analysis (PCA) with a quantitative distance criterion (e.g., samples beyond 3 SD from the centroid on PC1–PC2) or robust PCA could also be used for outlier detection — PCA-based criteria provide a reproducible, pre-specified exclusion rule that can be reported numerically, complementing the visual judgement inherent in dendrogram-based exclusion
  • Log-linear age transformation was applied prior to elastic net modeling to account for rapid methylation change before sexual maturity
    Could also: Generalized additive models (GAMs) or penalised spline regression could also model the non-linear age-methylation relationship without imposing a specific functional form — Data-driven smooth functions allow the trajectory shape to be estimated from the data rather than pre-specified, which may be useful when the exact form of age acceleration around sexual maturity is uncertain or species-specific
  • The DMA used a binary SOS-positive vs SOS-negative classification, excluding 'SOS suspect' samples from the positive group
    Could also: An ordinal or continuous severity score, or a sensitivity analysis treating 'SOS suspect' as positive, could also be used to capture graded disease severity — A graded variable would utilise intermediate phenotype information and could reveal dose-response methylation patterns across the spectrum of SOS severity, complementing the binary contrast
Software: R 4.0.1 · glmnet (R package) · WGCNA (R package; standardScreeningNumericTrait / standardScreeningBinaryTrait) · limma (R package; lmFit) · clusterProfiler (R package; enrichKEGG) · SeSaMe (normalization pipeline for IDAT files)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GPL28271 GEO in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
GSE310779 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41528985 (CheetahClock)

Paper: Investigating epigenetic biomarkers of age, sex, and disease in captive South African cheetahs. PLoS One 2026; DOI 10.1371/journal.pone.0336127. Code: https://github.com/mysrael/CheetahClock — single file CheetahClock_age_sex.Rmd. Data: GEO GSE310779 — 57 cheetah samples, Illumina HorvathMammalianMethylChip40 (Mammal40), RAW IDATs (42.5 MB) + series matrix. Per-sample tissue/sex/age are encoded in the GSM title (e.g. "Cheetah liver female 2.8y") and IDAT filenames carry GSM####_<barcode> → full GSM↔array mapping available.

Pipeline-derived results (the paper's computational outputs)

Result Pipeline In scope? Notes
CheetahClock age clock (LOOCV r, MAE, #CpGs) SeSaMe β → glmnet elastic net (α=0.5), log-linear age transform (ASM=2,k=0.2), LOOCV, λ.min YES (primary) Fully documented in the Rmd. Reproduce SeSaMe from GEO IDATs, then the exact EN procedure.
Sex clock (#CpGs, accuracy) glmnet binomial elastic net (α=0.5), LOOCV λ.min YES (secondary) Documented in the Rmd.
Test-set age predictions (blood/skin r,MAE) apply trained clock partial depends on training clock; report if clock reproduces.
FelidClock multi-species clock (46 CpGs, r=0.94) EN on cheetah+lion+tiger NO — out of scope Lion/tiger methylation are EXTERNAL data not in GSE310779; not obtainable from this accession.
SOS disease DMC (4269 CpGs), EWAS of age (per-tissue DMC counts) per-CpG differential methylation NO — deferred (hard 20%) Requires reconstructing the authors' exact contrasts/sample subsets & multiple-testing; not the low-hanging output. Skipped, stated.

Known fidelity gaps (will be stated, not hidden)

  • Authors start from intermediate RDS (beta_sesame_shifted.RDS, data_sesame.RDS) and a metadata CSV (cheetah_metadata_vod2.csv) that are NOT shipped in the repo. We re-derive β with SeSaMe openSesame(platform="Mammal40") from the GEO IDATs — a faithful but not byte-identical preprocessing.
  • A manual "shifted" correction on submission-1 betas and ComBat(batch=study) cannot be reproduced exactly (no shipped study labels / shift vector). We process all IDATs uniformly and (optionally) ComBat by tissue-derived study label.
  • Hand-removed outliers are listed by author-SID (e.g. ET0394TOX00092), which do not map to GEO barcodes without the unshipped CSV → we apply SeSaMe QC + NA filtering instead of the identical hand list.

Compute: trivial (SeSaMe on 57 IDAT pairs + glmnet on ~36k×~45 matrix) — minutes on one «our HPC» std node. All compute on «our HPC»; data on «infra».

Figures / tables: Fig 1AS2 TableS5 Table
C1
Reported
age clock LOOCV r = 0.97
Reproduced
0.71 (raw beta) / 0.667 (raw+ComBat), n=52
partial
C2
Reported
age clock LOOCV MAE = 0.86 yr
Reproduced
1.82 yr (raw) / 2.36 yr (ComBat)
partial
C3
Reported
age clock 52 CpG sites
Reproduced
82 (raw) / 27 (ComBat)
partial
C4
Reported
sex clock 67 CpG sites
Reproduced
96
partial
C5
Reported
sex clock accuracy 100%
Reproduced
90.2% LOOCV (46/51)
partial
META
Reported
training set 52 (liver+blood)
Reproduced
52 (45 liver + 7 blood)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 58/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is a solid partial reproduction: the identical public IDATs (GSE310779) were decoded and the repo's elastic-net age+sex clock was run verbatim on the exact same 52 liver+blood training samples, and the clock unambiguously works (age LOOCV r=0.71, p<1e-4; sex 90.2%, 46/51). The headline metrics were not matched (r 0.97→0.71, MAE 0.86→1.82 yr, sex 100%→90%, 52→82 CpGs), but every gap is explained by reproducible, declared deviations on the input/preprocessing side — the SeSaMe noob/pOOBAH normalization could not be run (Bioconductor OSN CDN globally down) and the authors' 'shifted' correction plus the SID→GSM outlier-mapping CSV were not deposited. The deviation is therefore moderate, sits in preprocessing rather than the modeling logic, and shows no fabrication signal; the central conclusion (a valid cheetah methylation clock exists in this data) holds in limited form.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

424.4 k
tokens (I/O) · 50.8 M incl. cache
72 min
runtime · 0.04 CPU-h
4.4 GB
peak RAM
6
HPC jobs
hummel
machine