Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Caloric Restriction Reprograms Adipose Tissues in Rhesus Monkeys.

Aging Cell · 2025
66/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
66/100
Reproducibility score
0.5 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 28% of all assessed papers rank 830 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL (strong, honest) reproduction of Clark et al. 2025, Aging Cell — rhesus adipose caloric restriction. Full faithful pipeline run on «our HPC»: Skewer->STAR/RSEM (Ensembl Mmul_10 r110)->edgeR on the paper's OWN data. DATA-ACCESSION CORRECTION: the brief's GSE186466 is WRONG (it is a HUMAN cancer-cachexia study); the paper's real data is BioProject PRJNA1337456 = 16 Macaca mulatta paired-end runs (SAT/VAT x Control/CR, 4 each), confirmed via ENA labels matching the design. KEY METHODOLOGICAL FINDING: the edgeR TEST CHOICE is the dominant degree of freedom. The paper states only 'edgeR v4.0.16' (test unspecified). The classic glmLRT/exactTest path reproduces the paper CLOSELY; the conservative glmQLFTest does NOT. Headline C3 (CR-response, both depots): reported 118 vs reproduced 119 (glmLRT) = near-exact. Within-tolerance: C2 (482 vs 459), C6 (834 vs 856), C9 total genes (14749 vs 15161). Partial: C1 (30 vs 24), C4 (1322 vs 1679), C5 SAT-adj (100 vs 62), C7 VAT-adj (3 vs 1; both ~0 = 'VAT nearly unresponsive' qualitatively reproduced), C8 (637 vs 560). NO outright mismatches. Central qualitative claim CONFIRMED: SAT responds far more strongly to CR than VAT (adj DE SAT 62 vs VAT 1; paper 100 vs 3). Grades: 4 within-tol, 5 partial, 0 mismatch. NOT ATTEMPTED (hard 20%, per 80/20): DEXseq exon usage, WGCNA modules, CALERIE cross-species, wet-lab phenotypes. No values fabricated; every reproduced count comes from the actual «our HPC» edgeR run.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 66
    assessed: 2026-06-19 ⛓ 00ce4a27a56d
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Adipose tissue changes are implicated in the health benefits of caloric restriction (CR) in primates, but the molecular details are unknown; the study tests whether life-long CR induces shared and depot-specific transcriptional adaptations in subcutaneous (SAT) and visceral (VAT) adipose tissue of aged rhesus monkeys, and whether these responses are conserved with humans.

Core claims
  • At baseline, SAT and VAT transcriptomes are highly similar, with only ~1% of genes (30 genes, adjusted p<0.05) differentially expressed between depots in Controls finding
  • CR induces depot-specific transcriptional responses, with SAT showing far more significantly DE genes (100) than VAT (3) finding
  • RNA processing and proteostasis-related pathways are enriched with CR in both depots, representing a shared adaptation finding
  • Metabolic, growth, and inflammatory pathway changes in response to CR are depot-specific rather than shared finding
  • At baseline, SAT is enriched for metabolic/homeostatic pathways (oxidative phosphorylation, lipid metabolism, lysosome) while VAT is enriched for growth and immune/inflammatory pathways finding
  • CR specifically engages RNA splicing/exon usage changes in SAT (156 exons/129 genes) but minimally in VAT (5 exons/5 genes) finding
  • CR alters adipose tissue cellular composition, with shared reductions in granulocytes and increases in endothelial cells across both depots, and depot-divergent changes in dendritic cells, monocytes, macrophages, T cells, and epithelial cells finding
  • Depot differences and CR responses identified in rhesus monkey adipose are highly conserved with human adipose tissue data finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq / differential gene expression analysis subcutaneous and visceral adipose tissue, male rhesus monkeys (~25 yrs) 30% caloric restriction (life-long) vs Control diet differentially expressed genes between depots and diets
GSEA with KEGG pathway mapping SAT and VAT transcriptomes, rhesus monkeys CR vs Control diet enriched metabolic/immune/growth pathways
xCell in silico cell-type deconvolution SAT and VAT transcriptomes, rhesus monkeys CR vs Control diet relative enrichment of immune/stromal cell types xCell
DEXseq differential exon usage analysis SAT and VAT, rhesus monkeys CR vs Control diet exon-level usage changes (splicing) independent of total transcript levels DEXseq
diffSplice (limma/edgeR) transcript isoform analysis SAT and VAT, rhesus monkeys CR vs Control diet differential exon usage via sequence read alignment divergence limma/edgeR
Weighted gene co-expression network analysis (WGCNA) SAT and VAT transcriptomes, rhesus monkeys (both diets) CR vs Control diet gene modules correlated with biometric/clinical traits
Multiple factor analysis (MFA) of biometric/clinical/disease-risk indices rhesus monkeys, whole-animal phenotyping CR vs Control diet integration of health/metabolic variables across individuals
Clinical blood chemistry/hematology panel rhesus monkeys, blood/plasma CR vs Control diet body weight, fat mass, %fat, insulin, cholesterol, glucose, blood counts, kidney/liver/muscle indices
Key results
  • 99% of SAT/VAT genes not differentially expressed at baseline; 30 genes DE at adjusted p<0.05, 482 at unadjusted p<0.05 ~1% DE (30/14,749 genes)
  • Combined-depot CR response yielded 118 DE genes (q<0.05) and 1322 (9.2%) at unadjusted p<0.05 118 genes q<0.05
  • SAT-only CR analysis identified 100 significant DE genes (834 unadjusted); VAT-only identified 3 significant DE genes (637 unadjusted) 100 vs 3 DE genes
  • CR animals had lower body weight, fat mass, and percent fat than Controls ~10 vs ~13 kg body weight; 18% vs ~35% body fat
  • SAT showed extensive differential exon usage with CR (156 exons/129 genes) versus minimal in VAT (5 exons/5 genes) 129 vs 5 genes
  • CR lowered granulocyte representation and raised endothelial cell representation in both depots; other immune cell types diverged by depot
  • CR lowered expression of inflammation-associated adipokine genes MMP7 and ANGPT1; MMP15 and ITLN1 significantly lower in SAT only; MMP3 upregulated in SAT only
  • Shared CR-enriched pathways (ribosome, drug metabolism) found in combined and both depots; TNF/JAK-STAT enriched in VAT only; spliceosome/proteosome enriched in SAT only
Key statistics
  • count 30 genes adjusted p<0.05 (482 unadjusted p<0.05) (baseline SAT vs VAT DE genes in Control monkeys)
  • count 118 DE genes q<0.05 (1322, 9.2% unadjusted p<0.05) (combined-depot CR vs Control DE genes)
  • count 100 DE genes significant (834 unadjusted) (SAT-only CR vs Control DE genes)
  • count 3 DE genes significant (637 unadjusted) (VAT-only CR vs Control DE genes)
  • count 156 exons / 129 genes (adjusted p<0.1) (SAT differential exon usage (DEXseq) with CR)
  • count 5 exons / 5 genes (VAT differential exon usage (DEXseq) with CR)
  • mean 18% vs ~35% body fat (percent body fat, CR vs Control)
  • count n = 4 per diet, both depots (cohort size for SAT/VAT sampling)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used bulk RNA-seq from SAT and VAT biopsies of n=4 CR and n=4 Control male rhesus monkeys (~25 years) to characterize transcriptional responses to lifelong caloric restriction. Differential expression (DE) was assessed at the gene level with FDR-adjusted q-values, and pathway enrichment was evaluated via GSEA and ORA using KEGG mappings, with relaxed thresholds (q<0.1) for some secondary analyses. Integrative approaches including WGCNA module-trait correlations, MFA of biometric/clinical data, DEXseq for exon-level differential usage, and xCell in silico cell-type deconvolution complemented the DE results. Results were reported primarily as log2 fold changes on MA plots with q-value significance thresholds rather than with exact p-values or effect-size confidence intervals.

Replicationbiological Sample sizen=4 per diet group stated; no formal power analysis or sample-size justification mentioned in the available text GroupsCR vs Control (between-animal); SAT vs VAT (within-animal, matched depots) Pairingmixed Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionFDR (q-values); specific algorithm not named in text
Statistical tests used
Test Applied to n Assumptions
Differential expression analysis (tool not explicitly named in text; FDR/q-values reported) SAT vs VAT (Control only); CR vs Control (combined depots, SAT only, VAT only) n=4 per diet group per depot; n=8 for combined-depot analysis not stated
Gene Set Enrichment Analysis (GSEA) with KEGG mapping Depot comparison (Figure 1B, q<0.1); CR vs Control per depot and combined (Figure 3A, q<0.05) n=4 per group not stated
Over-Representation Analysis (ORA) with KEGG mapping Genes with differential exon usage in SAT (Figure 3D,E); WGCNA module genes (Figure 4C) gene lists derived from n=4 per group not stated
DEXseq differential exon usage test Exon-level CR vs Control comparison in SAT and VAT separately (Figure 3C, adjusted p<0.1) n=4 per diet per depot not stated
diffSplice (limma/edgeR) for differential isoform usage Transcript isoform-level CR response in SAT (Figure 3E) n=4 per diet per depot not stated
xCell in silico cell-type enrichment (p<0.05 threshold) Cell-type composition comparison between depots and diets (Figure 2G) n=4 per group not stated
WGCNA module-trait correlation (method not further specified) Correlation of 36 co-expression modules with biometric/clinical variables (Figure 4B) n=8 animals (both depots, both diets) not stated
Multiple Factor Analysis (MFA) Integration of biometric and clinical indices across individuals (Figure 4A) n=8 animals not stated
Significance test for biometric/clinical variables (type not stated) Body weight, fat mass, % fat, insulin, cholesterol, glucose, blood counts between CR and Control (Figure S1) n=4 per diet group not stated
Approaches that could also have been used
  • The biometric comparisons (body weight, fat mass, insulin, etc.) between CR and Control used an unspecified significance test with n=4 per group and no dispersion measures reported
    Could also: A non-parametric test (e.g., Mann-Whitney U / Wilcoxon rank-sum) with reported median and IQR, or a permutation test, could also be used given the very small sample size; reporting with individual data points overlaid on the summary statistic would additionally convey the full distribution — With n=4 per group, normality assumptions are difficult to verify; non-parametric or permutation-based approaches make fewer distributional assumptions, and showing individual data points prevents summary statistics from masking the small-n context
  • SAT and VAT were analyzed both separately and combined across depots in parallel DE pipelines, producing multiple overlapping family-of-tests without an overarching multiplicity framework
    Could also: A linear mixed-effects model (e.g., via limma-voom or DESeq2's LRT with depot as a covariate or interaction term) could simultaneously model the diet × depot interaction in a single unified framework — A single model with a diet × depot interaction term would directly test whether the CR effect differs between depots, controlling for within-animal correlation (SAT and VAT from the same monkey), and would reduce the number of parallel analyses needing separate FDR control
  • Dispersion of gene expression and biometric outcomes was not reported (no SD, SEM, CI, or range given for any continuous measurement)
    Could also: Reporting SD or 95% CI alongside point estimates (mean or median) is a standard complement, particularly relevant for small n where variability is a key interpretive consideration — With n=4 per group, individual variation substantially influences the reliability of group-level estimates; dispersion measures allow readers to assess effect magnitude relative to within-group spread
  • Cell-type composition estimated by xCell in silico deconvolution was tested with a p<0.05 threshold without mention of multiple-comparison correction across the many cell types compared
    Could also: Applying FDR correction across cell-type comparisons (e.g., Benjamini-Hochberg) or using a method that jointly models all cell types (e.g., Dirichlet regression on deconvolved proportions) could also be used — Testing many cell types simultaneously inflates the chance of false positives; a correction method consistent with that used for the DE analyses would maintain coherent error-rate control across the paper
  • WGCNA module-trait correlations were computed across n=8 samples (4 CR + 4 Control, both depots treated as independent observations), without explicit accounting for the paired depot structure within each animal
    Could also: A mixed-effects correlation or permutation-based module-trait association that respects the within-animal pairing of SAT and VAT could also be applied — SAT and VAT from the same animal are not independent; treating them as independent in correlation analyses may underestimate uncertainty in module-trait associations given the small effective sample size
  • Pathway enrichment was performed via both GSEA (rank-based) and ORA (threshold-based gene lists) across several comparisons
    Could also: A single consistent enrichment approach (e.g., GSEA throughout, or fgsea for speed/reproducibility with small n) applied uniformly would also be a standard choice — ORA results depend on the choice of significance threshold used to define the 'hit list,' whereas GSEA uses the full ranked list; using one method consistently aids interpretability and avoids threshold-dependent results influencing pathway conclusions
Software: DEXseq · limma/edgeR (diffSplice) · GSEA · WGCNA · xCell · MFA (FactoMineR or equivalent)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41042069

Paper: Clark JP et al. Caloric Restriction Reprograms Adipose Tissues in Rhesus Monkeys. Aging Cell 2025. PMID 41042069 · PMCID PMC12686577 · DOI 10.1111/acel.70254.

Data accession correction (IMPORTANT)

The brief lists geo:GSE186466 as the dataset. That is wrong — GSE186466 is "Transcriptome-wide maps of subcutaneous/visceral adipose tissues in cancer-associated cachexia patients" (Homo sapiens, 6 samples, SRP342862 / PRJNA773957) — a different study. The rhesus CR paper's real data is BioProject PRJNA1337456 (Macaca mulatta, 16 paired-end runs, ~112 Gbases), recoverable from the paper's data-availability statement. All 16 ENA runs resolve with clear group labels (SAT/VAT × Control/Restriction, 4 each), exactly matching the paper's design. We reproduce against PRJNA1337456, not GSE186466.

Pipeline (from Methods)

Skewer v0.1.123 (trim) → STAR v2.5.0a (align, Macaca mulatta genome) → RSEM (quantify) → edgeR v4.0.16 (differential expression). DEXseq for exon usage; WGCNA for modules. Sequencing 2×100 bp HiSeq2000, ~75 M reads/sample.

In scope (pipeline-derived, attempted)

The clearly-specified, low-hanging outputs are the edgeR DEG counts per comparison and the total detected-gene count. All flow from one STAR+RSEM gene count matrix → edgeR, i.e. the paper's own pipeline run on the paper's own data.

id reported result paper location
C1_depot_adj 30 DE genes SAT-vs-VAT in Controls (adj p<0.05) Results, depot comparison
C2_depot_unadj 482 genes SAT-vs-VAT (unadj p<0.05) Results, depot comparison
C3_CR_adj 118 genes CR response, both depots (q<0.05, <1%) Results, CR response
C4_CR_unadj 1,322 genes CR response, both depots (unadj p<0.05, 9.2%) Results
C5_SAT_CR_adj 100 DE genes SAT CR response (adj) Results, SAT
C6_SAT_CR_unadj 834 genes SAT CR (unadj p<0.05) Results, SAT
C7_VAT_CR_adj 3 DE genes VAT CR response (adj) Results, VAT
C8_VAT_CR_unadj 637 genes VAT CR (unadj p<0.05) Results, VAT
C9_total_genes 14,749 genes identified across all specimens Results

The biologically salient, robust qualitative claim: SAT responds far more strongly to CR than VAT (100 vs 3 adj DE genes); VAT is nearly unresponsive.

Out of scope / not attempted (the hard ~20%)

  • DEXseq exon usage (SAT 156 exons/129 genes; VAT 5/5) — exon-level model, separate pipeline, underspecified; not attempted per 80/20.
  • WGCNA 36 modules — soft-threshold/module params not pinned; not attempted.
  • CALERIE human cross-species ~72% pathway congruence — needs external human dataset + pathway mapping; out of scope (external data).
  • Wet-lab/histology/metabolic-phenotype results — not pipeline-derived.

Degrees of freedom (expected to prevent bit-identical match → grade honestly)

  • Trimmer: paper uses Skewer; we match Skewer where feasible.
  • STAR 2.5.0a (2015) vs current; RSEM version; reference annotation version (we use Ensembl Mmul_10; paper says only "Macaca mulatta reference genome").
  • edgeR filtering / exact test vs QLF, dispersion estimation, design matrix for the "both depots combined" model. The paper underspecifies these.
  • Therefore exact DEG counts may differ; we report counts + grade exact/within-tol/partial/mismatch and emphasize the qualitative pattern.
C1_depot_adj
Reported
30
Reproduced
24 (glmLRT; glmQLFTest alt=2)
partial
C2_depot_unadj
Reported
482
Reproduced
459 (glmLRT; glmQLFTest alt=511)
within tolerance
C3_CR_adj
Reported
118
Reproduced
119 (glmLRT; glmQLFTest alt=6)
within tolerance
C4_CR_unadj
Reported
1322
Reproduced
1679 (glmLRT; glmQLFTest alt=1804)
partial
C5_SAT_CR_adj
Reported
100
Reproduced
62 (glmLRT; glmQLFTest alt=2)
partial
C6_SAT_CR_unadj
Reported
834
Reproduced
856 (glmLRT; glmQLFTest alt=867)
within tolerance
C7_VAT_CR_adj
Reported
3
Reproduced
1 (glmLRT; glmQLFTest alt=0)
partial
C8_VAT_CR_unadj
Reported
637
Reproduced
560 (glmLRT; glmQLFTest alt=487)
partial
C9_total_genes
Reported
14749
Reproduced
15161 (glmLRT; glmQLFTest alt=15161)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

723.6 k
tokens (I/O) · 50.2 M incl. cache
417 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.