Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Histone deacetylase SIRT6 regulates tryptophan catabolism and prevents metabolite imbalance associated with neurodegeneration.

Nat Commun · 2025
L1 93/100 PQI 98
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
93/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 85% of all assessed papers rank 154 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> 1:1. Authors' own R repo (github.com/SIRT6/Kaluski_et_al_2024 @28187de, Zenodo 17250775) ships both data and code; GEO GSE221077 not needed. Re-ran the Drosophila DESeq2 differential-expression pipeline (DESeqDataSetFromMatrix ~Factors + filterByExpr_custom + DESeq + 4 contrasts + lfcShrink ashr) from shipped raw_counts_drosophila.csv on «our HPC»; all four contrasts' DE-gene counts (up/down/total) and the 18751-gene test universe matched the committed notebook outputs EXACTLY, even though our conda env resolved DESeq2 1.50.2/R 4.5.3 vs the authors' 1.44.0/R 4.4.1 -- so the result is data-derived and version-robust. Fig2B hypergeometric phyper p reproduced exactly (deterministic). Fig1D rstatix t-tests reproduced (mESC/HeLa/SH-SY5Y) but the notebook never prints the numeric p (only figure significance brackets), so graded partial -- no printed reference, not a mismatch. NOT attempted: circadian wet-lab re-plots (non-pipeline, manually tabulated), GSEA tail of Suppl_Fig6_7 (permutation RNG), MetaboAnalyst GUI normalization upstream of Fig1B/C. No fabrication red flags.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 93
    assessed: 2026-06-16 ⛓ 09f91e99ffc0
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper tests whether the NAD+-dependent histone deacetylase SIRT6 regulates tryptophan catabolism in the brain, and whether its loss causes a metabolic shift toward neurotoxic kynurenine-pathway metabolites (at the expense of serotonin/melatonin) that drives circadian/sleep disruption and neurodegeneration.

Core claims
  • SIRT6 is an evolutionarily conserved regulator of tryptophan catabolism that balances tryptophan usage between the kynurenine and serotonin/melatonin pathways finding
  • SIRT6 deficiency shifts tryptophan metabolism toward the kynurenine pathway, elevating neurotoxic metabolites (KA, QA) while decreasing serotonin and melatonin mechanism
  • SIRT6 directly binds the promoter regions of TDO2, IDO1, and AANAT and transcriptionally regulates rate-limiting tryptophan-catabolism enzymes mechanism
  • SIRT6 loss disrupts melatonin oscillation, Aanat expression rhythm, and circadian/sleep quality finding
  • Redirecting tryptophan via TDO2 inhibition rescues neuromotor behavior defects and brain vacuolization in a SIRT6 KO Drosophila model finding
  • Metabolomic changes in SIRT6-deficient cells significantly overlap with CSF metabolome of human patients with AD, schizophrenia, and brain inflammatory diseases finding
  • SIRT6 KO increases expression of the brain tryptophan transporter Slc7a5 and upregulates Tdo2, Ido1, Ido2, Kynu while downregulating Kmo, Aanat, and Asmt finding
Experimental setups
Assay System Perturbation Readout Platform
Untargeted/targeted metabolomics (over-representation/PCA analysis) Mouse embryonic stem (ES) cells, SH-SY5Y, HeLa, ARPE19 (WT vs SIRT6 KO) SIRT6 KO Tryptophan and derivative metabolite abundances (Tryp, KA, Kyn, QA, AA, serotonin, NAD+); metabolite-set/disease enrichment
RNA-seq Brains of brS6KO (Nestin-Cre) mice and WT Brain-specific SIRT6 KO Gene expression changes (e.g., Slc7a5, tryptophan metabolism genes)
Microarray transcriptomics Brains of full-body SIRT6 KO mice and WT Full-body SIRT6 KO Tryptophan metabolism gene expression
Quantitative PCR (qPCR) Brains of brS6KO (Nestin-Cre) mice and WT Brain-specific SIRT6 KO mRNA levels of Tdo2, Ido1/Ido2, Kynu, Kmo, Aanat, Asmt, Tph2, etc. (incl. 24h oscillation)
ChIP-seq and ChIP-qPCR Cells (human SIRT6) none (chromatin binding) SIRT6 binding to TDO2, IDO1, AANAT promoter regions
ELISA Serum from WT and SIRT6 KO mice SIRT6 KO Serotonin concentration
Melatonin measurement (serum, 24h) Serum from WT and brS6KO mice Brain-specific SIRT6 KO Melatonin levels over light/dark cycle
Behavioral/neurodegeneration assay (neuromotor + histology) Drosophila melanogaster SIRT6 KO model SIRT6 KO + TDO2 inhibition rescue Neuromotor behavior, brain vacuolar formation
Key results
  • Tryptophan levels increased in SIRT6 KO ES, HeLa, and SH-SY5Y cells
  • Kynurenic acid (KA) increased in SIRT6 KO ES cells
  • Quinolinic acid (QA) increased in SIRT6 KO ES cells
  • NAD+ reduced in SIRT6 KO ES cells despite elevated precursor QA
  • Slc7a5 tryptophan transporter expression increased in brS6KO brains
  • Tdo2, Ido1 upregulated and Kmo downregulated in brS6KO brains; Aanat and Asmt downregulated
  • Serotonin decreased in serum of SIRT6 KO mice; serotonin/Tryp ratio reduced in HeLa and SH-SY5Y
  • Melatonin oscillation amplitude reduced and failed to increase during dark phase in brS6KO mice
Key statistics
  • pvalue FDR P = 3.86 × 10^-5 (Kynurenic acid change, ES WT vs SIRT6 KO, two-sided t-test)
  • pvalue FDR P = 4.72 × 10^-2 (Quinolinic acid change, ES WT vs SIRT6 KO)
  • pvalue FDR P = 1.52 × 10^-5 (NAD+ change, ES WT vs SIRT6 KO)
  • pvalue P = 8.83e-06 (Tryptophan increase in HeLa SIRT6 KO; ES FDR P=0.000842 (n=3); SH-SY5Y P=0.0276 (n=4))
  • pvalue FDR P = 6.46 × 10^-22 (Slc7a5 expression change, brS6KO brains RNA-seq, two-sided Wald test)
  • pvalue P = 3.14 × 10^-22 (Hypergeometric test for gene overlap in tryptophan metabolism between two mouse models)
  • pvalue P = 0.012 (Asmt expression WT vs brS6KO at ZT18 (WT n=5/ZT18 n=4, brS6KO n=4))
  • pvalue AA/Kyn P=0.0007; Ser/Tryp P=0.0087; KA/Kyn P=0.0251 (SH-SY5Y SIRT6 KO metabolite ratio changes, two-sided unpaired t-test)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This multi-model study (mouse ES cells, human cell lines, brain-specific mouse knockouts, and Drosophila) used metabolomics (targeted and untargeted), transcriptomics (microarray, RNA-seq, qPCR), and ChIP-seq to characterize SIRT6-dependent changes in tryptophan catabolism. Primary between-group comparisons relied on two-sided unpaired t-tests (with FDR adjustment for metabolomics screens) and two-sided Wald tests for RNA-seq differential expression. Results were consistently reported as mean ± SEM with exact P values; no effect sizes or confidence intervals were provided.

Replicationmixed Sample sizeSample sizes stated per figure/gene; no a priori power calculation or sample-size justification reported GroupsSIRT6 KO (or brS6KO) vs wild-type, across ES cells, HeLa, SH-SY5Y, ARPE19 cell lines, and mouse models Pairingunpaired Randomization/blindingnot stated DispersionSEM Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionFDR adjustment (method not specified; implemented via MetaboAnalyst)
Statistical tests used
Test Applied to n Assumptions
Hypergeometric test (one-tailed, FDR-adjusted) for over-representation analysis (ORA) Metabolite-set enrichment against disease-associated CSF metabolomes (Fig 1A) Not stated (metabolite list from ES SIRT6 KO vs WT comparison) not stated
Two-sided t-test with FDR adjustment Metabolite abundance changes in ES SIRT6 KO vs WT cells (Fig 1B, 1C); tryptophan levels across cell lines (Fig 1D) ES n=3, HeLa n=5, SH-SY5Y n=4 replicates not stated
Two-sided Wald test (RNA-seq differential expression) Slc7a5 tryptophan transporter expression in brS6KO vs WT mouse brains (Fig 1E) n=4 mice not stated
Two-sided unpaired t-test Metabolite ratios (AA/Kyn, Sero/Tryp, KA/Kyn) in HeLa and SH-SY5Y SIRT6 KO vs WT (Fig 1F) n=4 replicates not stated
Hypergeometric test Significance of gene expression overlap between full-body SIRT6 KO microarray and brS6KO RNA-seq datasets (Fig 2B) Not stated (gene counts from two transcriptomic datasets) not stated
Two-sided unpaired t-test qPCR of individual tryptophan-pathway genes in brS6KO vs WT mouse brains (Fig 2C) and 24-h oscillation time points (Fig 3C); ELISA serotonin in serum (Fig 3A) qPCR: n=7–12 mice per group (gene-dependent); ELISA: WT n=5, SIRT6 KO n=4 not stated
Approaches that could also have been used
  • Multiple individual gene qPCR comparisons (≥10 genes in Fig 2C; multiple circadian time points in Fig 3C) were each tested with separate two-sided unpaired t-tests
    Could also: A two-way ANOVA (genotype × gene or genotype × time) followed by a post-hoc correction (e.g., Tukey HSD or Benjamini-Hochberg) could also have been applied across this family of tests — Treating the gene panel or time-course as a family and correcting jointly would explicitly control the false-discovery rate across the set, which is especially relevant when many comparisons share a common biological question
  • Dispersion around means is reported as SEM throughout
    Could also: SD or 95% confidence intervals would also convey variability — With small group sizes (n=3–5 for cell lines, n=4–12 for mice), SD directly describes sample spread, and 95% CIs make the precision of each estimate explicit; SEM shrinks with larger n and can visually understate variability in small samples
  • Two-sided unpaired t-tests were used for cell-line comparisons with n=3–5 replicates
    Could also: A non-parametric alternative (e.g., Mann-Whitney U) could also have been applied — With very small n the central-limit-theorem basis for t-test normality is difficult to verify empirically; a non-parametric test makes no distributional assumption and is also widely accepted for metabolomics comparisons with small replication
  • Melatonin oscillation over 24 h was characterized with n=2 mice per genotype at each time point (Fig 3B), and gene expression oscillation was compared with separate unpaired t-tests at each time point (Fig 3C)
    Could also: A mixed-effects model (genotype × circadian time as fixed effects, mouse as random effect) or a two-way repeated-measures ANOVA could also have been used for the time-course data — These approaches jointly model the time structure and genotype effect, increasing power while naturally accounting for the temporal correlation among measurements; they would also yield a formal genotype-by-time interaction test
  • Principal component analysis (PCA) of untargeted metabolomics was used to describe metabolome-wide differences between SIRT6 KO and control cell lines (Supplementary Fig 1C–G)
    Could also: A permutation-based multivariate test such as PERMANOVA (e.g., via vegan in R) could also formally test whether the group centroids differ beyond chance — PCA is an unsupervised visualization tool; PERMANOVA provides an inferential P value for the overall metabolome separation, complementing the visual PCA and the per-metabolite t-tests
  • The FDR-adjustment method used in MetaboAnalyst for metabolite ORA and abundance comparisons is not named
    Could also: Reporting the specific FDR procedure (e.g., Benjamini-Hochberg, Benjamini-Yekutieli) and the universe of tests it covered would also be standard — Different FDR procedures have different assumptions about test independence; naming the method and defining the test family aids reproducibility and lets readers assess how conservative or liberal the correction was
Software: MetaboAnalyst · RNA-seq pipeline with Wald test (software not named; DESeq2 is standard for this test but not stated) · BioRender (figure preparation only)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
4
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE221077 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41345108

Paper: Kaluski-Kopatch et al. (2025) "Histone deacetylase SIRT6 regulates tryptophan catabolism and prevents metabolite imbalance associated with neurodegeneration." Nat Commun. DOI 10.1038/s41467-025-67021-y.

Code (P16, authors' own): https://github.com/SIRT6/Kaluski_et_al_2024 commit 28187de1165b12e04c87045e0b541516928a7bad (2025-10-02). Zenodo mirror DOI 10.5281/zenodo.17250775. R notebooks (IRkernel), data shipped in repo Data/.

Data: All inputs for the in-scope notebooks are shipped IN the repo (Data/). GEO GSE221077 holds the mouse RNA-seq deposit but is not needed for the in-scope computational outputs — the repo ships the relevant processed matrices and the Drosophila raw counts directly.

Pipeline-derived results (IN SCOPE)

The notebooks recompute statistics from shipped data; their committed cell outputs are the reported values we reproduce 1:1.

PRIMARY — Suppl_Fig6_7.ipynb : Drosophila DESeq2 differential expression

Pipeline: DESeq2 1.44.0 DESeqDataSetFromMatrix(design=~Factors) → custom filterByExpr_custom(min_samples=1,min_expression=1)DESeq() → 4 contrasts → lfcShrink(type="ashr") → count DE genes at padj<0.05 (|LFC|>0). Input: Data/Drosophila_experiment/raw_counts_drosophila.csv. Deterministic (ashr is deterministic; no RNG). Concrete reported counts (committed notebook outputs):

  • res_wt (WT_TDO2 vs WT_DMSO): up=10, down=6 ; nrow(sig_wt)=16
  • res_s6 (S6_TDO2 vs S6_DMSO): up=45, down=7 ; nrow(sig_s6)=52
  • res_tdo (S6_TDO2 vs WT_TDO2): up=832, down=628 ; nrow(sig_tdo)=1460
  • res_dmso (S6_DMSO vs WT_DMSO): up=901, down=721 ; nrow(sig_dmso)=1622
  • universe: "out of 18751 with nonzero total read count"; outliers=339

SECONDARY (cheap, deterministic) — Figure2B.ipynb

Hypergeometric enrichment p-value: phyper(7-1,10,15662-10,10,lower.tail=F) → reported 3.142146e-22.

SECONDARY — Figure1D.ipynb

Pairwise t-tests (rstatix, fdr) on tryptophan abundance, S6KO vs WT, in mESC / HeLa / SH-SY5Y from shipped CSVs. Deterministic.

OUT OF SCOPE (not attempted)

  • All wet-lab measurements (Western blots, metabolite LC-MS acquisition, circadian qPCR/ELISA series) — Circadian_oscillation_curves.ipynb only re-plots manually-tabulated experimental values (no pipeline). Out of scope.
  • GSEA/clusterProfiler enrichment plots (Suppl_Fig6_7 tail) — may carry permutation RNG; skipped under 80/20.
  • MetaboAnalyst PCA/volcano upstream (Fig1B/C uses precomputed volcano.csv) — the normalization was done in the MetaboAnalyst GUI, not re-runnable from code; we only re-derive the downstream counts/tests, not the GUI normalization.
Figures / tables: Fig6_7Fig 6Figure2BFigure1D
C1
Reported
WT TDO2-vs-DMSO DE (padj<0.05): up=10 down=6 total=16
Reproduced
up=10 down=6 total=16
exact
C2
Reported
S6 TDO2-vs-DMSO DE: up=45 down=7 total=52
Reproduced
up=45 down=7 total=52
exact
C3
Reported
S6-vs-WT under TDO2 DE: up=832 down=628 total=1460
Reproduced
up=832 down=628 total=1460
exact
C4
Reported
S6-vs-WT under DMSO DE: up=901 down=721 total=1622
Reproduced
up=901 down=721 total=1622
exact
C5
Reported
DESeq2 test universe = 18751 nonzero-count genes
Reproduced
18751
exact
C6
Reported
Fig2B hypergeometric p = 3.142146e-22
Reproduced
3.142146e-22
exact
C7
Reported
Fig1D trp S6KO-vs-WT t-test p (not printed; figure stars)
Reproduced
mESC=8.42e-04; HeLa=8.83e-06; SH-SY5Y=2.76e-02
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 93/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

This is a clean 1:1 reproduction: claims C1–C6 (four DESeq2 contrast DE-counts, the 18751-gene test universe, and the Fig2B hypergeometric p=3.142146e-22) all regenerated exactly from the authors' own shipped data and code, and remained exact even under a newer DESeq2/R toolchain — strong evidence the numbers are genuinely data-derived. The only non-exact item, C7, is graded partial purely because the notebook computes but never prints the Fig1D t-test p-values (only significance brackets); the reproduced values (8.42e-04/8.83e-06/2.76e-02) are consistent with the shown // stars, so this is a missing-reference issue, not a discrepancy. No fabrication red flags on any claim; the deviation, if any, is on neither our methodology nor the authors' side.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

100.1 k
tokens (I/O) · 6.8 M incl. cache
12 min
runtime · 0.02 CPU-h
3.8 GB
peak RAM
1
HPC jobs
hummel
machine