Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Tumor methionine metabolism drives T-cell exhaustion in hepatocellular carcinoma.

Nat Commun · 2021
L1 78/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 51% of all assessed papers rank 533 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction (healthy, compute ran end-to-end). Paper: Hung et al. Nat Commun 2021 (PMID 33674593), 'Tumor methionine metabolism drives T-cell exhaustion in HCC'. In-scope pipeline = ATAC-seq peak-calling with the paper's own third-party tool Genrich (github.com/jsh58/Genrich, verbatim Methods args) on the authors' OWN data GSE166213 (12 runs SRR13633825-836; the BRIEF's GSE98638 is a peripheral reused scRNA-seq atlas, NOT the pipeline data). Pipeline: cutadapt -> bowtie2 hg38 --very-sensitive -X2000 (align 96.2-98.5%) -> Picard MarkDuplicates [@RG] (C2) -> Genrich (C1/C3). This room RE-RAN the full pipeline from scratch (prior room's «infra» workdir was janitor-reclaimed) to recover the two items the prior room left un-fetched. RESULTS: C2 (72-79% non-dup) REPRODUCES the reported range almost exactly - observed 72.20-78.84%, all 12 samples in-band. C3 (SAM/MTA reduce CD8 accessibility) REPRODUCES cleanly - control(23k) >> SAM/MTA(13-14k) across all reps. C1 (headline 135,583 peaks) reproduces in regime/direction but not exactly - SUM of per-sample peaksets = 208,934 (~1.54x), UNION = 28,660 (rules out union/consensus); gap attributable to the paper pinning no Genrich version or peak thresholds. 2 of 3 claims reproduce; the single headline count is ~1.54x => honest overall = partial. NO fabrication concern (208,934 is a plausible Genrich output on this data). NOT attempted (out of scope): wet-lab/mechanistic (CRISPR/metabolomics/flow/mouse), TCGA + 13 reused public GEO survival/expression re-analyses, DiffBind/ChIPseeker/motif downstream. Version deltas vs paper (cutadapt 5.2/1.18, bowtie2 2.5.x/2.2.6, picard 3.4.0/2.18.26, Genrich 0.6.2/unstated) in environment.lock.

💻 Code ↗ 🗄 Data: GSE98638

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 75
    assessed: 2026-06-20 ⛓ 6469245c8b63
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether and how tumor cell (hepatocellular carcinoma) metabolism, specifically reprogrammed methionine recycling, actively drives CD8+ T-cell exhaustion and thereby promotes tumor immune evasion.

Core claims
  • A transcriptome-derived T-cell exhaustion score (ES) is prognostic for HCC patient survival independent of known clinical/molecular factors finding
  • Elevated tumor SAM and MTA, products of an oncogenically reprogrammed methionine salvage pathway, are tightly linked to T-cell exhaustion in HCC finding
  • SAM and MTA directly induce CD8+ T-cell dysfunction (reduced proliferation, increased exhaustion markers) in vitro finding
  • CRISPR-Cas9-mediated deletion of MAT2A (a key SAM-producing enzyme) reduces T-cell dysfunction and inhibits HCC tumor growth in mice finding
  • HCC tumors preferentially upregulate the methionine salvage pathway and downregulate the de novo pathway, raising the salvage-to-de novo expression ratio mechanism
  • Genes in the methionine salvage (MTAP, SRM, SMS, APIP) and de novo (MTR, BHMT, BHMT2) pathways show tumor-specific somatic copy number alterations correlated with SAM/MTA abundance and T-cell exhaustion mechanism
  • Serum MTA level correlates with tumor MTA content and predicts HCC patient survival, validated in an independent cohort finding
  • 82 T-cell exhaustion-specific genes derived from single-cell RNA-seq of HCC-associated T cells (Supplementary Data 1) serve as a resource for scoring exhaustion in bulk tumor data resource
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA-seq HCC-associated T cells (GSE98638, 5063 cells) none identification of T-cell exhaustion-specific genes
bulk transcriptome hierarchical clustering / Cox regression survival analysis HCC tumor tissue (TIGER-LC n=62, LCI n=247, TCGA-LIHC n=366 cohorts) none exhaustion score, clinical/molecular subgroup association, overall survival
metabolomics HCC tumor and paired non-tumor tissue (TIGER-LC cohort) none levels of 718 metabolites correlated with exhaustion score
single-cell RNA-seq malignant cells and associated T cells from 4 HCC patients (GEO125449, 5112 cells, 534 malignant cells, 175 T cells) none salvage vs de novo methionine metabolic status of tumor cells and T-cell exhaustion/methylation gene expression
somatic copy number alteration (SCNA) analysis HCC tumor vs matched non-tumor tissue (TIGER-LC, LCI, TCGA-LIHC cohorts) none copy number alterations in methionine salvage/de novo pathway genes
serum and tumor metabolomics HCC patient serum and tumor tissue (TIGER-LC n=51; validation cohort n=102 Chinese patients) none serum MTA level vs tumor MTA/salvage pathway activity and patient survival
in vitro T-cell functional assay (CFSE proliferation, flow-cytometric activation/exhaustion marker staining) human CD8+ T cells from healthy donors, anti-CD2/CD3/CD28 bead-stimulated drug treatment (SAM or MTA, up to 200 µM, or 100 µM for functional studies) cell viability, proliferation, CD28/CD44 activation markers, PD1/TIM3 exhaustion markers over time (Day 3, 7, 14)
CRISPR-Cas9 gene knockout / in vivo tumor growth assay HCC mouse model CRISPR-Cas9-mediated deletion of MAT2A T-cell dysfunction and HCC tumor growth CRISPR-Cas9
Key results
  • MTA level correlates with exhaustion score ρ=0.444, p=0.000685
  • SAM level correlates with exhaustion score ρ=0.437, p=0.000841
  • Exhaustion score predicts HCC survival independent of common molecular subtypes, staging, cirrhosis Hazard ratio=3.27 (95% CI 1.85-5.79), p<0.001
  • Serum MTA correlates with tumor MTA content Pearson r=0.331, p=0.0165
  • T cells from salvage-dominant tumors show higher exhaustion than those from de novo-dominant tumor p=0.04
  • SAM/MTA up to 200 µM do not cause acute CD8+ T-cell cytotoxicity at 3 days
  • SAM or MTA (100 µM) treatment reduces CD8+ T-cell proliferation at Day 7 and later, with increased exhaustion markers (PD1/TIM3+, CD28-)
  • SAM/MTA level positively associated with salvage-to-de novo pathway expression ratio in tumors
Key statistics
  • correlation ρ=0.444, p=0.000685 (MTA level vs exhaustion score, TIGER-LC cohort)
  • correlation ρ=0.437, p=0.000841 (SAM level vs exhaustion score, TIGER-LC cohort)
  • other Hazard ratio=3.27 (95% CI 1.85-5.79), p<0.001 (Exhaustion score predicts HCC survival, multivariant Cox regression, TIGER-LC cohort)
  • correlation Pearson r=0.3310, p=0.0165 (serum MTA vs tumor MTA level)
  • pvalue p=0.04 (T-cell exhaustion difference between salvage-dominant vs de novo-dominant tumors (t-test))
  • count 718 metabolites detected, 21 significantly correlated with ES (FDR<0.05) (tumor metabolomics, TIGER-LC cohort)
  • count 675 HCC patients (combined cohort size across ethnicities/etiologies for exhaustion-survival link)
  • count 82 T-cell exhaustion-specific genes identified from 5063 T cells (single-cell transcriptome analysis, GSE98638)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper combines observational analyses of multiple HCC patient cohorts (bulk and single-cell transcriptomics, metabolomics, copy-number data) with in vitro CD8+ T-cell functional assays. Associations between a derived exhaustion score, methionine metabolite levels, and gene expression were assessed mainly with Spearman/Pearson correlations, group differences with independent t-tests, and patient survival with Kaplan-Meier curves and Cox proportional hazards/log-rank tests; a metabolite screen used FDR correction.

Replicationbiological Sample sizeSample sizes are stated per analysis/cohort (e.g., TIGER-LC n=62, LCI n=247, TCGA-LIHC n=366, single-cell n=534/175, serum cohorts n=51/102), but no a priori power calculation is described in the excerpted text. GroupsExhaustion clusters/high vs low exhaustion score; tumor vs matched non-tumor tissue; salvage- vs de novo-dominant tumors/cells; high vs low serum MTA or salvage-to-de novo ratio Pairingmixed Randomization/blindingnot stated Dispersionmixed Exact p-valuesyes Effect sizesyes Confidence intervalsyes Multiplicity correctionFalse discovery rate (FDR) correction, threshold FDR < 0.05
Statistical tests used
Test Applied to n Assumptions
Two-sided log-rank test (Kaplan-Meier survival analysis) Survival by exhaustion cluster/score (Fig. 1c,1d), salvage-to-de novo ratio (Fig. 3a), and serum MTA level (Fig. 3d) Varies by cohort: TIGER-LC n=62, LCI n=247/102, TCGA-LIHC n=366, TIGER-LC serum n=51 not stated
Cox proportional hazards model / multivariate Cox regression Selection of exhaustion-score gene combination and independent prognostic value of the exhaustion score (Fig. 1a; Supplementary Tables 2-3) TIGER-LC cohort, n=62 not stated
Two-sided Spearman's rank correlation coefficient test Exhaustion score vs cytolytic score (Fig. 1e); exhaustion score vs metabolite levels (Fig. 2a); exhaustion score vs salvage/de novo pathway expression (Fig. 2d); SAM/MTA vs salvage-to-de novo ratio (Fig. 2e); serum vs tissue MTA (Fig. 3b) n=62 for tumor-based correlations; n=52 for serum/tissue MTA not stated
Pearson correlation Serum MTA vs tumor MTA level (Fig. 3b, r=0.3310, p=0.0165) n=52 not stated
Two-sided independent t-test T-cell exhaustion/methylation/metabolic gene expression in salvage- vs de novo-dominant tumors (Fig. 2g); serum MTA by tumor salvage-to-de novo status (Fig. 3c) 175 T cells (Fig. 2g); n=20 low / n=26 high salvage-to-de novo group (Fig. 3c) not stated
FDR-corrected correlation screen (multiple Spearman correlations, FDR<0.05) Correlation of exhaustion score with 718 detected tumor metabolites (Fig. 2a, Supplementary Table 5) TIGER-LC cohort metabolomics dataset not stated
Approaches that could also have been used
  • Many individual Spearman/Pearson correlations and t-tests are reported across different figures and cohorts, with an explicit FDR correction applied only to the 718-metabolite screen.
    Could also: A systematic multiple-testing correction (e.g., Benjamini-Hochberg FDR or Bonferroni) could also be applied uniformly across the full set of individual correlation and group comparisons presented in the paper. — This would provide an additional, consistent way to control the false discovery or family-wise error rate when many statistical comparisons are drawn from overlapping or related datasets.
  • Group differences in gene/metabolite levels (e.g., Fig. 2g, Fig. 3c) are assessed with two-sided independent t-tests.
    Could also: A non-parametric alternative such as the Mann-Whitney U test could also be used, especially for smaller samples or data that may not be normally distributed. — Non-parametric tests do not require an assumption of normally distributed data and can be more robust for biological measurements such as gene expression or metabolite concentrations.
  • Survival analyses split continuous biomarkers (exhaustion score, salvage-to-de novo ratio, serum MTA) into two groups at the median for Kaplan-Meier/log-rank testing.
    Could also: Modeling the biomarker as a continuous variable in a Cox regression (as was already done for the exhaustion score in Supplementary Table 3), or using data-driven cutpoint methods such as maximally selected rank statistics, could also be used. — This can retain more information than a single median split and better characterize a dose-response relationship between biomarker level and survival.
  • Associations between the exhaustion score, methionine metabolites, and gene expression are largely characterized through pairwise correlation coefficients.
    Could also: A multivariable regression model that incorporates several clinical and molecular covariates simultaneously could also be used to complement the pairwise correlation analyses. — This could help account for potential confounding or shared variance among the correlated variables being studied.
  • Fig. 3c summarizes serum MTA level distributions using a boxplot showing minimum/maximum, median, and 25th/75th percentiles.
    Could also: Reporting group means with SD or SEM, or means with 95% confidence intervals alongside individual data points, could also be used to convey central tendency and spread. — This offers a complementary summary style, particularly useful when comparing group means directly or when sample sizes are small.
  • Patients are grouped into three exhaustion clusters (ECs) using unsupervised hierarchical clustering of exhaustion-gene expression.
    Could also: Alternative unsupervised methods such as k-means, consensus clustering, or model-based clustering could also be applied to derive or cross-validate the cluster assignments. — Comparing results across multiple clustering approaches can provide additional information about the stability of the identified exhaustion subgroups.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

C1
Reported
135,583 ATAC-seq peaks identified from 12 samples
Reproduced
SUM of 12 per-sample Genrich peaksets = 208,934 (paper verbatim args); UNION (bedtools merge) = 28,660; pooled-12-BAM Genrich not run (bounded << 135,583)
partial
C2
Reported
72-79% non-duplicate mapped reads across samples
Reproduced
Observed non-dup range across 12 samples = 72.20% to 78.84% (every sample inside the reported 72-79% band); Picard MarkDuplicates with @RG
within tolerance
C3
Reported
Methionine metabolites (SAM/MTA) reduce CD8 T-cell chromatin accessibility
Reproduced
Per-condition mean Genrich peaks nonact=23,090 mock=19,428 SAM=13,918 MTA=13,209; clear control>>treated gradient across all 3 reps
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

Reproduction used the authors' own deposited data (GSE166213, 12/12), so data identity and endpoint comparability are strong (green). The central conclusion — SAM/MTA reduce CD8 chromatin accessibility — reproduces cleanly with a consistent control>>treated gradient across all replicates (q7 green). The headline peak count deviates ~1.54x (208,934 vs 135,583), but stays the same order of magnitude and is best explained by the paper's unpinned Genrich version and peak-call threshold plus our newer tool stack — a version/underspecification gap on the methods side, not fabrication (q5/q6/q8 yellow). C2 (72-79% non-dup) remains pending a «our HPC» fetch.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

<synthetic>

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

1.1 M
tokens (I/O) · 77.6 M incl. cache
710 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.