Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

TGF-β-mediated activation of fibroblasts in cervical cancer: implications for tumor microenvironment and prognosis.

PeerJ · 2025
L1 79/100 PQI 93
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
79/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 55% of all assessed papers rank 514 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for the SHIPPED-data portion. The headline prognostic model's downstream TCGA results reproduce 1:1: risk-score time-ROC AUC (0.75/0.78/0.76/0.74/0.72) and risk-group KM (p<1e-4, groups 145/146) and the TGF-score KM (p=0.0014) all match to rounding when recomputed from the repo's own per-sample tables with survival/timeROC. The 3-gene model (ITGA5/SHF/SNRPN) and its coefficients are taken from the shipped artifact (gene_coef.csv); they could NOT be independently re-derived because (a) the TCGA expression matrix CESC_TPM.txt is not shipped (only clinical+CNV) and (b) gene selection used a stochastic resampling loop with an un-recorded seed. The external validation on GSE44001 (genuinely downloaded + the published coefficients applied) is the weak point: high-risk does trend to worse DFS (KM p=0.043) but time-ROC AUC is only ~0.6 at 1-5y across fixed-coef/z-score/re-fit variants, below the paper's 'good AUC' claim, and ITGA5's weight collapses to ~0 when re-fit on GSE44001 -- flagged as a possible over-statement of external robustness. NOT attempted: de-novo LASSO/WGCNA (un-shipped TPM), scRNA 9-cell-type landscape (Seurat, hard 20%), and all wet-lab Fig 9 assays (out of scope). «host» holds only small result files + pointers; data/compute on «infra»/«our HPC».

💻 Code ↗ 🗄 Data: GSE44001

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 79
    assessed: 2026-06-14 ⛓ d9c6c45c1465
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

This study tests whether aberrant activation of TGF-β signaling in cancer-associated fibroblasts shapes the immunosuppressive tumor microenvironment and drives progression of cervical cancer, and whether TGF-β-related genes can yield a clinically useful prognostic model.

Core claims
  • TGF-β signaling activity is significantly increased in cervical cancer fibroblasts compared to normal fibroblasts and promotes their proliferation/differentiation. finding
  • Strong TGF-β-mediated communication between fibroblasts and macrophages and NK/T cells contributes to an immunosuppressive microenvironment. mechanism
  • A three-gene prognostic model (ITGA5, SHF, SNRPN) derived from TGF-β-related WGCNA modules predicts cervical cancer survival across multiple datasets. resource
  • ITGA5 and SNRPN are upregulated and SHF is downregulated in cervical cancer cells relative to normal cervical epithelial cells. finding
  • ITGA5 knockdown suppresses viability, migration, and invasion of cervical cancer cells. finding
  • Single-cell analysis of the cervical cancer TME identifies nine major cell types, with fibroblasts central to TGF-β activation. finding
  • WGCNA identifies gene modules significantly associated with the TGF-β signaling pathway used for downstream prognostic modeling. method
  • TGF-β signaling activity correlates with suppression of inflammatory pathways, immune escape, and metabolic regulation in fibroblasts. mechanism
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA-seq analysis (Seurat/Harmony/UMAP clustering) human cervical cancer TME (GSE208653: 2 normal + 3 HPV-infected CC samples) none cell-type identification and TGF-β signaling activity (AUCell score) 10x Genomics (Read10X); GSE208653
pseudo-time trajectory analysis fibroblasts from normal vs CC samples (GSE208653) none differentiation trajectory and gene expression dynamics Monocle
cell-cell communication analysis cervical cancer TME single-cell subpopulations none ligand-receptor TGF-β signaling interactions CellChat
WGCNA and Cox/LASSO prognostic modeling TCGA-CESC bulk RNA-seq (291 tumor samples) and GSE44001 (300 samples) none prognostic gene modules, risk score, survival (K-M, ROC) WGCNA, glmnet, timeROC
quantitative real-time PCR Hela CC cells and Ect1/E6E7 normal cervical epithelial cells none / si-ITGA5 knockdown mRNA levels of ITGA5, SHF, SNRPN (2^-ΔΔCT, GAPDH normalizer) SYBR Green qPCR (Beyotime)
CCK-8 cell viability assay Hela CC cells si-ITGA5 vs si-NC transfection cell viability (OD 450 nm) CCK-8 (Beyotime); Bio-Rad iMark reader
scratch/wound healing migration assay Hela CC cells si-ITGA5 vs si-NC transfection wound closure (%) / migration Olympus DP27 microscope
transwell invasion assay (Matrigel) Hela CC cells si-ITGA5 vs si-NC transfection number of invaded cells (crystal violet) Corning 8 µm transwell; Olympus DP27
Key results
  • TGF-β signaling AUCell score markedly higher in CC fibroblasts than normal fibroblasts
  • TGF-β-related genes ID2, PPP1R15A, SMAD7, XIAP significantly hyperexpressed in CC fibroblasts
  • CC fibroblasts located at the end of the pseudo-time differentiation trajectory while normal fibroblasts at the start
  • Three-gene model (ITGA5, SHF, SNRPN) shows good predictive ability across TCGA-CESC and GSE44001
  • ITGA5 and SNRPN higher and SHF lower expression in CC cells vs normal cervical epithelial cells
  • ITGA5 knockdown suppressed viability, migration and invasion of CC cells
  • Nine major cell types identified in the cervical cancer TME with fibroblast markers COL1A2, DCN, COL1A1 highly expressed
  • TGF-β signaling activity positively correlated with negative-regulation-of-immune inflammatory pathways
Key statistics
  • count 604,000 new cervical cancer cases worldwide in 2020 (WHO global CC incidence)
  • count 342,000 deaths (global CC mortality 2020)
  • count 291 tumor samples (TCGA-CESC samples used for analysis)
  • count 300 tumor samples (GSE44001 validation samples retained)
  • count 2 normal + 3 HPV-infected CC samples (GSE208653 single-cell dataset)
  • count nine cell types (major cell types identified in TME)
  • other >90% (share of global CC burden in less developed nations)
  • pvalue p < 0.05 (threshold defined as statistically significant)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This bioinformatics and translational study combined single-cell RNA-seq analysis (GSE208653; 5 samples), bulk RNA-seq prognostic modeling (TCGA-CESC; n=291 tumors split 70/30 training/validation), and an independent microarray cohort (GSE44001; n=300 tumors) with in vitro functional validation in HeLa cells. Pathway activity was scored via AUCell and gene co-expression modules were identified by WGCNA; a three-gene prognostic risk score was constructed by sequential univariate Cox, LASSO Cox, and multivariate Cox regression. Survival differences between risk groups were assessed by Kaplan–Meier analysis with log-rank test, and model discrimination was characterized by time-dependent ROC curves.

Replicationmixed Sample sizeTCGA-CESC n=291 tumors (70% training, 30% validation by random split); GSE44001 n=300 tumors; scRNA-seq GSE208653 n=5 samples (2 normal, 3 CC); in vitro HeLa and Ect1/E6E7 cell-line experiments with biological replicate n not explicitly stated; no power calculation mentioned GroupsNormal vs CC fibroblasts (scRNA-seq); high TGF-β vs low TGF-β signaling score subgroups (scRNA-seq); high-risk vs low-risk patients (prognostic model); si-ITGA5 vs si-NC HeLa cells (in vitro functional assays); HeLa vs Ect1/E6E7 (qPCR expression) Pairingunpaired Randomization/blindingnot stated Dispersionunclear Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Wilcoxon rank-sum test Two-group comparisons of continuous variables (e.g., TGF-β AUCell scores in normal vs CC fibroblasts) not stated per comparison not stated
Student's t-test Two-group comparisons of continuous variables (stated alongside Wilcoxon in statistical analysis section) not stated per comparison not stated
Pearson correlation Correlation of TGF-β AUCell scores with inflammatory, proliferative, and metabolic pathway AUCell scores in CC fibroblasts not stated not stated
Spearman correlation General correlations as stated in the statistical analysis section not stated not stated
Univariate Cox proportional hazards regression Screening prognostic relevance of WGCNA module genes in TCGA-CESC training set ~204 (70% of 291) not stated
LASSO Cox regression (glmnet) Penalized feature selection from univariate Cox candidate genes ~204 (70% of 291) not stated
Multivariate Cox regression Final risk score coefficient estimation for ITGA5, SHF, SNRPN ~204 (70% of 291) not stated
Kaplan–Meier analysis with log-rank test Overall survival comparison between high- and low-risk groups in TCGA-CESC training set, TCGA-CESC validation set, and GSE44001 291 (TCGA-CESC total), 300 (GSE44001) not stated
Time-dependent ROC curve (timeROC package) Discriminative performance of the prognostic model at multiple survival time horizons not stated per dataset na
ssGSEA (GSVA package) 28-immune-cell type infiltration scoring across TCGA-CESC samples stratified by risk group 291 na
AUCell scoring TGF-β signaling pathway activity per cell in scRNA-seq data; also applied to inflammatory, proliferative, and metabolic gene sets 5 samples (GSE208653: 2 normal, 3 CC) na
Approaches that could also have been used
  • Pathway activity in individual cells was quantified with AUCell scores, which use an area-under-the-recovery-curve statistic over ranked genes
    Could also: UCell or single-sample GSVA could also score pathway activity per cell — UCell uses a Mann–Whitney U-based rank statistic that is not sensitive to dataset size or expression-level normalization choices; GSVA is the established bulk-RNA counterpart extended to single-cell contexts — comparing two or more scoring methods can reveal whether conclusions are robust across approaches
  • Pearson correlation was used to relate TGF-β AUCell scores to other pathway scores in fibroblasts, while Spearman was stated as the general correlation method elsewhere in the same paper
    Could also: Spearman rank correlation could also be applied to AUCell score correlations for consistency — AUCell scores are bounded and may be non-normally distributed; Spearman is more robust to non-normality and outliers and is already the paper's stated default, so applying it uniformly would make the correlation analyses internally consistent
  • Multiple individual gene-level comparisons between normal and CC groups (and between high/low TGF-β subgroups) were made without a stated multiplicity correction
    Could also: Benjamini–Hochberg false discovery rate (FDR) correction could also be applied across the family of simultaneous gene-level tests — When many genes are tested simultaneously, FDR control at a chosen threshold (commonly 5%) is a standard approach in omics research that keeps the expected proportion of false positives interpretable without being as conservative as Bonferroni
  • Prognostic gene selection used a sequential pipeline of univariate Cox → LASSO Cox → multivariate Cox regression
    Could also: Elastic net Cox regression (combining L1 and L2 penalties) could also perform joint feature selection and coefficient shrinkage in a single step — Elastic net can handle correlated predictors more stably than pure LASSO, which may arbitrarily retain one gene among a group of highly correlated candidates; it is a frequently used alternative when collinearity among gene expression features is expected
  • Internal validation used a single 70/30 random split of the TCGA-CESC dataset
    Could also: Repeated k-fold cross-validation (e.g., 5- or 10-fold, repeated multiple times) or bootstrap-based optimism correction could also estimate internal model performance — A single random split can produce variable performance estimates depending on the specific partition drawn; cross-validation averages over multiple splits to reduce this variance and is a common complement to external cohort validation
  • Survival differences between risk groups were assessed with the standard (unweighted) log-rank test
    Could also: A Cox proportional hazards model comparison, or a weighted log-rank test (e.g., Fleming–Harrington), could also be used — The standard log-rank test is most powerful when hazard ratios are proportional and constant over time; a Cox model provides an effect size estimate (hazard ratio with confidence interval) in addition to a p-value, and weighted variants can better detect early or late divergence in survival curves — both are common in cancer survival studies
Software: R 3.6.0 · GraphPad Prism 8.0.2 · Seurat · harmony · AUCell · Monocle · CellChat · WGCNA · glmnet (LASSO Cox) · timeROC · GSVA (ssGSEA) · ESTIMATE · TIMER

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
10
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE208653 GEO in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet
GSE44001 GEO in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet

Downstream reach in the literature

95 downstream papers · 2 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40124621

Paper: TGF-β-mediated activation of fibroblasts in cervical cancer: implications for tumor microenvironment and prognosis. PeerJ 2024; DOI 10.7717/peerj.19072. PMID 40124621 · PMCID PMC11929507. Code: https://github.com/21kunzhang/raw-data (pushed 2024-12-20, public, no license). Data: GEO GSE208653 (scRNA, 5 samples), GSE44001 (microarray, 300 tumors, GPL14951, DFS endpoint), TCGA-CESC (TPM expression + clinical + CNV).

Datasets — what is / isn't shipped in the repo

  • scRNA GSE208653 RAW (10x: barcodes/features per sample) — SHIPPED under 00_origin_datas/GEO/GSE208653_RAW/ (5 samples: 3 CA_HPV, 2 NO_HPV).
  • TCGA-CESC clinical + gene-level CNV — SHIPPED (00_origin_datas/TCGA/).
  • TCGA-CESC expression CESC_TPM.txtNOT shipped (script reads it; absent). → any de-novo step needing TCGA expression (LASSO gene selection, WGCNA, risk score from expression) requires GDC/Xena download.
  • GSE44001 expression/clinicalNOT shipped; script downloads via a custom getGEOExpData() helper → reproduce via GEOquery.
  • Many intermediate result tables are SHIPPED (gene_coef.csv, tcga.risktype.cli.txt, TGF.score.txt, WGCNA_Modules.csv, aucs.csv, …) — these let us reproduce downstream steps without the missing raw expression.

Pipeline stages (numbered folders) and scope decision

stage analysis pipeline in scope?
01_landscape Seurat scRNA QC/cluster/UMAP, 9 cell types Seurat+harmony partial (C8, heavy — hard 20%)
02_AUCell AUCell TGF/escape scoring, DEG AUCell descriptive, no single pinnable number → not attempted
03_pseudotime monocle BEAM trajectory monocle descriptive → not attempted
04_immu_metab DEG + GO/KEGG + Pearson corr clusterProfiler descriptive → not attempted
05_CellChat ligand–receptor TGF communication CellChat descriptive → not attempted
06_TGF TGF ssGSEA score → KM(OS) survival IN SCOPE (C5) p=0.0014 shipped
07_WGCNA WGCNA modules vs TGF score WGCNA IN SCOPE (C7), heavier, needs TCGA TPM
09_model univ-Cox→LASSO-Cox→risk score, KM, ROC glmnet/survival/timeROC IN SCOPE (C1-C4, C6) core
10_immu ssGSEA immune vs risk group GSVA descriptive → not attempted
Raw experimental data EdU, CCK-8, transwell invasion, wound healing, qPCR (ITGA5 KD in Hela) wet-lab OUT OF SCOPE (manual/experimental)

What we attempt (80/20), in priority order

  1. C1/C2 — confirm the 3-gene model + coefficients (from shipped gene_coef.csv; note de-novo LASSO is stochastic resampling and needs un-shipped TPM → not exactly re-derivable, flagged).
  2. C3/C4 — reproduce TCGA risk-score time-ROC AUC (1–5y) and risk-group KM p-value from the shipped per-sample risk scores (tcga.risktype.cli.txt). Pure downstream recompute, no download → strongest internal check.
  3. C5 — reproduce TGF-score KM(OS) p from shipped TGF.score.txt + OS, median split.
  4. C6 — external validation on GSE44001: apply the published 3-gene coefficients, compute risk score, time-ROC (DFS) + KM. Genuine external reproduction (needs GEO DL).
  5. C7 (optional/heavier) — WGCNA on TCGA-CESC expression: confirm 17 merged modules at β=6 and brown↔TGF positive correlation. Needs TCGA TPM download.
  6. C8 (hard 20%) — scRNA 9 cell types: clustering-resolution sensitive; attempt only if budget remains.

Out of scope (not attempted), with reason

  • All Fig 9 cellular validation assays (EdU/CCK-8/transwell/wound/qPCR) — wet-lab, not pipeline-derived.
  • AUCell / pseudotime / CellChat / DEG-enrichment figures — descriptive, no single reported scalar to compare 1:1 (would be qualitative only).

Hard-rule compliance

All compute on «our HPC» SLURM; all downloads (GSE44001, optional TCGA TPM) land on «infra» reproductions/pmid-40124621/. «host» holds only small result files + pointers. See AUDIT.md for honesty flags.

Figures / tables: Fig 6Fig 7Fig 8Fig 5Fig 1
C1
Reported
LASSO model = ITGA5, SHF, SNRPN
Reproduced
same 3 genes (shipped artifact; selection not re-derived)
partial
C2
Reported
coef 0.539 / 0.403 / -0.311
Reproduced
0.5392 / 0.4028 / -0.3113
exact
C3
Reported
TCGA risk ROC AUC 0.75/0.78/0.76/0.74/0.72 (1-5y)
Reproduced
0.750/0.784/0.761/0.743/0.722
exact
C4
Reported
TCGA risk KM p<0.0001; High 145/Low 146
Reproduced
logrank p=2.5e-6; High 145/Low 146
exact
C5
Reported
TGF-score KM p=0.0014
Reproduced
logrank p=0.0014
exact
C6
Reported
GSE44001 external validation 'good AUC values'
Reproduced
KM p=0.043 (direction correct); time-ROC AUC ~0.53-0.65 across all variants; ITGA5 re-fit coef ~0
partial
C7
Reported
WGCNA 17 modules, beta=6, brown~TGF
Reproduced
17 modules confirmed in shipped table; de-novo not re-run (TPM unavailable)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 79/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

The headline computational claims reproduce essentially 1:1: recomputed from the repo's own shipped per-sample tables, TCGA risk-score time-ROC AUC (0.750/0.784/0.761/0.743/0.722), risk-group KM (p=2.5e-6, groups 145/146) and TGF-score KM (p=0.0014) all match the reported values to rounding, with no fabrication signal on the core. The genuine deviation is on the authors' side in external validation: GSE44001 reproduces only AUC ~0.53-0.65 (across three methods) against the paper's 'good AUC' claim, and ITGA5's weight collapses to ~0 — the direction/significance still hold (KM p=0.043) so it is a moderate over-statement, not a flipped conclusion. Two availability caveats (unshipped CESC_TPM.txt + un-recorded resampling seed) mean the upstream model selection is not byte-reproducible, but that is a data-availability limit, not a defect. Overall: solid reproduction with one explainable, documented over-statement of external robustness → yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

171.7 k
tokens (I/O) · 13.1 M incl. cache
18 min
runtime · 0.02 CPU-h
2.8 GB
peak RAM
3
HPC jobs
hummel
machine