Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Unveiling the immunometabolic landscape of colorectal cancer through PANoptosis-related gene expression.

Front Immunol · 2026
L1 41/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
41/100
Reproducibility score
1.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 4% of all assessed papers rank 1124 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the HEADLINE DEG result 1:1-class. The paper's 'code' (IOBR) is a generic third-party toolkit, not authors' scripts, so per P16 we ran the named methods (limma + sva::ComBat on the 3 GPL570 GEO discovery sets GSE41328/GSE22598/GSE23878; limma-voom on TCGA COAD+READ via recount3) on the paper's named data with the stated thresholds (|logFC|>1 & p<0.05). C1: reported 908 common DEGs vs reproduced 926 (+2.0%) despite fully independent tooling (recount3 read counts vs authors' GDC; symbol-alias matching) -> within-tol; 13/15 canonical CRC markers tested are in the intersection, confirming biological correctness. NOT attempted (the hard ~20%): consensus clustering C1=418/C2=128, LASSO 11-gene CPAN-index + survival AUC 0.62/0.60/0.62, CIBERSORT/ESTIMATE/TIDE immune scores, GSEA, and all single-cell analyses (multi-dataset, many free parameters). FLAG for human review: the stated PANoptosis-list composition 87.4/9.7/2.9% equals exactly 90/10/3 of a 103-gene list and is NOT derivable from the cited ref-23 set of 277 non-redundant genes (259/27/15) -> possible mis-citation; C2's '15-gene overlap' rests on this ambiguous list and was not reproduced rather than fabricated.

💻 Code ↗ 🗄 Data: GSE41328

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 41
    assessed: 2026-06-14 ⛓ 20c784954228
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Do PANoptosis-related genes drive colorectal cancer (CRC) progression by reshaping the immune microenvironment and metabolic reprogramming, and can their differential expression be used to build a prognostic risk-stratification tool (CPAN-index) for CRC patients?

Core claims
  • A CPAN-index prognostic model built from 11 PANoptosis-related differentially expressed genes (CPAN_DEGs) stratifies CRC patients into high-risk and low-risk groups with distinct survival and immunophenotypes. resource
  • The high-risk CRC group exhibits an 'invasion-metabolism-immunosuppressive' phenotype with immune tolerance and non-classical immune escape. finding
  • CDKN2A is upregulated in CRC cell lines and its knockdown inhibits proliferation and promotes apoptosis in vitro. finding
  • CDKN2A-mediated PANoptosis signaling drives CRC progression by reshaping the immune microenvironment and metabolic reprogramming. mechanism
  • PANoptosis-related DEGs between CRC and normal tissue are significantly enriched in metabolism, apoptosis, and immune regulation pathways. finding
  • scRNA-Seq shows a higher proportion of CPAN-index-positive immune cells and a lower proportion of CPAN-index-positive tumor cells, indicating a role in the tumor immune microenvironment. finding
  • Apoptosis-related genes constitute the majority (87.4%) of the PANoptosis gene list. finding
  • Unsupervised consensus clustering of CPAN_DEGs divides CRC into two clusters (C1, C2) with distinct immune scores and immune checkpoint expression. method
Experimental setups
Assay System Perturbation Readout Platform
Bulk RNA-seq differential expression and prognostic modeling (limma, LASSO, Cox) CRC tissue (TCGA COAD/READ + GEO cohorts: GSE41328, GSE22598, GSE23878, GSE39582, GSE72970, GSE17536) none differentially expressed PANoptosis genes, CPAN-index risk score, survival/ROC
Single-cell RNA-seq (scRNA-Seq) cell-type analysis and module scoring CRC samples (TISCH datasets GSE146771, EMTAB8107, GSE166555) none PANoptosis gene expression per cell type, proportion of CPAN-index-positive cells TISCH database; Seurat/Harmony/SingleR/UMAP
Immune microenvironment deconvolution merged CRC bulk dataset none immune/stromal/ESTIMATE scores, immune cell fractions, TIDE/dysfunction/CD274/IFNG/Merck18/CD8 scores CIBERSORT, ESTIMATE, TIDE
RT-qPCR CRC cell lines (LoVo, HT-29, HCT116, SW480) and normal colon epithelium HIEC-6 CDKN2A siRNA knockdown vs negative control CDKN2A mRNA expression PrimeScript RT kit, SYBR GreenER Supermix, Roche480 PCR system
CCK-8 proliferation assay HCT116 and SW480 CRC cells CDKN2A siRNA knockdown cell proliferation (absorbance at 450nm) CCK8 (KeyGEN), microplate reader
Flow cytometry apoptosis assay HCT116 and SW480 CRC cells CDKN2A siRNA knockdown apoptosis (Annexin/FITC and PI staining) FITC/PI (Biosharp)
Transwell migration assay HCT116 and SW480 CRC cells CDKN2A siRNA knockdown migrated/invaded cells (crystal violet)
Wound healing (scratch) assay HCT116 and SW480 CRC cells (5×10^5 cells/well) CDKN2A siRNA knockdown wound closure at 0h and 48h
Key results
  • Apoptosis-related genes comprised 87.4% of the PANoptosis gene list (pyroptosis 9.7%, necroptosis 2.9%) 87.4%
  • Of 908 DEGs common to CRC and GSE datasets, 15 genes overlapped with the PANoptosis gene list (CPAN_DEGs) 15 of 908
  • CDKN2A was upregulated in CRC cell lines; knockdown inhibited proliferation and promoted apoptosis in vitro
  • CPAN-index built from 11 CPAN_DEGs distinguished high-risk vs low-risk CRC patients with the high-risk group showing immunosuppressive phenotype 11 genes
  • Immune score and microenvironment (ESTIMATE) score were significantly higher in C2 cluster than C1
  • Immune checkpoints (TIGIT, PDCD1LG2, PDCD1, LAG3, HAVCR2, CTLA4, CD274) significantly higher in C2 than C1, indicating immunosuppressive state
  • scRNA-Seq showed higher proportion of CPAN-index-positive immune cells and lower proportion of CPAN-index-positive tumor cells
  • Consensus clustering identified k=2 as optimal, dividing the CRC cohort into C1 (n=418) and C2 (n=128) with no significant survival difference p=0.28
Key statistics
  • count 87.4% apoptosis, 9.7% pyroptosis, 2.9% necroptosis (composition of PANoptosis gene list)
  • count 908 common DEGs; 15 in PANoptosis list (Venn overlap of CRC and GSE DEGs)
  • count 11 CPAN_DEGs (genes in the CPAN-index prognostic model)
  • count C1 n=418, C2 n=128 (consensus cluster sizes of CRC cohort)
  • pvalue p=0.28 (p>0.05) (Kaplan-Meier survival difference between C1 and C2 clusters)
  • pvalue p<0.05 (immune score and microenvironment score higher in C2 vs C1)
  • other |logFC|>1 & p.Val<0.05 (DEG selection threshold (limma))
  • count 4000 cells per well; 40,000 per well (transwell); 5×10^5 per well (wound healing) (cell seeding densities for functional assays)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study integrates public bulk RNA-seq data (TCGA and GEO cohorts) with in vitro cell-line experiments and public scRNA-seq data to characterize PANoptosis-related gene expression in colorectal cancer. Differentially expressed genes were identified with the limma R package, consensus clustering was used to stratify patients, and a LASSO-penalized Cox proportional hazards model generated a prognostic index (CPAN-index) that was externally validated in three independent GEO cohorts via ROC/AUC analysis. In vitro functional experiments (CCK-8, flow cytometry, wound healing, transwell) used at least three biological replicates per group, with comparisons by two-sided t-test and results expressed as mean ± SD.

Replicationmixed Sample sizeIn vitro experiments stated to use ≥3 biological replicates; consensus cluster sizes reported (C1 n=418, C2 n=128); validation cohort sample sizes not stated in the text; no formal power calculation described GroupsCRC vs normal colon tissue; C1 vs C2 consensus clusters; CPAN-index high-risk vs low-risk; CDKN2A siRNA knockdown vs negative control in HCT116 and SW480 cell lines Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionAdjusted p-values (p.adj) explicitly reported for GO/KEGG enrichment; limma's default BH FDR adjustment implied by the pipeline but not named; no correction stated for the multiple t-tests across immune scores, checkpoint molecules, and functional assay endpoints
Statistical tests used
Test Applied to n Assumptions
limma moderated t-test (via 'limma' R package) Differential expression between CRC and normal tissues in TCGA and combined GEO datasets (|logFC|>1, p.Val<0.05); also between C1 and C2 consensus clusters C1 n=418, C2 n=128 for cluster DEGs; TCGA+GEO merged dataset size not stated not stated
ConsensusClusterPlus unsupervised clustering evaluated by PAC and CDF curves Determination of optimal cluster number (k=2 selected) and patient stratification into C1 and C2 subtypes C1 n=418, C2 n=128 not stated
Kaplan-Meier survival analysis (log-rank test implied but not named) Overall survival comparison of C1 vs C2 clusters (p>0.05); high-risk vs low-risk CPAN-index groups in training and validation datasets C1 n=418, C2 n=128; validation cohort sizes not stated not stated
Univariate Cox proportional hazards regression Screening CPAN_DEGs for prognostic value prior to LASSO selection null not stated
LASSO (L1-penalized) Cox regression Dimensionality reduction to select candidate genes for the CPAN-index; optimal lambda chosen by minimizing binomial deviance null not stated
Multivariate Cox proportional hazards regression (maximum partial likelihood estimation) Construction of CPAN-index from 15 LASSO-selected candidate genes null not stated
ROC curve analysis (time-dependent AUC) Validation of CPAN-index predictive performance in GSE39582, GSE17536, and GSE72970 null na
GO/KEGG overrepresentation enrichment analysis (Fisher's exact or hypergeometric, adjusted p-values reported as p.adj) Functional annotation of 15 CPAN_DEGs: GO-BP, GO-MF, KEGG pathways (Figure 3) 15 CPAN_DEGs not stated
GSEA (Gene Set Enrichment Analysis, via 'GseaVis' R package) Pathway enrichment comparison between high-risk and low-risk CPAN-index groups null not stated
CIBERSORT algorithm (immune deconvolution) Immune cell proportion estimation in C1 vs C2 clusters and high-risk vs low-risk groups null na
ESTIMATE algorithm (immune/stromal scoring) Immune score, matrix score, and microenvironment score in C1 vs C2 clusters (Figure 5) null na
Two-sided t-test All between-group comparisons stated in Section 2.13; covers in vitro assays (CCK-8, flow cytometry, wound healing, transwell) and bulk-data group comparisons (immune/checkpoint scores between risk groups) ≥3 biological replicates per group (stated) not stated
Seurat FindMarkers (underlying statistical test not specified; Wilcoxon rank-sum by default) Identification of genes with significant expression variation across cell types in integrated scRNA-seq data null not stated
Seurat AddModuleScore Quantification of CPAN-index gene module activity per cell in scRNA-seq data; proportion of CPAN-index-positive cells per subpopulation calculated null na
Approaches that could also have been used
  • Group comparisons in cell-line experiments used two-sided t-tests with n≥3 biological replicates per group
    Could also: A non-parametric Wilcoxon rank-sum (Mann-Whitney U) test could also be used — With only three replicates per group, normality cannot be formally assessed; non-parametric alternatives make fewer distributional assumptions and are commonly applied in small-sample experimental biology
  • Feature selection for the prognostic model used LASSO-penalized Cox regression (L1 penalty only)
    Could also: Elastic net Cox regression (combining L1 and L2 penalties, alpha between 0 and 1) could also be applied — When predictors are correlated — as co-expressed genes frequently are — elastic net can improve coefficient stability by grouping correlated features rather than arbitrarily selecting one; it also tends to produce more reproducible gene sets across validation cohorts
  • Dispersion for in vitro assay data is reported as mean ± SD
    Could also: Individual data points overlaid on bar or box plots, or 95% confidence intervals, could also convey spread — With n=3 replicates, showing individual points alongside the mean makes the raw data transparent; 95% CIs communicate both variability and estimation precision, and are increasingly recommended by journals for small-n experimental data
  • Immune cell deconvolution relied on the CIBERSORT algorithm alone
    Could also: A second deconvolution algorithm such as TIMER2, xCell, or EPIC could also be applied in parallel — Different algorithms use different reference gene signatures and statistical assumptions; cross-validating immune composition estimates across two algorithms is a common approach for assessing robustness of findings
  • Many pairwise t-tests were conducted across immune scores, checkpoint molecules, and functional assay endpoints without a stated multiple-comparison correction
    Could also: A Benjamini-Hochberg FDR or Holm correction could also be applied across the family of simultaneous comparisons — When many endpoints are evaluated concurrently — as in the immune checkpoint panel — a family-wise correction controls the expected proportion of false discoveries across all tests in that family
  • The log-rank test underlying the Kaplan-Meier curve comparisons is implied by the reported p-values but is not named in the text
    Could also: Explicitly naming the log-rank test, or supplementing it with a Cox log-rank score test or restricted mean survival time (RMST) analysis, could also be reported — Naming the test aids reproducibility; RMST provides an effect-size estimate (difference in mean survival time up to a horizon) that complements the p-value and does not assume proportional hazards
Software: R 4.4.1 · R/limma · R/ConsensusClusterPlus · R/sva (ComBat batch correction) · R/IOBR (remove_batcheffect) · R/GseaVis · R/Seurat (FindMarkers, AddModuleScore, UMAP) · R/harmony (scRNA-seq batch correction) · R/SingleR (automated cell annotation) · CIBERSORT (web/algorithm) · ESTIMATE (algorithm) · TIDE (online scoring tool)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41601652

Paper: He X, Wang W, Li L, Yin Y, Ding S. Unveiling the immunometabolic landscape of colorectal cancer through PANoptosis-related gene expression. Front Immunol 2025/2026. DOI 10.3389/fimmu.2025.1615022. PMCID PMC12832467.

Nature of the "code" link. The brief lists the code as https://github.com/IOBR/IOBR — IOBR is a generic third-party immuno-oncology R toolkit (TME deconvolution, signature scoring, batch removal), not the authors' own analysis pipeline. Per brief rule P16, applying the described third-party tools to the paper's own data is an equally valid reproduction. There is no authors' repo with parameter-pinned scripts; we reproduce by running the named tools (limma, sva/ComBat, recount3-sourced TCGA) on the paper's named data with the paper's stated thresholds.

Datasets named in the paper

  • Bulk discovery (GEO, all GPL570 / Affy HG-U133 Plus 2.0): GSE41328 (n=20), GSE22598 (n=38), GSE23878 (n=59) — CRC tumour vs normal.
  • TCGA COAD + READ bulk RNA-seq (tumour vs normal).
  • Validation cohorts: GSE39582, GSE72970, GSE17536 (survival).
  • scRNA-seq: GSE146771, GSE166555, E-MTAB-8107.
  • PANoptosis gene list: from ref 23 = Song et al. 2023, Front Immunol, PMC10311484 (HCC HPAN-index), reported there as 277 non-redundant genes (apoptosis 259 / pyroptosis 27 / necroptosis 15 before de-duplication).

In scope (pipeline-derived, attempted)

# Reported result Location Pipeline Status
C1 908 DEGs common to TCGA and GEO (|logFC|>1, p<0.05, limma) Results/DEG limma on GEO 3-set (ComBat-merged) ∩ limma-voom on TCGA COAD+READ attempted (primary)
C2 15 of the 908 DEGs are PANoptosis genes Results/DEG C1 ∩ Song-2023 277-gene list secondary (list provenance risk)
C3 PANoptosis list composition 87.4% apop / 9.7% pyro / 2.9% necro Results/Fig count categories of the cited list flagged — does not match the cited 277-gene set; 87.4/9.7/2.9 = 90/10/3 of 103 genes, implying a different list than ref 23 (see AUDIT)

Out of scope (not attempted — why)

  • Consensus clustering C1=418 / C2=128 (k=2): depends on exact TCGA sample filtering + PANoptosis-DEG feature set; high specification ambiguity.
  • LASSO 11-gene CPAN-index + survival AUC 0.62/0.60/0.62: needs TCGA survival + exact LASSO seed/lambda; not pinnable (the hard last 20%).
  • CIBERSORT / ESTIMATE / TIDE immune scores, GSEA, single-cell (Seurat/Harmony/ SingleR) AddModuleScore proportions (NK 70.3% etc.): multi-tool, many free parameters, several datasets — out of scope for a clear 1:1.
  • CDKN2A knockdown (apoptosis/proliferation/migration p<0.0001): wet-lab, not computational — out of scope.

Strategy (80/20)

Reproduce the headline DEG intersection (C1) faithfully with the named tools and thresholds; report the GEO-side and TCGA-side DEG counts separately too (the paper prints only the intersection). Treat C2 as best-effort and C3 as an auditable inconsistency rather than a clean target. Do not chase the clustering/LASSO/immune/scRNA tail.

Figures / tables: figure
C1
Reported
908 DEGs common to TCGA and GEO (limma, |logFC|>1 & p<0.05)
Reproduced
926
within tolerance
C1a
Reported
GEO discovery DEGs (not printed separately)
Reproduced
998
partial
C1b
Reported
TCGA COAD+READ DEGs (not printed separately)
Reproduced
8834
partial
C2
Reported
15 of the common DEGs are PANoptosis-list genes
Reproduced
not reproduced (list provenance ambiguous)
did not match
C3
Reported
PANoptosis list composition 87.4% / 9.7% / 2.9%
Reproduced
inconsistent with cited 277-gene source (= 90/10/3 of 103)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 41/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

The headline DEG result reproduced well — 926 common DEGs vs the reported 908 (+2.0%) despite fully independent tooling (recount3/limma-voom vs the authors' GDC/IOBR pipeline), with 13/15 canonical CRC markers present, so the DEG backbone is credible. The genuine concern is on the authors' side: the reported PANoptosis composition 87.4/9.7/2.9% equals exactly 90/10/3 of a 103-gene list and is not derivable from the cited 277-gene ref-23 set (86.0/9.0/5.0%), making the 15-gene overlap (C2) unverifiable — a possible mis-citation/too-perfect flag rather than confirmed fabrication. Severity is moderate: the central DEG finding holds while the PANoptosis-specific framing and all downstream analyses (clustering, LASSO, immune, single-cell) remain unverified, so this is a yellow case flagged for human review.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

212.5 k
tokens (I/O) · 11.6 M incl. cache
19 min
runtime · 0.03 CPU-h
7 GB
peak RAM
1
HPC jobs
hummel
machine