Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Assessment tool based on fatty acid metabolic signatures for predicting the prognosis and treatment response in bladder cancer.

Heliyon · 2023
L1 71/100 PQI 90
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
71/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 38% of all assessed papers rank 694 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for a FAITHFUL 1:1 reproduction of the central validation result. PRIMARY target = GSE32894 external-validation panel (Fig 3 G/H), reproduced on the authors' own shipped data (GEO/GSE32894_exp.rds + GSE32894_cli.RData) via «our HPC» SLURM «job». timeROC AUC reproduced to within rounding: 1y 0.797 vs 0.80, 3y 0.837 vs 0.84, 5y 0.857 vs 0.86; KM low-vs-high RiskScore p=2.0e-6 with low-risk better (matches text); 5-gene set (SRC,CTLA4,CTSE,LAMA2,ADAMTSL4) confirmed. NOT attempted (the hard 20%): TCGA training cohort (3 FAM subtypes via ConsensusClusterPlus, 13-gene LASSO @ lambda=0.0309, the TCGA-trained 5-gene coefficients, and TCGA AUC 0.71/0.68/0.65) -- these need the TCGA expression matrix BLCA_TPM.txt which the repo READS but does NOT ship, plus un-seeded stochastic steps (cv.glmnet, ConsensusClusterPlus resampling) the authors did not pin, so the exact panel is non-deterministic. HONESTY FLAGS for human audit: (1) paper claims 316 GSE32894 samples but the authors' shipped data has 216 -- not derivable from deposited artifacts; (2) validation AUC (0.80-0.86) exceeds training AUC (0.65-0.71) because the validation RiskScore Cox is re-fit on GSE32894 itself rather than applying TCGA coefficients -> optimistic in-sample bias (reproducible, but not true external validation); (3) repo not runnable as-shipped (undefined helpers ggplotTimeROC/writeMatrix, mismatched load() filenames) -- numeric results obtained by re-implementing the well-specified steps.

💻 Code ↗ 🗄 Data: GSE32894

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 71
    assessed: 2026-06-14 ⛓ 3137ac802552
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Fatty acid metabolism (FAM)-related gene expression patterns can be used to define molecular subtypes and build a prognostic risk model that predicts prognosis and treatment (immunotherapy/chemotherapy) response in bladder cancer (BLCA) patients.

Core claims
  • Consensus clustering of prognosis-related fatty acid metabolism genes (FAMGs) identifies three molecular subtypes of BLCA (FAMC1, FAMC2, FAMC3) with distinct prognoses and tumor microenvironments finding
  • FAMC1 subtype shows the most favorable prognosis, highest FAM score, and a less immunosuppressive tumor microenvironment finding
  • FAMC3 subtype shows the greatest immunosuppression (highest immune checkpoint expression, TIDE/Exclusion/Dysfunction scores) and worse predicted immunotherapy benefit finding
  • A five-gene RiskScore (SRC, CTLA4, CTSE, LAMA2, ADAMTSL4), derived via LASSO and multivariate Cox regression on FAMGs from a PPI network module, predicts BLCA prognosis method
  • Patients with low RiskScore have significantly better survival, more favorable immune microenvironment, and better predicted immunotherapy response than high RiskScore patients finding
  • The RiskScore's prognostic performance is validated in an independent GEO cohort (GSE32894) finding
  • RiskScore correlates with predicted chemotherapy drug sensitivity (IC50) for eight chemotherapeutic agents finding
  • A Nomogram integrating RiskScore and clinical variables provides strong prognostic value for BLCA, confirmed by calibration, decision curve, and ROC analyses method
Experimental setups
Assay System Perturbation Readout Platform
Bulk RNA-seq / high-throughput expression data (bioinformatic analysis) TCGA-BLCA cohort (403 tumor samples) none FAMG expression, univariate Cox prognostic association TCGA GDC portal
Consensus clustering (pam algorithm, spearman distance) TCGA-BLCA tumor samples none molecular subtype assignment (FAMC1-3) SangerBox / consensus clustering software
Gene set enrichment analysis (GSEA/GSVA, HALLMARK gene sets) FAMC1-3 subtype groups, TCGA-BLCA none differentially activated biological pathways MSigDB h.all.v7.5.1 gene sets
Immune deconvolution (CIBERSORT, ESTIMATE, MCP-Count, ssGSEA) TCGA-BLCA tumor samples none immune/stromal cell infiltration scores, immune checkpoint gene expression CIBERSORT web tool; ESTIMATE; MCP-Count; GSVA package
Protein-protein interaction network construction and MCODE module analysis Differentially expressed FAMGs across FAMC1-3 none PPI network modules, GO/KEGG enrichment STRING database v11.5; Cytoscape 3.9.1
LASSO and multivariate Cox regression (RiskScore construction) TCGA-BLCA prognostic FAMGs (Module 1) none gene coefficients, RiskScore formula R (pRRophetic-adjacent Cox modeling)
External validation of RiskScore (survival analysis, ROC) GSE32894 cohort (316 samples) none AUC for 1/3/5-year survival prediction GEO dataset GSE32894
Chemotherapy drug sensitivity prediction (ridge regression) and immunotherapy response scoring (TIDE, T-cell-inflamed GEP) TCGA-BLCA RiskScore subgroups none predicted IC50 for 8 agents; TIDE score; GEP score pRRophetic R package; TIDE web tool
Key results
  • Optimal clustering at k=3 defined three FAM molecular subtypes (FAMC1-3) based on consensus matrix/CDF
  • FAMC1 patients had the most favorable prognosis and the highest FAM score
  • Immune checkpoint genes CD80, PDCD1LG2, HAVCR2, CD28, CTLA4, and PDCD1 showed sequential increase from FAMC1 to FAMC3
  • FAMC3 had significantly higher Exclusion, TIDE, and Dysfunction scores than FAMC1/FAMC2
  • LASSO Cox model selected 13 FAMGs at optimal lambda; multivariate Cox narrowed this to 5 genes (SRC, CTLA4, CTSE, LAMA2, ADAMTSL4) forming the RiskScore lambda = 0.0309
  • RiskScore predicted 1-, 3-, 5-year survival in TCGA-BLCA cohort AUC 0.71 / 0.68 / 0.65
  • RiskScore predicted 1-, 3-, 5-year survival in GSE32894 validation cohort AUC 0.8 / 0.84 / 0.86
  • Low RiskScore patients had significantly higher T-cell-inflamed GEP score, indicating better predicted immunotherapy response
Key statistics
  • count 403 tumor samples (TCGA-BLCA training cohort size)
  • count 316 samples (GSE32894 validation cohort size)
  • count 30 prognosis-related FAMGs (univariate Cox screening, FDR<0.05 & |log2FC|>1)
  • count 55 FAMGs with prognostic impact (p<0.05) in Module 1 (univariate Cox on PPI module genes)
  • other lambda = 0.0309, 13 FAMGs retained (10-fold cross-validated LASSO Cox model)
  • other RiskScore = -0.171*SRC - 0.303*CTLA4 - 0.092*CTSE + 0.29*LAMA2 + 0.12*ADAMTSL4 (final multivariate Cox RiskScore formula)
  • other AUC 0.71 (1yr), 0.68 (3yr), 0.65 (5yr) (ROC for RiskScore in TCGA-BLCA cohort)
  • other AUC 0.8 (1yr), 0.84 (3yr), 0.86 (5yr) (ROC for RiskScore in GSE32894 validation cohort)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This bioinformatics study used TCGA-BLCA (n=403) as a training cohort and GSE32894 (n=316) as a validation cohort to characterize fatty acid metabolism in bladder cancer. Molecular subtypes were derived via consensus clustering of prognostic fatty acid metabolism genes, and a five-gene RiskScore was constructed through a sequential univariate Cox → LASSO Cox → multivariate Cox pipeline. Survival differences were visualized with Kaplan-Meier curves and model discrimination was summarized by time-dependent ROC/AUC at 1, 3, and 5 years; immune infiltration and treatment-response differences across groups were characterized using CIBERSORT, ESTIMATE, MCP-Count, ssGSEA, and TIDE scores.

Replicationunclear Sample sizeTraining cohort: 403 TCGA-BLCA tumor samples retained after filtering for follow-up and gene symbols; validation cohort: 316 GSE32894 samples; no formal power calculation described GroupsThree fatty acid molecular subtypes (FAMC1/2/3); high vs. low RiskScore (median split); clinicopathological subgroups (age, sex, stage, grade, TNM) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionFDR (specific procedure not named; BH is the default in limma and GSEA but is not stated explicitly)
Statistical tests used
Test Applied to n Assumptions
Univariate Cox proportional-hazards regression Initial screening of prognostic FAMGs from TCGA-BLCA (p<0.05; results section also applies FDR<0.05 & |log2FC|>1); secondary screening of 55 DEGs from PPI Module 1 403 not stated
LASSO Cox regression with 10-fold cross-validation Feature selection reducing 55 candidate FAMGs to 13 (optimal lambda=0.0309) 403 not stated
Multivariate Cox proportional-hazards regression Final selection of 5-gene prognostic model (SRC, CTLA4, CTSE, LAMA2, ADAMTSL4) and Nomogram construction 403 not stated
Kaplan-Meier survival analysis (log-rank test implied, not explicitly named) Prognosis comparison across FAMC1/2/3 subtypes and high vs. low RiskScore groups in both TCGA-BLCA and GSE32894 cohorts; subgroup analyses by age, sex, stage, grade, TNM 403 (training); 316 (validation) not stated
limma linear model with empirical Bayes moderation (differential expression) Identification of DEGs across FAMC1-3 subtypes (FDR<0.05, |log2FC|>1) 403 not stated
Gene Set Enrichment Analysis (GSEA) and single-sample GSEA (ssGSEA via GSVA package) Biological pathway enrichment in molecular subtypes and RiskScore groups (FDR<0.05); immune cell and TME signature scoring; FAM pathway activity score; T-cell-inflamed GEP score 403 not stated
Time-dependent ROC / AUC Predictive discrimination of RiskScore at 1, 3, and 5 years in TCGA-BLCA (AUC 0.71/0.68/0.65) and GSE32894 (AUC 0.80/0.84/0.86); Nomogram calibration and decision curve also reported 403 (training); 316 (validation) na
Ridge regression (pRRophetic R package) Estimation of IC50 values for eight chemotherapeutic agents in high vs. low RiskScore groups 403 not stated
Approaches that could also have been used
  • Model discrimination was summarized by time-dependent AUC at three fixed time points (1, 3, 5 years)
    Could also: Harrell's concordance index (C-index) or Uno's C-statistic could also quantify discrimination for a Cox-based survival model — The C-index integrates discrimination across the entire observed follow-up rather than at selected time points and is the canonical performance metric for Cox models, making results more directly comparable with published prognostic signatures
  • Patients were dichotomized into high/low RiskScore groups at the median
    Could also: The continuous RiskScore, or tertile/quartile groupings, could also be used in survival analyses — Median dichotomization discards within-group prognostic variation; retaining the score as continuous (or using finer groupings) preserves that information and avoids the sensitivity of results to the chosen cut-point
  • LASSO Cox was used for variable selection among 55 candidate genes
    Could also: Elastic-net Cox regression (mixing L1 and L2 penalties) could also be applied for variable selection — Pure LASSO selects arbitrarily among correlated genes; elastic net tends to produce more stable gene signatures when predictors are correlated, which is common in co-expressed transcriptomic data
  • Consensus clustering used the PAM algorithm with Spearman correlation distance over 500 bootstraps
    Could also: Non-negative Matrix Factorization (NMF) or hierarchical clustering with Ward linkage could also be used to derive molecular subtypes — Different algorithms can produce different subtype solutions; NMF is widely used in transcriptomic cancer subtyping and yields interpretable metagene patterns, while Ward hierarchical clustering is another established approach — reporting sensitivity across methods would characterize subtype robustness
  • Four parallel immune deconvolution methods (CIBERSORT, ESTIMATE, MCP-Count, ssGSEA) were applied concurrently to characterize TME
    Could also: xCell or TIMER2.0 could also estimate immune cell composition, or a single method could be pre-selected based on a published benchmark for bulk RNA-seq — Different deconvolution tools use different reference matrices and assumptions; applying multiple tools without a primary pre-specified method can complicate interpretation — selecting one tool based on benchmarking, or reporting concordance across tools, makes the results more interpretable
  • The initial univariate Cox screening of FAMGs used an uncorrected p<0.05 threshold across multiple genes before LASSO refinement
    Could also: An FDR-adjusted threshold (e.g., BH-corrected p<0.05 or p<0.10) could also be applied at this first screening step — Applying FDR correction at the univariate screening stage reduces entry of false-positive associations into the LASSO step, potentially improving the stability and reproducibility of the final gene signature
Software: SangerBox (online data processing platform) · R/limma · R/clusterProfiler · R/GSVA (ssGSEA) · R/pRRophetic · Cytoscape (PPI network visualization) 3.9.1 · STRING (PPI database) 11.5 · CIBERSORT (web-based immune deconvolution) · TIDE (web-based immunotherapy response tool)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
4
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

Downstream reach in the literature

100 downstream papers · 1 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 38076064 (BLCA Fatty-acid metabolism RiskScore)

Paper: Assessment tool based on fatty acid metabolic signatures for predicting the prognosis and treatment response in bladder cancer. Heliyon 2023;e22768. Repo: https://github.com/xchen1212/BLCA_Fatty_acid (commit on main, pushed 2023-08-15). Data: TCGA-BLCA (training) + GEO GSE32894 (validation).

Pipeline-derived results (candidate for reproduction)

# Result Pipeline Data shipped in repo? In scope?
R1 3 FAM molecular subtypes (FAMC1/2/3) ssGSEA(HALLMARK FAM) → ConsensusClusterPlus (pam) k=3 on TCGA TCGA expr NOT shipped (only clinical + CNV) → needs GDC; consensus has stochastic resampling 20% / partial
R2 13-gene LASSO panel @ lambda=0.0309 univariate Cox → cv.glmnet LASSO on TCGA needs TCGA expr + upstream DEG/PPI; cv.glmnet random, no seed in code 20% (not byte-reproducible)
R3 5-gene RiskScore = −0.171·SRC −0.303·CTLA4 −0.092·CTSE +0.29·LAMA2 +0.12·ADAMTSL4 stepwise Cox on the 13 needs TCGA expr; depends on R2 20%
R4 GSE32894 validation: timeROC AUC@1/3/5y = 0.80 / 0.84 / 0.86 (Fig 3H) coxph(5 genes) on GSE32894 → riskscore z → timeROC YES — GSE32894_exp.rds + GSE32894_cli.RData both shipped PRIMARY (deterministic)
R5 GSE32894 KM high vs low RiskScore, logrank p<0.05 (Fig 3G) median-split risk, survfit logrank YES (shipped) PRIMARY (deterministic)
R6 TCGA timeROC AUC@1/3/5y = 0.71 / 0.68 / 0.65 (Fig 3E) coxph(5 genes) on TCGA → timeROC needs GDC TCGA TPM; 5-gene coefs given SECONDARY (if time)

Decision (80/20)

PRIMARY = R4 + R5 — the GSE32894 validation cohort. Both inputs are shipped in the repo, the model genes are explicitly stated, and the computation (coxph linear predictor → median split → timeROC/logrank) is fully deterministic. This is a central reported result (Fig 3 G/H/I) and self-contained.

SECONDARY = R6 (TCGA training AUC) if the GDC download + 5-gene scoring is quick.

NOT attempted (the hard, stochastic, or non-shipped 20%): R1–R3 — they require the un-shipped TCGA expression matrix (BLCA_TPM.txt is read by the script but not in the repo) AND involve un-seeded stochastic steps (ConsensusClusterPlus resampling, cv.glmnet fold assignment) that the authors did not pin, so the exact 13-gene panel / k=3 partition is not byte-reproducible by construction.

Reproducibility-surface notes (honesty)

  • The repo's scripts/scripts.R calls several undefined helper functions (ggplotTimeROC, writeMatrix) not present in base.R and not from any declared package → the script is not runnable as-is. We re-implement the well-specified parts (timeROC via the timeROC CRAN pkg; coxph linear predictor per get_riskscore).
  • The shipped GEO file names differ from those the script load()s (GSE32894_exp.rds shipped vs GSE32894_exp.RData loaded; GSE32894_cli.RData shipped vs GSE32894.RData loaded) → minor adaptation.
  • Methodological flag: GSE32894 RiskScore Cox is re-fit on the validation data itself (not applying the TCGA-trained coefficients), which inflates validation AUC above the training AUC (0.80–0.86 > 0.65–0.71). Recorded as a possible optimistic-bias note for the human auditor.
Figures / tables: Fig 3HFig 3GFig 3CFig 3E
R4_AUC1
Reported
0.80
Reproduced
0.797
within tolerance
R4_AUC3
Reported
0.84
Reproduced
0.837
within tolerance
R4_AUC5
Reported
0.86
Reproduced
0.857
within tolerance
R5_KM
Reported
logrank p<0.05; low-risk better OS
Reproduced
logrank p=1.98e-06; low better (death 0.9% vs 22.2%)
exact
R3_GENES
Reported
SRC,CTLA4,CTSE,LAMA2,ADAMTSL4
Reproduced
all 5 present & used
exact
R3_MODEL
Reported
-0.171*SRC -0.303*CTLA4 -0.092*CTSE +0.29*LAMA2 +0.12*ADAMTSL4
Reproduced
GSE32894-refit coefs differ (val cohort re-fit by design); 4/5 signs match
partial
COHORT_GSE
Reported
316
Reproduced
216 (authors' shipped data)
did not match
R6_TCGA_AUC
Reported
0.71/0.68/0.65 (1/3/5y)
Reproduced
not attempted (un-shipped GDC TCGA TPM; 80/20 deferral)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 71/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

The primary validation target reproduces cleanly: GSE32894 timeROC AUCs (0.797/0.837/0.857) match the reported 0.80/0.84/0.86 within rounding, KM separation is strongly significant (p=1.98e-06, low-risk better), and the 5-gene set is confirmed — so the prognostic core claim holds. However, three flags keep this at yellow: (1) the paper's n=316 cohort is not derivable from the authors' own shipped 216-sample data, and the TCGA training AUCs depend on an un-deposited TPM matrix; (2) the 'external validation' is actually an in-sample Cox re-fit on GSE32894 (validation AUC > training AUC), an optimistic bias that limits the central external-validation claim; (3) the repo is not runnable as shipped. The unresolved anomalies sit on the authors'/data-availability side rather than our method, with no fabrication evidence for the reproduced values.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

123.6 k
tokens (I/O) · 9.9 M incl. cache
15 min
runtime · 0.01 CPU-h
1.7 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine