Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Evaluating Distribution and Prognostic Value of New Tumor-Infiltrating Lymphocytes in HCC Based on a scRNA-Seq Study With CIBERSORTx.

Front Med (Lausanne) · 2020
L1 78/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 51% of all assessed papers rank 533 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the CORE DEG pipeline; the headline immune method (CIBERSORTx) and enrichment/prognostic steps are web-tool/stochastic and out of scope. C1 (exact): the authors' shipped DESeq2 table (All_results.csv, 25,299 genes) filtered at |log2FC|>1 & padj<0.05 reproduces the reported 755 DEGs / 323 up / 432 down EXACTLY; DEG-genesymbol.txt = exactly 755 genes; group.txt = 269 samples (135/134). C2 (within-tol): re-running DESeq2 High-vs-Low from a COMPLETELY INDEPENDENT quantification (recount3 STAR/gencode-v26 raw counts) with the shipped group labels matched 269/269 TCGA-LIHC patients and reproduced the shipped per-gene log2FoldChange at Pearson r=0.932 / Spearman 0.920 across 21,870 genes, recovering 558 of 696 shipped DEGs (Jaccard 0.733). Strong positive evidence the DEG analysis is genuine and reproducible (cross-pipeline, so not bit-identical). NOT attempted: CIBERSORTx signature matrix + cell fractions (Stanford web server, registration-gated; outputs shipped, which enabled the DEG repro); GO/pathway enrichment (PANTHER + Reactome/KOBAS web tools, only plotting code shipped; flagged discrepancy: enrichment input = 651 genes, not 755 DEGs); LASSO Cox prognostic model C-index 0.775/0.787 (stochastic CV, gene-expression model input not cleanly shipped). Documentation gaps: no shipped DESeq2 code (only inputs/outputs); TCGA download script has an undefined-variable bug (matrixl); group.txt High/Low labels swapped; volcano script uses |log2FC|>2 vs the |log2FC|>1 of the 755 headline.

💻 Code ↗ 🗄 Data: GSE76427

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 78
    assessed: 2026-06-15 ⛓ d19a06e4ba07
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

This study tests whether immune cell subsets newly identified by scRNA-seq (particularly the CD4-GZMA tumor-infiltrating lymphocyte subset) show distinct tumor-infiltration patterns and carry prognostic value in large bulk HCC cohorts when deconvolved with CIBERSORTx.

Core claims
  • CIBERSORTx can combine scRNA-seq-derived signatures with bulk HCC transcriptomes to estimate proportions of 11 TIL subsets and infer cell-type-specific gene expression. method
  • TIL subsets are distributed differently between tumor and normal tissue, with CD4-GZMA and CD8-LAYN higher in normal tissue and CD8-LEF1 higher in tumor. finding
  • Low proportion of CD4-GZMA cells is significantly associated with poor overall survival in HCC. finding
  • An immune risk score model built from CD4-GZMA cell-specific DEGs via LASSO Cox regression predicts HCC patient prognosis and validates across TCGA and ICGC cohorts. resource
  • CD4-GZMA cells and their molecular characteristics may yield therapeutic benefit in HCC immunotherapy. mechanism
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA-seq (Smart-seq2) HCC patient T cells from peripheral blood, tumor, and adjacent normal tissue (GSE98638) none single-cell transcriptome / 11 immune cell subset clusters used to build signature matrix Illumina HiSeq2500 and HiSeq4000
bulk RNA-seq deconvolution (CIBERSORTx) TCGA LIHC HCC cohort (269 patients) and ICGC LIRI cohort (229 patients) none relative proportions of 11 TIL subsets in tumor vs normal samples CIBERSORTx online platform (B-mode batch correction, 1000 permutations)
microarray/gene expression deconvolution (CIBERSORTx) GEO HCC datasets GSE76427, GSE64041, GSE36376, GSE14520 none immune subset proportions tumor vs normal
differential gene expression (DESeq2) TCGA HCC RNA-seq counts (269 patients), CD4-GZMA high vs low groups CD4-GZMA proportion median split 755 differentially expressed genes (|log2FC|>1.0, adj.p<0.05)
CIBERSORTx group-mode cell-type-specific expression imputation TCGA HCC tumor vs normal classes none 384 CD4-GZMA cell-specific DEGs (FDR<0.05, |log2FC|>1.0) CIBERSORTx group-mode
survival analysis (univariable Cox, LASSO Cox, Kaplan-Meier, ROC, nomogram) TCGA and ICGC HCC cohorts none prognostic genes, immune risk score, overall survival prediction R (glmnet, survminer, survival, rms, survcomp, ROCR, rmda)
Key results
  • Low CD4-GZMA cell proportion associated with poor overall survival (TCGA, validated in two other datasets) log-rank p<0.05
  • High CD4-CTLA4 cell proportion associated with poor overall survival in TCGA log-rank p<0.05
  • CD4-GZMA and CD8-LAYN cells higher in normal tissue; CD8-LEF1 higher in tumor tissue
  • 19 CD4-GZMA cell-specific genes associated with prognosis by univariable Cox regression p<0.05
  • Immune risk score nomogram showed strong predictive accuracy in TCGA cohort C-index 0.775
  • Immune risk score nomogram validated in independent ICGC cohort C-index 0.787
  • Time-dependent ROC for risk score in TCGA at 2, 3, 4, 5 years AUC 0.748, 0.759, 0.752, 0.747
  • Time-dependent ROC for risk score in ICGC at 2, 3, 4, 5 years AUC 0.686, 0.686, 0.748, 0.762
Key statistics
  • count 4,070 cells across 11 immune cell subsets (cells retained after filtering for signature matrix)
  • count 2,527 genes in signature matrix of 11 cell clusters (CIBERSORTx signature matrix)
  • count 755 DEGs (|log2FC|>1.0, adj.p<0.05) (CD4-GZMA high vs low TCGA groups)
  • count 384 CD4-GZMA cell-specific DEGs (FDR<0.05, |log2FC|>1.0) (CIBERSORTx group-mode tumor vs normal)
  • other C-index 0.775 (nomogram predictive accuracy in TCGA)
  • other C-index 0.787 (nomogram predictive accuracy in ICGC)
  • fold_change FNDC4 HR 1.3 (95% CI 1.1–1.6), p=0.0041 (univariable Cox, top prognostic gene)
  • count 13 genes with non-zero LASSO coefficients in ≥900/1000 evaluations (final immune risk score model genes)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper applied CIBERSORTx to deconvolve bulk RNA-seq profiles from TCGA (n=269), ICGC (n=229), and four GEO cohorts into proportions of 11 scRNA-seq-derived TIL subsets, compared distributions between tumor and normal samples via Wilcoxon tests, and assessed prognostic relevance of each subset using Kaplan-Meier/log-rank analysis. DESeq2 and CIBERSORTx group-mode were used to identify differentially expressed genes in CD4-GZMA-high vs. -low tumors; univariable Cox regression screened 19 prognostic genes, and LASSO Cox regression (10-fold CV, repeated 1,000 times) built a 13-gene immune risk score. Model performance was reported as time-dependent AUC, C-index, calibration curves, and Decision Curve Analysis, with independent validation in the ICGC cohort.

Replicationbiological Sample sizeTCGA n=269, ICGC n=229, GSE14520 n=213 stated in Table 2; scRNA-seq reference dataset n=4,070 cells stated; no formal power calculation or sample size justification mentioned GroupsTumor vs. normal tissue (immune infiltration); CD4-GZMA high vs. low proportion groups (DEG analysis); high vs. low immune risk score (survival); TCGA (model building) vs. ICGC (validation) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsyes Multiplicity correctionBenjamini-Hochberg FDR: implicit via DESeq2 adjusted p<0.05 for bulk DEG analysis; explicit FDR<0.05 for CIBERSORTx group-mode DEGs and GO/pathway enrichment; no correction stated for the family of Wilcoxon tests across 11 subsets and 6 datasets or for the multiple log-rank survival tests
Statistical tests used
Test Applied to n Assumptions
Wilcoxon rank-sum test Comparison of 11 immune cell subset proportions between tumor and normal samples across 6 datasets (TCGA, ICGC, GSE76427, GSE64041, GSE36376, GSE14520); violin plots in Figure 2 Varies by dataset: TCGA n=269, ICGC n=229, GSE14520 n=213; exact n for other GEO datasets not stated in text not stated
Log-rank test Kaplan-Meier overall survival comparisons for each of 6 immune cell subsets (Figure 3) and for high vs. low immune risk score groups (Figures 6A, 7A) in TCGA and ICGC cohorts TCGA n=269; ICGC n=229; effective n after CIBERSORTx p<0.05 filter not separately stated not stated
Univariable Cox proportional hazards regression (Wald test) Association of each CD4-GZMA cell-specific DEG with overall survival in TCGA cohort (Table 3) TCGA n=269 (or CIBERSORTx-filtered subset; exact effective n not stated) not stated
DESeq2 negative binomial GLM with Wald test Differentially expressed genes between CD4-GZMA high and low proportion groups defined by median fraction in TCGA (Figure 4) TCGA n=269 (count-level RNA-seq data) not stated
LASSO Cox regression with 10-fold cross-validation (repeated 1,000 times) Construction of 13-gene immune risk score model in TCGA cohort (Figure 6); genes selected if non-zero coefficient in ≥900/1,000 evaluations TCGA n=269 (or CIBERSORTx-filtered subset; exact n not restated here) not stated
Time-dependent ROC / AUC (ROCR package) Discrimination performance of immune risk score at 2, 3, 4, and 5 years in TCGA and ICGC cohorts (Figures 6B, 7B) TCGA n=269; ICGC n=229 na
Approaches that could also have been used
  • Overall survival groups were defined by dichotomizing CIBERSORTx-estimated cell fraction at the median for Kaplan-Meier analysis
    Could also: Cox proportional hazards regression treating cell fraction as a continuous predictor — or modeled via restricted cubic splines — could also characterize the association — Continuous modeling retains information across the full range of values and does not require selection of a threshold; it also allows the shape of the dose-response relationship to be visualized and tested
  • Log-rank tests were performed separately for each of the 11 immune cell subsets across multiple datasets without a stated multiplicity correction
    Could also: A Benjamini-Hochberg FDR or Bonferroni correction applied across the family of survival tests could also be reported alongside the uncorrected p-values — When many parallel log-rank tests share the same outcome data, a familywise or FDR correction quantifies the expected proportion of false discoveries across the full set of comparisons
  • Nineteen DEGs were pre-selected by univariable Cox regression (p<0.05) before being entered into LASSO Cox regression
    Could also: LASSO Cox regression applied directly to the full set of CD4-GZMA-specific DEGs (FDR<0.05, |log2FC|>1) without prior univariate filtering could also be used — Two-stage filtering (univariate then penalized) can exclude variables that contribute to prognosis only in combination; applying the penalty directly to the DEG set allows the regularization term to perform variable selection in a single stage
  • Samples with a CIBERSORTx deconvolution p-value ≥0.05 were excluded from downstream analyses
    Could also: Sensitivity analyses retaining all deconvolved samples (or using alternative confidence thresholds) alongside the filtered analysis could also be reported — If excluded samples are systematically different in clinical or expression characteristics, reporting results under multiple thresholds communicates how sensitive conclusions are to the filtering choice
  • Model discrimination was summarized using time-dependent AUC at selected time points (2, 3, 4, and 5 years)
    Could also: Harrell's concordance index (C-statistic) or Uno's C-statistic could also summarize overall discrimination across the entire follow-up in a single value — A global C-statistic integrates discrimination over all observed event times rather than at pre-specified points, providing a complementary summary particularly when the censoring pattern varies across cohorts
  • Bulk-level DEG analysis (DESeq2 on TCGA counts comparing high vs. low CD4-GZMA fraction groups) and cell-type-specific DEG analysis (CIBERSORTx group-mode) were conducted as parallel, largely separate analyses
    Could also: A formal overlap analysis reporting concordance of direction and magnitude of log-fold-changes between the two DEG lists could also be presented alongside the Venn diagram — Quantifying the agreement (or divergence) between bulk and cell-type-deconvolved gene expression estimates helps readers assess how much the cell-type-specific deconvolution step adds beyond standard bulk-level differential expression
Software: R 3.6.1 · DESeq2 · CIBERSORTx (online platform) · glmnet (LASSO Cox regression) · survminer (optimal cut-off, KM plots) · survival (Kaplan-Meier, log-rank, Cox) · rms (nomogram) · survcomp (C-index) · ROCR (time-dependent ROC/AUC) · rmda (Decision Curve Analysis) · KOBAS (pathway enrichment) 3.0

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
31
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33043022

Paper: Li et al. 2020, Front Med 7:451. "Evaluating Distribution and Prognostic Value of New Tumor-Infiltrating Lymphocytes in HCC Based on a scRNA-Seq Study With CIBERSORTx." PMID 33043022 / PMC7527443 / DOI 10.3389/fmed.2020.00451.

Code: https://github.com/szlilixing/Transcriptome-analysis @ commit 1668b28ab1c30cfd550276268d7e4d5b4734ad46 (pushed 2020-06-28). Authors' own R code, organized as folders each holding a *code.txt script plus the shipped input/output data file(s) for that figure.

Data: TCGA-LIHC (RNA-seq, n=269 tumor), ICGC LIRI-JP (n=229), GEO GSE76427 / GSE64041 / GSE36376 / GSE14520 (microarray, immune-infiltration validation), GSE98638 (scRNA-seq, 4,070 T cells, signature-matrix source).

Pipeline map (what produces each reported result)

Result (reported) Pipeline Shipped artifact In scope?
scRNA signature matrix (2,527 genes × 11 clusters) from GSE98638 CIBERSORTx web "Create Signature Matrix" (B-mode, Smart-seq2) none (only used downstream) OUT — CIBERSORTx is a registration-gated web app (cibersortx.stanford.edu), not runnable headlessly; signature matrix not shipped
Immune-cell fractions per bulk sample (TCGA/ICGC/GEO) CIBERSORTx web "Impute Cell Fractions" (1000 perms, B-mode) Violin plot/LIHC-Fraction.txt (fractions) OUT for re-running CIBERSORTx; the fractions themselves are shipped → downstream steps start from them
755 DEGs (323 up / 432 down), |log2FC|>1 & padj<0.05, High vs Low CD4-GZMA group, TCGA-LIHC DESeq2 on TCGA-LIHC counts, High/Low split Volcano plot/All_results.csv (full DESeq2 table), Heatmap/ntdeseq_res.txt, Heatmap/group.txt (269 grp labels), Heatmap/DEG-genesymbol.txt IN — (a) internal-consistency of shipped table vs reported counts; (b) re-run DESeq2 on TCGA-LIHC counts + shipped group labels
Volcano / Heatmap figures ggplot2 / pheatmap on the DEG table scripts + data shipped IN (derived from DEG table; figure-level, lower priority)
GO + KEGG/pathway enrichment of the 755 DEGs clusterProfiler / KOBAS 3.0 Bubble plot of GO analysis/F3-GO.txt, Bar plot of Pathway analysis/F3-Pathway.txt IN (secondary)
Univariable Cox (Table 3, 19 genes w/ HR,p) + LASSO Cox 13-gene model, C-index 0.775 (TCGA) / 0.787 (ICGC), time-dependent AUCs survival/glmnet on clinical+expression Construction of clinical model code.txt, Kaplan-Meier analysis/clinical-info-KM.txt IN (tertiary; LASSO has CV randomness → approximate)
GEO GSE76427 download → probe→symbol matrices GEOquery + illuminaHumanv4.db Date processing scripts partial scope — deterministic but not a headline number; low priority
Kaplan-Meier per immune cell type (CD4-GZMA, CD4-CTLA4, …) survminer on fractions + clinical KM-survminer code.txt, clinical-info-KM.txt IN (secondary)

Out-of-scope, with reason

  • CIBERSORTx signature-matrix construction and cell-fraction imputation — the central method — runs only on the Stanford CIBERSORTx web server (account/token gated, no headless CLI without a Docker token + the scRNA signature). Not attempted. Its outputs (cell fractions) are shipped (LIHC-Fraction.txt), so downstream pipeline steps can be reproduced starting from the shipped fractions/groups.
  • scRNA-seq clustering into 11 subsets (from the source scRNA paper, Zheng et al. GSE98638) — upstream, not re-derived here.

Reproduction priority (80/20)

  1. C1 (cheap): internal consistency — does shipped All_results.csv filtered at |log2FC|>1 & padj<0.05 yield 755 DEGs / 323 up / 432 down? Do group.txt (135 High/134 Low) and DEG-genesymbol.txt (755) match?
  2. C2 (core): re-run DESeq2 on TCGA-LIHC raw counts + shipped High/Low group labels; compare log2FC/padj and DEG counts to All_results.csv.
  3. C3 (secondary): GO/KEGG enrichment on the 755-gene list vs shipped F3-GO / F3-Pathwa
Figures / tables: Fig 3
C1_deg_total
Reported
755 DEGs
Reproduced
755
exact
C1_deg_up_down
Reported
323 up / 432 down
Reproduced
323 up / 432 down
exact
C1_deg_list_size
Reported
755-gene DEG list
Reproduced
755 genes (DEG-genesymbol.txt)
exact
C2_log2FC_corr
Reported
DEG analysis (High vs Low, TCGA-LIHC)
Reproduced
indep. DESeq2 on recount3 counts: log2FC Pearson r=0.932, Spearman 0.920 (21,870 genes)
within tolerance
C2_deg_overlap
Reported
755 DEGs
Reproduced
558/696 shipped-sig DEGs recovered, Jaccard 0.733; 269/269 barcodes matched
within tolerance
C3_go_enrichment
Reported
GO/Reactome enrichment of DEGs
Reproduced
not attempted (PANTHER+Reactome web tools); input 651 genes != 755 DEGs (discrepancy)
partial
C4_prognostic_cindex
Reported
C-index 0.775 (TCGA) / 0.787 (ICGC)
Reproduced
not attempted (stochastic LASSO; model input not cleanly shipped)
partial
CIBERSORTx_fractions
Reported
scRNA signature + immune cell fractions
Reproduced
not attempted (registration-gated web server)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

The core DEG pipeline reproduces cleanly: the reported 755/323/432 DEGs derive exactly from the shipped DESeq2 table, and a completely independent recount3 quantification concords strongly (log2FC Pearson r=0.932, 558/696 shipped DEGs recovered) with 269/269 barcodes matched — strong evidence the DEG analysis is genuine, not fabricated. Deviations are on our/technical side (the 937-vs-755 count gap is purely because recount3 tests ~2x more genes) plus authors'-side documentation issues (no DESeq2 code, 651≠755 enrichment input, |log2FC|>2-vs->1 threshold mismatch, swapped labels). However, the paper's central claims — CIBERSORTx immune-cell distribution and the prognostic C-index 0.775/0.787 — were untestable (registration-gated web server / stochastic LASSO with un-shipped input), so the headline conclusion is confirmed only in limited part, not refuted. Overall yellow: solid, reproducible DEG core with explainable deviations but unverified central methodology.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

167.4 k
tokens (I/O) · 12.7 M incl. cache
19 min
runtime · 0.05 CPU-h
5.1 GB
peak RAM
3 (1 failed)
HPC jobs
hummel
machine