Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genetically engineered stem cell-derived retinal grafts for improved retinal reconstruction after transplantation.

iScience · 2021
L1 84/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
84/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 63% of all assessed papers rank 392 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for a faithful 1:1 reproduction of the ONE module with public data. The paper is overwhelmingly wet-lab: 8 of 9 analysis modules are Bayesian (Stan) re-analyses of measurement tables (cell counts, ERG/MEA, fluorescence, behaviour, human IHC) that are NOT deposited in the repo or GEO and so cannot be reproduced. The single module with public data is the mouse retinal-organoid microarray (GEO GSE178653, 9 samples = 3 lines x 3 differentiation days). Re-running the repo's described pipeline (Agilent normalized signal -> drop controls -> log2(KO/wt) per day -> 4-fold cut) on the deposited data reproduces the paper's reported marker-gene differences: Bhlhb4-/- shows lower photoreceptor markers (Cnga1/2/3, Gucy2d, Rho) at all days -> 5/5 exact; Islet1-/- shows higher markers (Cabp4, Cnga1, Cngb1, Gnat2, Gngt1, Rho, Rcvrn, Sag) at the mature day DD23 -> 8/8 (5/8 if naively averaged over all days, because they are lower at the immature DD10 -> a developmental-timing artifact, not a contradiction). The fourfold threshold and the 'largely unaltered' overall pattern both reproduce. NOT attempted: the GO-enrichment bar charts of SFig2 (need un-shipped helper files go.obo/take.R/genes.xlsx/Agilent annotation csv) and all non-microarray modules (no deposited data). No fabrication concerns: every checked microarray claim is derivable from the deposited data and reproduces in the reported direction.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 84
    assessed: 2026-06-15 ⛓ 094b5cfea54e
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can genetically engineered ESC/iPSC-retinal sheets with reduced secondary retinal neurons (rod/cone bipolar cells via Bhlhb4 or Islet1 knockout) but intact photoreceptor layers improve graft-host neural integration and visual recovery after subretinal transplantation into end-stage degenerate (rd1) mouse retinas?

Core claims
  • Bhlhb4−/− and Islet1−/− grafts have markedly reduced bipolar cell populations after transplantation while retaining intact photoreceptor cell layers finding
  • Photoreceptors in bipolar-cell-KO organoids/grafts can functionally mature in vivo finding
  • Bhlhb4−/− and Islet1−/− grafts form more de novo synapses per host bipolar cell, improving neural integration finding
  • KO grafts better suppress heightened spontaneous spiking in the degenerated host retina finding
  • Mice transplanted with KO grafts perform better in light-guided behavior tests finding
  • Using genetically engineered cell lines (Bhlhb4−/−, Islet1−/−) to deplete bipolar cells is a strategy to enhance retinal sheet transplantation outcomes method
  • Bhlhb4−/− and Islet1−/− cell lines differentiate into retinal organoids with similar potency, timing, and gene expression as wild-type finding
  • Reduced intragraft bipolar cells leave photoreceptor output directly available to contact host bipolar cells, reducing spontaneous activity and improving visual recovery mechanism
Experimental setups
Assay System Perturbation Readout Platform
Retinal organoid differentiation with Nrl-GFP reporter live fluorescence imaging wt, Bhlhb4−/−, Islet1−/− mouse ESC/iPSC lines (Tg Nrl-GFP) Bhlhb4 KO / Islet1 KO / none (wt) Nrl-GFP fluorescence intensity over DD19–DD33; organoid morphology
RT-PCR / real-time PCR of key retinal genes wt, Bhlhb4−/−, Islet1−/− iPSC- and ESC-derived organoids (DD10, DD16, DD23) Bhlhb4 KO / Islet1 KO / none mRNA expression of Bhlhb4, Islet1 and other retinal markers
Microarray gene expression analysis Nrl-GFP;Ribeye-reporter iPSC-derived organoids (DD10, DD16, DD23) Bhlhb4 KO / Islet1 KO / none genome-wide expression patterns; differentially expressed genes / GO enrichment
Immunohistochemistry of in vitro organoids wt, Bhlhb4−/−, Islet1−/− organoids DD15 and DD29 Bhlhb4 KO / Islet1 KO / none PKCα, Chx10, Rx, GS, Nrl-GFP marker expression
Subretinal transplantation followed by immunohistochemistry of grafts rd1 mouse host retina with wt/Bhlhb4−/−/Islet1−/− grafts (4–5 weeks post-transplant) Bhlhb4 KO / Islet1 KO / none graft genotype counts of PKCα+, SCGN+, Chx10+, Calretinin+, Calbindin+ inner cells; rhodopsin+ photoreceptor content
Flat-mount confocal imaging with orthogonal reconstruction transplanted rd1 retinas (wt, Bhlhb4−/−, Islet1−/− grafts) graft genotype number of photoreceptor cells per graft
De novo synapse quantification (immunostaining of pre/postsynaptic markers) rd1;L7-GFP host mice transplanted with Ribeye-reporter-derived wt/Bhlhb4−/−/Islet1−/− grafts graft genotype; host sex number of host-graft synapses (RIBEYE + Cacna1s pairs) per L7-GFP host rod bipolar cell
Key results
  • PKCα+ rod bipolar cells consistently and markedly reduced in KO grafts vs wt wt 0.14 (95% CI 0.04–0.38) vs Bhlhb4−/− 0.04 (0.01–0.16) and Islet1−/− 0.02 (0.00–0.09) of inner cells
  • Cone bipolar (SCGN+) cells decreased in Islet1−/− but not clearly in Bhlhb4−/− wt 0.28 (0.16–0.42) vs Islet1−/− 0.19 (0.09–0.30), Bhlhb4−/− 0.25 (0.14–0.36)
  • Average number of de novo synapses per host bipolar cell notably increased in KO lines confidence in difference 98% Bhlhb4−/−-wt and 98% Islet1−/−-wt
  • Photoreceptor content not significantly different across genotypes ~0.6 rod photoreceptor content for all lines
  • Nrl-GFP differentiation time courses and maximum intensities broadly similar across wt and KO lines
  • Probability of host bipolar cells forming any synapse largely overlapping across genotypes (only modest improvement in Bhlhb4−/−) predicted probabilities 0.08–0.23; confidence in difference 84% wt-Bhlhb4−/−
  • Host sex affects synapse outcome with more synapses per bipolar in females confidence in difference 91% male-female
  • Chx10+ cells reduced in Bhlhb4−/− and slightly less in Islet1−/−; Islet1−/− Chx10+ cells showed weaker signal confidence in difference 99% Bhlhb4−/−-wt, 84% Islet1−/−-wt
Key statistics
  • mean 0.14 (95% CI 0.04–0.38) PKCα+ fraction of inner cells in wt grafts (rod bipolar fraction wt grafts (3 wt, 3 Bhlhb4−/−, 4 Islet1−/− mice))
  • mean 0.04 (95% CI 0.01–0.16) Bhlhb4−/−; 0.02 (0.00–0.09) Islet1−/− (PKCα+ rod bipolar fraction in KO grafts)
  • other confidence 97% Bhlhb4−/−-wt, 99% Islet1−/−-wt, 84% Bhlhb4−/−-Islet1−/− (confidence in graft genotype effect on PKCα+ cells)
  • other confidence 98% Bhlhb4−/−-wt, 98% Islet1−/−-wt, 64% Bhlhb4−/−-Islet1−/− (genotype effect on average synapses per host bipolar (9 wt, 9 Bhlhb4−/−, 9 Islet1−/− mice))
  • count 10,354 L7-GFP bipolar cell observations from 27 mice (synapse quantification dataset)
  • count 170 organoids (54 wt, 61 Bhlhb4−/−, 55 Islet1−/−) (Nrl-GFP growth curve analysis)
  • mean ~0.6 rod photoreceptor content (photoreceptor ratio all genotypes (2 wt, 3 Bhlhb4−/−, 3 Islet1−/− mice))
  • other confidence 91% male-female (effect of host sex on synapses per bipolar cell)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a Bayesian modeling framework to characterize stem-cell-derived retinal grafts. Continuous and count outcomes were fit with explicit generative models (e.g., exponential growth curves for Nrl-GFP intensity, binomial models for marker-positive cell fractions, and a zero-inflated Poisson model for synapses per host bipolar cell), and effects of graft genotype (and host sex) were reported as posterior distributions. Results were summarized with modes and 95% compatibility (credible) intervals, and group differences were expressed as 'confidence in the difference' (the posterior probability/fraction of the difference distribution above zero) rather than as null-hypothesis p-values.

Replicationmixed Sample sizeSample sizes reported as number of organoids, mice, and observations per analysis (e.g., 170 organoids; 8–11 mice per IHC marker; 10,354 cells from 27 mice); no formal power/sample-size calculation described Groupswt vs Bhlhb4-/- vs Islet1-/- grafts (with host sex as an additional factor for synapses) Pairingunpaired Randomization/blindingnot stated DispersionCI Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionInference based on posterior distributions/compatibility intervals rather than multiple-testing correction; none stated for the modeled comparisons
Statistical tests used
Test Applied to n Assumptions
Bayesian exponential growth-curve model (Nrl-GFP fluorescence over differentiation days) Figure 1B, 1C, S1; organoid GFP signal across wt/Bhlhb4-/-/Islet1-/- 170 organoids (54 wt, 61 Bhlhb4-/-, 55 Islet1-/-) stated
Bayesian binomial model of marker-positive fraction vs total inner cells Figure 2B; PKCα+, SCGN+, Chx10+ (and Calbindin+/Calretinin+ in S2) cell counts by genotype 11 mice PKCα (3 wt,3 Bhlhb4-/-,4 Islet1-/-); 11 SCGN (3,4,4); 9 Chx10 (2,3,4) stated
Bayesian model of photoreceptor cell number as a function of graft cell number Figure 2D; rod photoreceptor content by genotype 8 mice (2 wt, 3 Bhlhb4-/-, 3 Islet1-/-) stated
Zero-inflated Poisson model (probability of forming a synapse θ and mean synapses λ) Figure 3J–L, S4; synapses per L7-GFP host bipolar cell by graft genotype and host sex 10,354 L7-GFP bipolar cells from 27 mice (9 wt, 9 Bhlhb4-/-, 9 Islet1-/-) stated
Microarray differential expression comparison with GO enrichment analysis Figures 1E, 1F, S2; DD10/DD16/DD23 organoid expression and genotype comparisons not stated
Approaches that could also have been used
  • Group differences were summarized as the posterior 'confidence in the difference' (fraction of the difference distribution above zero) with 95% compatibility intervals.
    Could also: A frequentist mixed-effects model (e.g., generalized linear mixed model) with reported p-values and confidence intervals could also be used. — A frequentist summary would provide the p-values and effect estimates many readers and meta-analyses expect, complementing the Bayesian credible intervals already reported.
  • Counts of marker-positive cells were modeled with a binomial model and synapse counts with a zero-inflated Poisson model.
    Could also: A negative-binomial or zero-inflated negative-binomial formulation could also be applied to count outcomes. — A negative-binomial component would accommodate overdispersion beyond the Poisson mean–variance assumption, which can be useful when count variability is large.
  • Software and package details for the Bayesian models were not stated in the provided text.
    Could also: Reporting the specific tools and versions (e.g., R with brms/Stan, priors, and convergence diagnostics) could also be included. — Naming software, priors, and diagnostics (R-hat, effective sample size) aids exact reproducibility and lets readers assess model fit.
  • Inference relied on posterior distributions without an explicit multiplicity adjustment across the family of comparisons.
    Could also: For the microarray/GO analyses, a Benjamini-Hochberg FDR (or hierarchical shrinkage) approach could also be reported. — An explicit FDR or hierarchical/partial-pooling scheme would document control of false discoveries across the many genes and contrasts tested.
  • Sample sizes were described as numbers of organoids, mice, and observations per analysis.
    Could also: A brief statement of how sample size was determined (or an a priori power/precision rationale) could also be provided. — Describing the basis for n helps readers gauge the precision the design was intended to achieve, especially for the smaller per-genotype mouse groups.
  • Central estimates were reported as the posterior mode with 95% compatibility intervals.
    Could also: The posterior mean or median with the interval could also be reported. — Median/mean summaries are less sensitive to the shape of skewed posteriors and are a common convention, making cross-study comparison straightforward.

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
35
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE178653 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
RRID:AB_10003372 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_10013382 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_10842442 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2034062 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2068336 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2068506 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2069582 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2110656 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2783559 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_399431 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34409267 (KO-graft / Matsuyama et al. 2021, iScience)

Paper: "Genetically engineered stem cell-derived retinal grafts for improved retinal reconstruction after transplantation." Repo: https://github.com/matsutakehoyo/KO-graft @ commit f48d71a2206910b9585b2184e6baccdcd3b39180 (2021-07-29). Public data: GEO GSE178653 ("Microarray analysis of mouse retinal organoids", Agilent SurePrint G3 Mouse GE 8x60K v2.0 = GPL21810, 9 samples GSM5395271–GSM5395279 = 3 lines {wt, Bhlhb4−/−, Islet1−/−} × 3 days {DD10,DD16,DD23}).

How the paper splits into modules (from mouse/README.md)

Module Figures Pipeline Input data In scope?
organoid fluorescence Fig1, SFig1 R + Stan growth model image-derived fluorescence table (NOT shipped) OUT (no data)
organoid microarray Fig1, SFig1–2 R (tidyverse) load→filter→log2FC→GO GSE178653 (PUBLIC) IN
cell type Fig2, SFig3 Stan binomial cell-count table (NOT shipped) OUT (no data)
synapse Fig3, SFig4 Stan ZIP hierarchical synapse counts (NOT shipped) OUT (no data)
mERG Fig4, SFig6–7 R + Stan (MED64 MEA) MEA recordings (NOT shipped) OUT (no data)
two photon SFig8 Stan gamma-hurdle Ca-imaging table (NOT shipped) OUT (no data)
RGC_1s Fig5–6, SFig10–11 Stan clogit/lognormal MEA recordings (NOT shipped) OUT (no data)
SAS (behaviour) Fig7, SFig12 Stan binomial behaviour table (NOT shipped) OUT (no data)
human organoids Fig4,6,7 Stan (ordered_probit etc.) IHC/MEA tables (NOT shipped) OUT (no data)

Only the microarray module has public input data (GSE178653). All other modules are Bayesian (Stan) re-analyses of wet-lab measurement tables that are not deposited in the repo or GEO, so they cannot be reproduced (would require the authors' private measurement spreadsheets). This is an honest scope limit, not a quality judgement: the heavy-stats core of the paper is wet-lab-data-bound.

In-scope reproduction target (microarray, Fig 1 / SFig 1–2)

The repo scripts load micro array data.R + gene comparison.R implement: load 9-sample Agilent normalized signal → drop control probes + probes below background → average duplicate probes → compute log2(KO/wt) per DD → flag genes with >4-fold change (|log2FC|>2) → GO enrichment.

Concrete, falsifiable paper claims to reproduce (from Results/Fig 1 text):

  • C1 Platform: 56,605 probes covering 37,074 genes (sanity check vs GPL21810 / data).
  • C2 Threshold = fourfold (|log2FC|>2) — matches code filter(log2_exp > 2 / < -2).
  • C3 Bhlhb4−/− shows lower expression of photoreceptor markers Cnga1, Cnga2, Cnga3, Gucy2d, Rho (vs wt).
  • C4 Islet1−/− shows higher expression of Cabp4, Cnga1, Cngb1, Gnat2, Gngt1, Rho, Rcvrn, Sag (vs wt).
  • C5 Overall: expression pattern of retinal genes "largely unaltered" between genotypes (few genes pass the 4-fold cut) — quantify by per-line up/down counts.

Out of the hard last 20%: exact GO-enrichment level1/level2 bar charts (SFig2) need un-shipped helper files (go.obo, take.R, gen-log.R, genes.xlsx, Agilent AllAnnotations_*.csv); not attempted — direction-of-marker-genes (C3/C4) is the auditable core and is attempted instead.

Compute: all on «our HPC» SLURM (GEOquery + tidyverse conda env). Data on «infra».

Figures / tables: Fig 1
C1
Reported
56,605 probes / 37,074 genes (Agilent 8x60K)
Reproduced
62,976 features; 59,305 non-control; 27,306 gene symbols (GSE178653 / GPL21810)
partial
C2
Reported
fourfold threshold (|log2FC|>2)
Reproduced
|log2FC|>2 applied identically
exact
C3
Reported
Bhlhb4-/- LOWER: Cnga1,Cnga2,Cnga3,Gucy2d,Rho
Reproduced
all 5 DOWN vs wt at every DD (mean log2FC -0.44..-1.02)
exact
C4
Reported
Islet1-/- HIGHER: Cabp4,Cnga1,Cngb1,Gnat2,Gngt1,Rho,Rcvrn,Sag
Reproduced
8/8 HIGHER at DD23 (mature; log2FC +0.40..+1.03); 5/8 on cross-DD mean
within tolerance
C5
Reported
expression largely unaltered between genotypes
Reproduced
~0.3-1.8% of 27,306 genes pass 4-fold per line/day
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 84/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

Only the mouse retinal-organoid microarray module (Fig 1, SFig 1-2) has public input data (GEO GSE178653), and on that module the reproduction is faithful: Bhlhb4-/- lower markers (5/5), Islet1-/- higher markers (8/8 at mature DD23), and the 'largely unaltered' pattern all reproduce in the reported direction with no fabrication signal. The only deviations are on our/technical side — C1's Agilent design-count nomenclature (56,605 vs 27,306 symbols) and C4's wash-out when naively averaging across differentiation days, both explainable. The dominant limitation is data availability: 8 of 9 modules are wet-lab Stan re-analyses with no deposited data, so they are out of scope, not quality failures. Overall a solid reproduction with explainable deviations → yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

106.4 k
tokens (I/O) · 8.2 M incl. cache
13 min
runtime · 0.01 CPU-h
2.5 GB
peak RAM
4
HPC jobs
hummel
machine