Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Positive Relationships between Intraocular Pressure and Thyroid Function in Graves' Ophthalmopathy: A Three-Aspect Evidence Analysis.

Yonsei Med J · 2025
L1 68/100 PQI 89
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
68/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 32% of all assessed papers rank 765 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce ONE of three aspects. The paper's own code link is bogus (github.com/PowerShell/PowerShell) and no authors' code exists; per P16 I reproduced the named third-party pipeline (GEO2R = limma-voom/edgeR TMM) on the paper's public data (GEO GSE199688), all compute on «our HPC», data on «infra». KEY RESULT 1:1: the central mechanistic claim reproduces - LOXL1-AS1 is significantly UP in TGF-b1-activated GO orbital fibroblasts (logFC +0.89, adj.P 5.9e-4), matching the paper's 'higher in TGF-b1-activated' (LOXL1 itself also up). DIFFERENT/partial: total DEG count 9706 (adj.P<0.05) vs reported 7673 - same order of magnitude (~half the genome), with 7673 falling between our no-cut (9706) and |logFC|>=1 (4289) thresholds; the paper underspecifies its GEO2R threshold/normalization so an exact match is not expected, and the reported down>up imbalance only reproduces under a logFC cut. Two flags for human audit (possible misrepresentation, not proven fraud): (1) code link is a text-mining false positive; (2) GSE199688 is a dihydroartemisinin drug study (PMID 35663306) with NO healthy controls - the paper's 'GO patients vs control subjects' framing is inaccurate; the real contrast is TGF-b1-treated vs untreated cells from the SAME GO patients, which is what I ran. NOT ATTEMPTED: SMR (LOXL1-AS1 FDR=0.04/HEIDI=0.25; hard 20%, needs GTEx V8 besd + BioBank Japan IOP GWAS + 1000G LD), DAVID GO(224)/KEGG(135) enrichment (depends on exact DEG list), clinical cohort (private/IRB), NHANES validation (manual SPSS, no code).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 68
    assessed: 2026-06-14 ⛓ 9ab3950dc5ee
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

This study investigates whether and how thyroid function influences intraocular pressure (IOP) in patients with Graves' ophthalmopathy (GO), and seeks to identify the molecular mechanism linking the two.

Core claims
  • Increased thyroid function is a risk factor for elevated IOP in patients with Graves' ophthalmopathy. finding
  • CAS, FT3 concentration, and sex are independent factors associated with high IOP (multiple regression r=0.46). finding
  • Higher LOXL1-AS1 expression has a causal effect on high IOP, as shown by SMR analysis, and is upregulated in GO patients, mechanistically connecting thyroid function and IOP. mechanism
  • IOP is linearly associated with GO duration, CAS, and FT3, FT4, and TSHRAb concentrations. finding
  • Thyroid function parameters (FT3, FT4, TT3) differ significantly between high and normal IOP groups in the NHANES validation dataset. finding
  • A combined three-aspect approach (clinical data, NHANES validation, SMR/GEO analysis) can establish the thyroid-IOP relationship in GO. method
Experimental setups
Assay System Perturbation Readout Platform
Retrospective clinical chart review with correlation/regression analysis Human GO patients (270 patients, 392 hospitalizations), Third Xiangya Hospital none IOP (non-contact tonometer) vs disease duration, CAS, exophthalmos, thyroid function parameters GraphPad Prism 10; SPSS v27
Cross-sectional epidemiological survey analysis NHANES 2007-2008 cohort (3848 individuals, US general population) none Thyroid function parameters (FT3, FT4, TT3, TT4, TSH, Tg, TgAb, TPOAb) compared between high vs normal IOP groups R v4.3.2; GraphPad Prism 10 (Mann-Whitney/Mann-Whitney test)
Summary-data-based Mendelian randomization (SMR) of eQTL GTEx V8 whole blood eQTL (670 donors); BioBank Japan IOP GWAS (177351 samples; 8448 cases, 168903 controls) none (genetic instrumental variables) Causal effect of gene expression on IOP (FDR, p-HEIDI) SMR software; 1000 Genomes European reference; Windows PowerShell 7.5.2
Differential gene expression / transcriptome analysis (GEO2R) GEO dataset GSE199688, TGF-β1-activated fibroblasts from GO patients vs controls GO disease state (TGF-β1 activation) Differentially expressed genes; LOXL1-AS1 expression; GO enrichment and KEGG pathways GEO2R; DAVID Bioinformatics
Key results
  • 142/392 (36.22%) hospitalizations had elevated IOP; durations 27.12 vs 38.56 months for high vs normal IOP groups 36.22%
  • Multiple linear regression model with CAS, FT3, and sex showed highest correlation coefficient r=0.46
  • LOXL1-AS1 identified by SMR in whole blood eQTL as positively associated with high IOP occurrence FDR=0.04, p-HEIDI=0.25
  • FT3 concentration significantly lower in high IOP group vs normal IOP group (NHANES) 2.98±0.35 vs 3.08±0.43 pg/mL, p<0.01
  • FT4 concentration significantly higher in high IOP group vs normal IOP group (NHANES) 0.81±0.19 vs 0.79±0.17 ng/dL, p<0.05
  • TT3 concentration significantly lower in high IOP group vs normal IOP group (NHANES) 103.96±21.98 vs 108.99±25.76 nM, p<0.01
  • 7673 differentially expressed genes identified in GO patients (3990 downregulated, 3683 upregulated); LOXL1-AS1 higher in GO 7673 DEGs
  • IOP linearly related to disease duration, CAS, FT3, FT4, and TSHRAb but not Tg, TSH, TPOAb, TgAb, exophthalmos, or severity
Key statistics
  • correlation r=0.46 (multiple linear regression model (CAS, FT3, sex) for IOP)
  • other FDR=0.04, p-HEIDI=0.25 (SMR identification of LOXL1-AS1 effect on IOP)
  • mean 2.98±0.35 vs 3.08±0.43 pg/mL, p<0.01 (FT3 in high vs normal IOP, NHANES)
  • mean 0.81±0.19 vs 0.79±0.17 ng/dL, p<0.05 (FT4 in high vs normal IOP, NHANES)
  • mean 103.96±21.98 vs 108.99±25.76 nM, p<0.01 (TT3 in high vs normal IOP, NHANES)
  • count 142 (36.22%) increased IOP; 250 (63.78%) normal IOP (clinical cohort IOP distribution)
  • count 7673 DEGs (3990 down, 3683 up) (GSE199688 GO vs control differential expression)
  • count 248 high IOP vs 3600 normal IOP (NHANES IOP group sizes)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper investigated associations between intraocular pressure (IOP) and thyroid function in Graves' ophthalmopathy (GO) using three complementary evidence streams: retrospective clinical data from 392 hospitalizations of 270 patients analyzed with Spearman rank correlation, simple linear regression, Kruskal-Wallis, and multiple linear stepwise regression; external validation in 3,848 NHANES 2007–2008 participants using Mann-Whitney U tests; and causal inference via summary-data-based Mendelian randomization (SMR) using GTEx V8 whole-blood eQTL data and a BioBank Japan IOP GWAS (n=177,351), supplemented by differential gene expression and pathway enrichment analysis of GEO dataset GSE199688. Results were reported primarily as p-value threshold inequalities, with dispersion metrics that were described inconsistently across sections.

Replicationunclear Sample size392 hospitalizations from 270 patients (142 high IOP, 250 normal IOP); NHANES validation n=3848; SMR outcome GWAS n=177,351; no formal power calculation or sample size justification described GroupsHigh IOP (≥21 mm Hg) vs. normal IOP (<21 mm Hg) in GO inpatients; same dichotomy applied to NHANES general-population sample Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionFDR <0.05 applied for SMR gene-expression–IOP causal tests and for GO/KEGG enrichment terms; no correction stated for the family of clinical Spearman correlations and simple linear regressions
Statistical tests used
Test Applied to n Assumptions
Chi-square test Comparison of categorical variables (sex, disease severity, prior glucocorticoid use, orbital decompression, smoking, alcohol) between high-IOP and normal-IOP clinical groups 392 hospitalizations (142 high IOP, 250 normal IOP) not stated
Spearman's rank correlation Univariate associations of IOP with disease duration, CAS, exophthalmos, and thyroid markers (FT3, FT4, Tg, TSH, TSHRAb, TPOAb, TgAb); also pairwise correlations among all predictor variables 392 hospitalizations stated
Simple linear regression Linear relationships of IOP with disease duration, CAS, exophthalmos, and thyroid function markers in the clinical dataset 392 hospitalizations not stated
Kruskal-Wallis test Association of IOP with disease severity (categorical: mild/moderate/severe) in the clinical dataset 392 hospitalizations stated
Multiple linear stepwise regression Identifying independent predictors of IOP from 11 candidate variables (disease duration, CAS, FT3, FT4, TSHRAb, sex, age, glucocorticoid use, orbital decompression, smoking, alcohol) 392 hospitalizations stated
Mann-Whitney U test Comparison of thyroid function parameters (TSH, TT3, FT3, TT4, FT4, Tg, TgAb, TPOAb) between high-IOP and normal-IOP groups in NHANES 2007–2008 3848 participants (248 high IOP, 3600 normal IOP) not stated
Summary-data-based Mendelian randomization (SMR) with HEIDI heterogeneity test Causal effect of whole-blood gene expression on IOP, using GTEx V8 eQTL (n=670 donors) and BioBank Japan IOP GWAS as exposure and outcome datasets respectively 177,351 (8,448 cases, 168,903 controls) for IOP GWAS; 670 donors for eQTL stated
Differential gene expression analysis (GEO2R, implicit limma-based) Comparison of gene expression between GO patients and controls in GEO dataset GSE199688 null not stated
Gene ontology and KEGG pathway enrichment analysis (DAVID Bioinformatics, FDR-corrected) Functional annotation of 7,673 differentially expressed genes from GSE199688 null na
Approaches that could also have been used
  • 392 hospitalizations from 270 patients were analyzed with standard regression treating each hospitalization as an independent observation
    Could also: A linear mixed-effects model with patient as a random intercept could also be used to account for within-patient correlation across repeated hospitalizations — Mixed-effects models explicitly partition between-patient and within-patient variance, which produces more accurate standard errors and inference for the fixed-effect estimates when the same individual contributes multiple observations
  • Eleven separate Spearman correlations and simple linear regressions were run across predictor variables in the clinical dataset without a stated family-wise or FDR correction
    Could also: Applying a Benjamini-Hochberg FDR or Holm-Bonferroni step-down correction across the family of univariate tests could also be used — With multiple simultaneous comparisons, the expected number of false discoveries under the null increases; a stated multiplicity adjustment makes the effective significance threshold explicit and aids interpretation
  • NHANES data were analyzed with Mann-Whitney U tests without accounting for the complex survey design (sampling weights, strata, and clusters)
    Could also: Survey-weighted logistic regression using NHANES sampling weights (e.g., via R's survey package) could also be used to produce population-representative estimates with appropriate standard errors — NHANES employs a stratified, multistage cluster sampling design; unweighted analyses can produce biased prevalence estimates and underestimated standard errors when the goal is inference to the U.S. population
  • High IOP in NHANES was operationalized as a binary self-report (ever told by an eye doctor you have glaucoma/high pressure), and groups were compared with non-parametric two-group tests
    Could also: Multivariable logistic regression adjusting for age, sex, race, and BMI could also be applied to this binary outcome to estimate adjusted odds ratios for each thyroid marker — Covariate-adjusted logistic regression would yield odds ratios with confidence intervals for each thyroid parameter while controlling simultaneously for the demographic confounders available in NHANES
  • The SMR analysis used GTEx whole-blood eQTL as the exposure tissue
    Could also: Tissue-matched eQTL data from orbital fibroblast, adipose, or ocular tissue could also be used if available; or a multi-tissue SMR approach (e.g., MT-SMR) integrating evidence across tissues could be applied — Whole-blood eQTL captures regulatory variation in a tissue that may differ from the orbital tissue most directly implicated in GO-related IOP elevation; tissue-matched eQTL can improve the biological specificity of the causal estimate
  • Continuous IOP values were available in the clinical cohort but were dichotomized at 21 mm Hg for group comparisons and for the NHANES validation
    Could also: Quantile regression or retaining IOP as a continuous outcome throughout all analyses could also be used — Dichotomization of a continuous outcome discards information about the magnitude of IOP elevation within each group; continuous-outcome models generally retain greater statistical power and can describe dose-response relationships
Software: GraphPad Prism 10 · SPSS 27 · R 4.3.2 · SMR (Yang Lab, yanglab.westlake.edu.cn) null · GEO2R (NCBI) null · DAVID Bioinformatics null

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE199688 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

Downstream reach in the literature

2 downstream papers · 1 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

GSE199688 GEO reused by 3 papers in the literature
Most-cited downstream papers:

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40992768

Paper: Zhang W et al., "Positive Relationships between Intraocular Pressure and Thyroid Function in Graves' Ophthalmopathy: A Three-Aspect Evidence Analysis." Yonsei Med J 2025. PMID 40992768 · PMCID PMC12479196 · DOI 10.3349/ymj.2024.0468.

The three aspects of evidence

  1. Clinical/observational — retrospective cohort, Third Xiangya Hospital (392 hospitalizations / 270 patients; IRB #Kuai 24368). Spearman correlation, Mann-Whitney/Kruskal-Wallis, chi-square, multiple linear stepwise regression (SPSS v27). → OUT OF SCOPE: private clinical data, not public, manual stats. non_pipeline + data_restricted.

  2. Population validation — NHANES 2007–2008 (n=3,848). Group comparison of thyroid hormones by IOP status. → OUT OF SCOPE (not attempted here): public data but a manual SPSS group-comparison, no shipped code, marginal pipeline character. Could be revisited but low value vs. effort.

  3. Genetic + transcriptomic — a. SMR (Summary-data-based Mendelian randomization): GTEx V8 whole-blood eQTL (670 donors) × BioBank Japan IOP GWAS (177,351; 8,448 cases). Reported result: LOXL1-AS1, FDR=0.04, p-HEIDI=0.25. Tool: SMR (Yang lab), 1000G EUR LD. → HARD 20%, deprioritized: requires large GWAS summary stats + GTEx besd + LD panel; feasible but heavy. Attempt only if primary target lands and budget allows. b. Transcriptomics on GSE199688 via GEO2R + DAVID/KEGG. Reported: 7673 DEGs (3990 down, 3683 up); 224 GO terms (BP 109, CC 85, MF 30); 135 KEGG pathways; LOXL1-AS1 higher in TGF-β1-activated fibroblasts. → IN SCOPE — PRIMARY TARGET. GEO2R = standard third-party tool on fully public GEO data (valid per P16; the paper's own "code" link is bogus, see below).

Code link is bogus

code/code.jsonhttps://github.com/PowerShell/PowerShell (Microsoft's PowerShell). This is a text-mining false positive / placeholder — NOT the authors' analysis code. No authors' code exists. Per P16 we proceed by applying the named third-party tool (GEO2R = limma-voom/edgeR) to the paper's public data (GSE199688).

CRITICAL dataset-mismatch finding (provisional, for human audit)

GSE199688 is NOT a "GO patients vs healthy controls" dataset. Its actual design (GEO record; original study PMID 35663306, Yang et al., about dihydroartemisinin (DHA) antifibrotic effects) is:

  • 9 samples, 3 groups × n=3, all from GO-derived orbital fibroblasts:
    • control (untreated): GSM5981744/745/746
    • TGFβ (TGF-β1 10 ng/mL, 48h): GSM5981747/748/749
    • DHA+TGFβ (DHA 20 µM + TGF-β1): GSM5981750/751/752

The paper states LOXL1-AS1 is "higher in TGF-β1-activated fibroblasts of patients with GO compared to control subjects." There are no healthy control subjects in GSE199688 — the "controls" are the same GO patients' untreated cells. The only comparison yielding "TGF-β1-activated vs control" is TGFβ (747-749) vs untreated control (744-746). We reproduce that comparison and flag the "control subjects" framing as a possible misrepresentation.

Reproduction plan (primary)

  • Comparison: TGFβ vs untreated control (n=3 vs n=3) on GSE199688 submitter count matrix (GSE199688_counts_anno.xls.gz).
  • Engine: limma-voom + edgeR TMM (the GEO2R RNA-seq workflow). Significance: GEO2R default adj.P.Val < 0.05 (Benjamini-Hochberg), no logFC cut → compare to 7673 / 3990 down / 3683 up. Also report at |logFC|>1 to bracket.
  • Extract LOXL1-AS1 logFC sign + adj.P.Val → compare to "higher in TGFβ".
  • All compute on «our HPC»; data on «infra»; only small results to «host».

Not attempted (declared)

  • Clinical cohort (private), NHANES validation (manual stats), SMR (heavy GWAS inputs), DAVID GO/KEGG enrichment counts (depends on exact DEG list + DAVID version; report only if DEG list reproduces cleanly).
Figures / tables: Fig 4
LOXL1AS1_direction
Reported
LOXL1-AS1 higher in TGF-b1-activated GO fibroblasts
Reproduced
logFC=+0.888, P=1.0e-4, adj.P=5.9e-4 (UP) [GeneID 100287616]
within tolerance
DEG_total
Reported
7673 DEGs (3990 down, 3683 up)
Reproduced
9706 @adj.P<0.05 (4952 up / 4754 down); 4289 @adj.P<0.05 & |logFC|>=1 (1767 up / 2522 down)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 68/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

The one feasible in-scope result — LOXL1-AS1 significantly up in TGF-b1-activated GO orbital fibroblasts (logFC +0.888, adj.P 5.9e-4) — reproduces 1:1 on the public GEO data, confirming the mechanistic bridge. The total-DEG count differs moderately (9706 vs reported 7673) and is explainable by an unreported GEO2R threshold, with the down>up split only reproducing under a |logFC|>=1 cut. Two authors'-side flags weigh against the paper: a bogus code link (github.com/PowerShell/PowerShell) and a dataset misattribution — GSE199688 is a DHA antifibrotic drug study with no healthy controls, contradicting the 'patients vs control subjects' framing. Net: yellow — a solid partial reproduction with explainable deviations and documentation/attribution problems on the authors' side, not fabrication.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

155.3 k
tokens (I/O) · 11.6 M incl. cache
16 min
runtime · 0 CPU-h
0.2 GB
peak RAM
6 (4 failed)
HPC jobs
hummel
machine