Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Large-Scale Phylogenomics of the Lactobacillus casei Group Highlights Taxonomic Inconsistencies and Reveals Novel Clade-Associated Features.

mSystems · 2017
L1 94/100 PQI 88
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Same input data as the authors
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
94/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 87% of all assessed papers rank 133 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the light, clearly-specified pipeline outputs 1:1. Method: downloaded the exact 183 NCBI assemblies listed in supplementary Table S1 (datasets CLI, on a «our HPC» compute node, «infra»), recomputed GC% per genome, and ran fastANI all-vs-all on the same genomes. GENOME SET + GC CONTENT (stage 04b) reproduce essentially exactly: per-species mean GC L. paracasei 46.30 (reported 46.3), L. rhamnosus 46.67 (46.7), L. zeae 47.75 (47.8), and crucially the L. casei GC is bimodal with EXACTLY 5 high-GC genomes (incl. type strain ATCC393) as the paper reports; original NCBI annotation counts 92/36/38/2 match exactly. ANI three-clade species separation (stage 05; paper used pyANI ANIb) was reproduced with fastANI (a faster third-party tool on the paper's identical data, explicitly valid per the brief): within-clade mins 97.6/94.1/96.5 and between-clade max 82.2 reproduce the same structure and the same clade-B-borderline signal, though absolute values differ because fastANI(mapping) != ANIb(BLAST). Authentic pyANI ANIb was launched («job») then cancelled in the makeblastdb stage as the hard 20% (all-vs-all BLAST on 183 ~3Mb genomes ~ 12h node time for a proxy->authentic refinement). NOT ATTEMPTED (deliberately, the heavy downstream chain): SPAdes assembly of AMBR2, Prokka annotation, Roary pangenome / 776 marker genes, RAxML phylogeny (Fig 2), OrthoFinder (5915 orthogroups / 521567 genes), and the interest-driven catalase / SecA2-SecY2 / glycosyltransferase analyses; the headline biological claims (only 10 L. casei after reclassification, catalase-positivity, 6/10 SecA2/SecY2) depend on that chain and are recorded as reported-only, not reproduced. No fabrication signals: every reproduced value is derivable from the shipped Table S1 accession list and matches closely. Grades are provisional; a human auditor decides (see AUDIT.md).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 94
    assessed: 2026-06-16 ⛓ 148c5b2d0b43
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The taxonomy of the Lactobacillus casei group (L. casei, L. paracasei, L. rhamnosus) is inconsistent; large-scale comparative phylogenomics of 184 genomes can resolve the group's structure, correct misclassifications, and reveal clade-associated functional features.

Core claims
  • The L. casei group resolves into three distinct clades (A, B, C) supported by phylogeny, GC content, ANI, and TETRA, and many strains are misclassified relative to their nearest type strain. finding
  • Reclassifying genomes to their most closely related type strain leaves only 10 genomes as L. casei sensu stricto (clade B), the smallest clade. finding
  • All 10 L. casei sensu stricto (clade B) strains carry a heme-dependent catalase gene and are catalase positive, making L. casei the first described catalase-positive species in the Lactobacillus genus. finding
  • L. casei AMBR2 was isolated from the human upper respiratory tract, the first L. casei strain from this niche. resource
  • 6 of 10 L. casei genomes contain a SecA2/SecY2 gene cluster encoding two putative glycosylated surface adhesins. finding
  • L. casei and L. paracasei are paraphyletic taxa under current annotations, being completely intermixed in clade A, whereas L. rhamnosus is monophyletic (clade C). finding
  • A high-quality maximum-likelihood phylogenetic tree was built from 776 conserved single-copy marker genes with L. nasuensis JCM 17158 as outgroup. method
  • Superoxide dismutase (SOD)-encoding genes are present only in clade A strains (69 of 70 genomes), restricting SOD to a single species rather than two. finding
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome sequencing and assembly L. casei AMBR2, human upper respiratory tract isolate none genome assembly / sequence
Comparative genomics / phylogenomics (maximum-likelihood tree) 184 L. casei group genome assemblies plus L. nasuensis JCM 17158 outgroup none phylogenetic tree from 776 single-copy marker genes roary
Pairwise genome comparison (ANIb and TETRA) all L. casei group genomes none ANI and tetranucleotide frequency distances / species delimitation BLAST-based ANIb
Orthogroup / gene content analysis 184 L. casei group genomes none core and accessory orthogroups, genes per genome OrthoFinder
Functional category mapping and PCoA all orthogroups of L. casei group genomes none predicted functional capacity across 18 categories eggNOG database v4.5
Gene/protein domain identification (catalase, SOD) L. casei group genomes (clade B catalase; clade A SOD) none presence of catalase (Pfam PF00199, PF05067) and SOD genes HMMER, Pfam
16S rRNA gene screening unclassified Lactobacillus sp. NCBI genomes none identification of additional L. casei group members RDP database v11
Catalase activity microbiology test 12 selected strains (4 per clade) grown in Weissella medium broth under anaerobic/aerobic conditions ± hemin/menaquinone H2O2 exposure; respiration-promoting supplementation bubble/froth formation indicating catalase activity 3% H2O2 standard lab test
Key results
  • After reclassification, only 10 genomes remain as L. casei sensu stricto (clade B), the smallest clade 10 genomes
  • Heme-dependent catalase gene found in all 10 clade B genomes; manganese-dependent catalase in 7 of 10 10/10 and 7/10
  • Only clade B strains tested (particularly AMBR2) were experimentally catalase positive na
  • 6 of 10 L. casei genomes contain a SecA2/SecY2 cluster with two putative glycosylated surface adhesins 6/10
  • SOD-encoding genes present only in clade A strains (69 of 70 genomes) 69/70
  • Inter-clade ANI values ≤85.1% while intra-clade ANI ≥96.1% (A), ≥93.6% (B), ≥96.3% (C), supporting three-clade separation ≤85.1% vs ≥93.6-96.3%
  • Clade B genomes show elevated GC content (47.74-47.76%) similar to L. zeae versus ~46.1-46.6% in L. paracasei-range 47.74-47.76% vs 46.10-46.60%
  • At least 17 assemblies show very high similarity to L. rhamnosus GG, suggesting repeated sequencing of the same strain 17 assemblies
Key statistics
  • count 184 L. casei group strains studied (183 public + AMBR2) (total genomes analyzed)
  • count 776 single-copy marker genes (phylogenetic tree construction)
  • correlation ANI inter-clade ≤85.1%; intra-clade ≥96.1% (A), ≥93.6% (B), ≥96.3% (C) (ANIb species delimitation)
  • other TETRA ≥0.9941 (A), ≥0.9872 (B), ≥0.9882 (C) (tetranucleotide frequency within clades)
  • count 521,567 genes total; 2,827 ± 141 genes/genome average (gene content; clustered into 5,915 orthogroups (1,814 core, 4,101 accessory))
  • count GC contents: L. paracasei 46.3%, L. rhamnosus 46.7%, L. zeae 47.8% (average GC content by species)
  • count 2,662 orthogroups mapped to category S (function unknown) (eggNOG functional annotation)
  • other heme-dependent catalase 487 amino acids; manganese-dependent catalase 269 amino acids (two catalase types identified in clade B)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a comparative/phylogenomics study of 184 Lactobacillus casei group genomes rather than a hypothesis-testing experimental study. The overall approach is descriptive: a maximum-likelihood phylogenetic tree built from 776 single-copy marker genes (with bootstrap support thresholds), pairwise genome-distance metrics (ANIb and TETRA) compared against published species-delimitation cutoffs, orthogroup/gene-content tabulation, and ordination (PCoA) of eggNOG functional categories. Results are reported as counts, ranges, threshold comparisons, and means ± SD, with one experimental catalase assay summarized qualitatively in a supplemental table; no formal inferential significance tests or P values are reported.

Replicationunclear Sample sizeSample size described as a census of available genomes meeting quality control (183 public + 15 unclassified + 1 new isolate = 184); no power/sample-size calculation GroupsThree phylogenetic clades (A, B, C) of the L. casei group; species annotations Pairingna Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Maximum-likelihood phylogenetic inference with bootstrap support (clades indicated at bootstrap >70) Fig. 2A-D, whole group and clade subtrees based on 776 single-copy marker genes 184 genomes (plus L. nasuensis outgroup); 776 marker genes not stated
Average nucleotide identity (ANIb) compared to published species cutoffs (95-96%) Fig. 3, pairwise within- and between-clade comparisons all pairwise comparisons among L. casei group genomes na
Tetranucleotide frequency correlation (TETRA) Fig. 3, within-clade pairwise values all pairwise comparisons among L. casei group genomes na
Principal-coordinate analysis (PCoA) of functional-category gene content Fig. 4, clustering of clades by eggNOG functional category orthogroups mapped per genome across 184 genomes na
Orthogroup clustering / core vs accessory gene-content tabulation Table 1, gene content per clade 521,567 genes; 5,915 orthogroups across 184 genomes na
Qualitative catalase activity assay (bubble/froth on H2O2 exposure) Table S2, 12 strains (4 per clade) under 4 growth conditions 12 strains tested na
Approaches that could also have been used
  • Species boundaries were assessed by comparing ANIb and TETRA values against previously published fixed cutoffs (e.g., 95-96% ANI).
    Could also: A clustering or gap-statistic / threshold-optimization approach (or model-based delineation such as those in GGDC or autoMLST) could also be applied to the pairwise distance matrix. — Such approaches would let the species boundary emerge from the data distribution itself and could quantify how robust the clade assignments are to the chosen threshold, which is useful when within-clade values fall near the cutoff (as noted for clade B).
  • Clade separation in functional capacity was shown via PCoA visualization of eggNOG category profiles.
    Could also: A complementary permutational test of group separation (e.g., PERMANOVA/adonis) or ANOSIM on the same distance matrix could also be reported. — This would attach a quantitative measure of between-clade versus within-clade variation to the visual clustering, giving readers an explicit summary statistic alongside the ordination plot.
  • Per-genome gene and orthogroup counts were summarized as mean ± SD per clade in Table 1.
    Could also: A 95% confidence interval or the full range/IQR could also accompany the mean. — For unequal and sometimes small group sizes (e.g., clade B with 10 genomes), a CI or range conveys uncertainty in the mean directly and is often preferred for comparing groups of differing n.
  • Branch support was conveyed by marking clades with bootstrap support >70.
    Could also: Reporting the full bootstrap values (or alternative supports such as SH-aLRT or Bayesian posterior probabilities) could also be presented. — Showing the continuous support values, rather than a single threshold dichotomy, lets readers gauge the strength of each split and compare relative confidence across nodes.
  • The catalase phenotype was assessed qualitatively (presence/absence of bubble formation) for 12 selected strains.
    Could also: A semi-quantitative or replicated readout with a defined scoring scheme could also be used. — Replicated, scaled measurements would allow the phenotype to be related to genotype (e.g., catalase gene copy presence) with an explicit comparison and would document measurement reproducibility.
  • The analysis used the full available set of genomes, including apparently redundant assemblies (e.g., ~17 highly similar to L. rhamnosus GG, duplicate L. zeae sequencing).
    Could also: A dereplication or down-weighting step (e.g., collapsing near-identical genomes) could also be applied before summary statistics. — Accounting for sampling redundancy would help ensure per-clade averages and ordination patterns are not influenced by repeatedly sequenced strains, which the authors already flag as a feature of the dataset.
Software: roary (single-copy marker gene identification) · OrthoFinder (orthogroup clustering) · HMMER (Pfam mapping, families PF00199/PF05067) · eggNOG database (functional category mapping) 4.5 · RDP database (16S rRNA screening) 11 · BLAST (ANIb implementation)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
91
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GCA_000414365 GCA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GCA_001063295 GCA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GCA_001434705 GCA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
HM070825 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PF00080 Pfam in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PF00081 Pfam in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PF00199 Pfam in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PF02109 Pfam in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PF02397 Pfam in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PF02777 Pfam in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PF04756 Pfam in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PF05067 Pfam in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PF09055 Pfam in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PRJEB21025 BioProject in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Reproduction scope — pmid-28845461

Paper: Wuyts et al. 2017, mSystems 2:e00061-17. "Large-Scale Phylogenomics of the Lactobacillus casei Group..." PMID 28845461 / PMC5566788. Code: https://github.com/LebeerLab/caseiGroup_mSystems_pipeline @ commit 78f013f861d28080d1ed11c4ceb236752381cafe (public, master, last push 2018-05-14). A collection of stage-numbered shell/R scripts (authors' own pipeline, P16-valid). Data: 183 public NCBI genome assemblies (exact accessions in supplementary Table S1 = sys004172127st1.csv) + 1 isolate (L. casei AMBR2, PRJEB21025).

Pipeline stages (from repo) and reproduction status

stage tool output in scope?
01 assembly SPAdes AMBR2 assembly from reads NO — single isolate, raw-read assembly, hard 20%
02 download genomes wget NCBI the 183 .fna set YES (we download the exact Table S1 accessions)
04a quality control QUAST N75/length/#contigs partial — QC thresholds (N75>10kb) are a selection rule, not re-run
04b GC content custom (per-genome GC%) per-species GC averages YES — PRIMARY, light compute
05 ANI pyani ANIb/TETRA within/between-clade ANI ranges YES — STRETCH, heavy all-vs-all BLAST
06 gene prediction Prokka .gff annotations NO — heavy, feeds 07/09
07 core alignment Roary 776 marker genes NO — heavy, depends on 06
08 phylogeny RAxML tree (Fig 2) NO — heavy, depends on 07
09 orthofinder OrthoFinder 5915 orthogroups, 521567 genes NO — heavy, depends on 06
10–11 gene content R / HMMER / BLAST catalase, SecA2/SecY2, GTs NO — interest-driven, downstream

In-scope targets (clearly-specified, pipeline-derived)

  1. Genome-set composition (Results, para 1): 183 public genomes; original NCBI annotation = 92 L. rhamnosus, 36 L. casei, 38 L. paracasei, 2 L. zeae (= 168) + 15 Lactobacillus sp. Audited directly against Table S1 (metadata), no compute. Provisional grade only — human confirms.

  2. GC content per species (Results "GC content"; Fig 1; methods stage 04b): reported averages L. paracasei 46.3 %, L. rhamnosus 46.7 %, L. zeae 47.8 %; L. casei splits into a large group (46.10–46.60 %) and a small group of 5 genomes (47.74–47.76 %). Reproduce by downloading the exact 183 Table S1 assemblies and recomputing GC% per genome. This is the clean 1:1 (≈183 data points → 4 per-species averages).

  3. ANIb three-clade separation (Results; Fig 3; methods "ANIb and TETRA", tool = pyani): within-clade ANIb min ≥96.1 % (A), ≥93.6 % (B), ≥96.3 % (C); between-clade ANIb ≤85.1 %. STRETCH — pyani ANIb all-vs-all on 183 genomes is a heavy BLAST job; attempted on «our HPC» only if budget allows.

Out of scope (and why)

Assembly, Prokka, Roary, RAxML, OrthoFinder, HMMER/eggNOG, the interest-driven catalase/SecA2/SecY2/glycosyltransferase analyses (stages 06–11): each is heavy, multi-tool, and chained off de-novo annotation — the hard last 20 %. The headline biological claims (only 10 L. casei after reclassification; all catalase-positive; 6/10 with SecA2/SecY2) depend on this full chain and are NOT attempted; they are recorded as reported-only.

Figures / tables: Fig2AFig1Fig3
C1_total_genomes
Reported
184 genomes (183 public + isolate AMBR2)
Reproduced
183 accessioned genomes in Table S1, all downloaded from NCBI; +AMBR2 = 184
exact
C2_orig_annotation
Reported
92 L. rhamnosus, 36 L. casei, 38 L. paracasei, 2 L. zeae (+15 Lactobacillus sp.)
Reproduced
92 / 36 / 38 / 2 / 15 sp. (exact)
exact
C3_gc_paracasei
Reported
46.3
Reproduced
46.297 (n=38)
exact
C4_gc_rhamnosus
Reported
46.7
Reproduced
46.671 (n=92)
exact
C5_gc_zeae
Reported
47.8
Reproduced
47.748 (n=2)
within tolerance
C6_gc_casei_split
Reported
L. casei GC bimodal: large group 46.10-46.60, small group of 5 genomes 47.74-47.76
Reproduced
exactly 5 high-GC L. casei genomes (47.75-47.93, incl. type strain ATCC393, all cladeB); other 31 at 46.14-46.53
exact
C7_ani_within_clade
Reported
>=96.1% (A), >=93.6% (B), >=96.3% (C) ANIb
Reproduced
fastANI within-clade min A 97.61 / B 94.06 / C 96.45
within tolerance
C8_ani_between_clade
Reported
<=85.1% ANIb between clades
Reproduced
fastANI between-clade max 82.21 (all below species threshold)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 94/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

What we could check reproduces cleanly: genome-set composition and GC content (C1–C6) match the paper essentially 1:1 from the shipped Table S1 accession list, including the diagnostic 5-genome high-GC L. casei split with type strain ATCC393 — no fabrication signal. The ANI three-clade species separation (C7–C8) was reproduced with fastANI as a proxy for the paper's pyANI ANIb, so absolute values differ (82.21 vs ≤85.1) on our methodological choice while the biological conclusion holds. The deviation is on our side (tool substitution) and benign, not authors'. Overall yellow because the paper's headline biological conclusions depend on a heavy downstream chain that was deliberately not run, leaving the central claim only partially confirmed.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

178.6 k
tokens (I/O) · 17.1 M incl. cache
28 min
runtime · 4.85 CPU-h
3 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine