Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

A curated collection of transcriptome datasets to investigate the molecular mechanisms of immunoglobulin E-mediated atopic diseases.

Database (Oxford) · 2019
L1 62/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Reported values were directly comparable
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
62/100
Reproducibility score
0.7 SD below mean
vs. all fields · 1187 studies
🎯 Scores higher than 23% of all assessed papers rank 897 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Curation/web-tool paper (GXB Gene Expression Browser over 33 GEO datasets). The linked repo BenaroyaResearch/gxbrowser @5a63c361 is the viewer WEB APP (Tomcat6/Grails2.1/Java7/MySQL/Mongo/R) and produces no number in the paper, so it was NOT stood up (out of scope, obsolete stack). The only pipeline-derived result is GXB's mean two-group fold-change (FC = mean(exp)/mean(ctrl), linear scale, genes ranked by FC), which the authors validate against two literature FCs. Per P16 I reimplemented that exact described method on the paper's own GEO data («our HPC» «job»: R4.3.3 + GEOquery2.70 on «infra», getGEO GSE8507/GSE88796; local python FC arithmetic). C1 (CD151/GSE8507, Job's PBMC vs healthy): single probe 204306_s_at, linear Affy signal, disease label from GSM description (Job's 56 / healthy 85). Baseline PBMC contrast -> FC 1.81 vs reported 1.7 (within-tol, delta 0.11/6.5%; reported value is itself rounded; FC>1 up-in-Job's in every grouping, range 1.45-1.81). C2 (CEACAM1/GSE88796, egg-allergic vs tolerant): 3 Illumina probes, log2->linearized; reported 1.69 is bracketed only by single probe ILMN_1716815 in egg-stimulated allergic-vs-control contrasts (1.68-1.74), while the mean of all 3 probes gives 1.06-1.47 -> graded partial because the paper pins neither the exact sample subset (which stimulation/control group) nor the probe-selection, so 1.69 is plausible/derivable but not uniquely reproducible. NOT attempted (hard 20% / out of scope): deploying the GXB web app, re-curating all 33 datasets, and independently recomputing the 33-dataset/1860-profile curation tallies (manual curation, not a pipeline). No fabrication signs: both reported GXB FCs are derivable from the public GEO data and agree in direction and order of magnitude; C2's residual gap is text ambiguity, not a non-derivable number. Described well enough for a 1:1 on the single-probe Affy case (C1); under-specified on the multi-probe stimulation case (C2).

💻 Code ↗ 🗄 Data: GSE87399

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 62
    assessed: 2026-06-14 ⛓ 0b035a2c6677
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

That publicly available GEO transcriptomic datasets relevant to IgE-mediated atopic diseases can be systematically curated, compiled, and displayed via the Gene Expression Browser (GXB) to create a validated resource facilitating mechanistic discovery and biomarker identification for allergy and related primary immunodeficiencies.

Core claims
  • A curated collection of 33 GEO transcriptome datasets (1860 profiles) relevant to IgE-mediated atopic disease was assembled and made available on the Gene Expression Browser (GXB) platform resource
  • GXB-displayed gene expression trends and fold-changes show good concordance with results reported in the datasets' associated original publications finding
  • A seven-strategy GEO search approach combined with manual curation can identify disease-relevant datasets despite heterogeneous study designs method
  • GXB is an open-source web application enabling query, visualization and annotation of ranked gene expression data across multiple compiled datasets resource
  • IgE plays roles beyond allergy, including modulation of vascular permeability, airway conduction, GI motility, and parasite defense mechanism
  • Allergic disease heritability is estimated at 33%-76%, indicating a strong genetic component alongside environmental factors finding
  • IgE dysregulation is implicated in several primary immunodeficiencies (e.g., DOCK8 deficiency, AD-HIES/Job's syndrome, CD40/CD40L deficiency) with heterogeneous phenotypes finding
Experimental setups
Assay System Perturbation Readout Platform
gene expression microarray/RNA-seq profiling PBMC/T cells allergen-specific immunotherapy circulating Tfh/Tfr cell balance gene expression Illumina HiSeq 2500 (GSE87399)
gene expression microarray profiling bronchial biopsy none (allergic asthma vs healthy) differentially expressed genes defining allergic asthma pattern Affymetrix Human Genome U133A Array (GSE41649)
gene expression microarray profiling human nasal epithelium cells IL-4 stimulation / disease state epithelial gene expression phenotype in rhinitis/asthma Affymetrix Human Gene 1.0 ST Array (GSE19190)
gene expression microarray profiling T cells (skin biopsy explants) intradermal immunotherapy (IDIT) gene expression profile changes with treatment Illumina HumanHT-12 V4.0 (GSE72324)
gene expression microarray profiling epidermis disease vs control (atopic dermatitis, psoriasis) differential gene expression Sentrix HumanRef-8 Expression BeadChip (GSE26952)
gene expression microarray profiling mast cells dexamethasone and FK506 drug treatment chemokine gene induction Affymetrix Human Genome U133 Plus 2.0 Array (GSE15174)
RNA-seq gene expression profiling B cells house dust mite (HDM) allergy status IL4R and related gene expression Illumina HiSeq 2000 (GSE52742)
gene expression microarray profiling neutrophils and PBMC Job's syndrome (AD-HIES) vs healthy control differential gene expression Affymetrix Human Genome U133 Plus 2.0 Array (GSE8507)
Key results
  • 33 datasets encompassing 1860 transcriptome profiles were curated and uploaded to GXB 33 datasets / 1860 profiles
  • Search and filtering funnel reduced initial candidate datasets from 435 to a final curated set of 33 435→313→115→53→33
  • Multiple datasets (e.g. GSE87399, GSE41649, GSE72324, GSE26952, GSE15823, GSE22528, GSE56681, GSE36842) showed 'good' to 'strong' validation concordance between GXB trend/fold-change and the associated publication
  • Several datasets (e.g. GSE19301, GSE72542, GSE64639, GSE51587, GSE44956, GSE54522, GSE70760) had data not available or reported differently in the associated publication, limiting validation
  • One dataset (GSE70050) had a discrepancy in sample names between GEO and SOFT files such that the reported group comparison could not be replicated
  • Allergic disease prevalence reaches ~20% in developed nations with sensitization rates approaching 40-50% in school-age children ~20%; 40-50%
  • Heritability of allergic disease and atopy is estimated at 33%-76% 33%-76%
Key statistics
  • count 33 datasets (final curated dataset collection size on GXB)
  • count 1860 transcriptome profiles (total samples across the 33 curated datasets)
  • count 435 datasets (combined results from seven independent search strategies before deduplication)
  • count 196 and 117 datasets (results of the two merged NCBI queries after duplicate removal)
  • count 115 datasets (after filtering for human samples and microarray/NGS platform)
  • count 53 datasets (after filtering for relevance to IgE-related atopic disease)
  • other ~20% of population (allergy prevalence in developed countries)
  • other 33%-76% (heritability estimate of allergic disease and atopy)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a database/resource paper describing the curation, upload, and web-based visualization of 33 publicly available GEO transcriptome datasets (1860 profiles) relevant to IgE-mediated atopic diseases. The paper's own analytical content is limited to a systematic search-and-filter workflow for dataset selection and a qualitative concordance check comparing gene expression trends and fold-changes in GXB against values reported in the original publications. No new inferential statistical analyses were performed by the authors; statistical methods used to generate the underlying datasets belonged to the original depositing studies.

Replicationunclear Sample sizeTotal collection described as 33 datasets encompassing 1860 transcriptome profiles; individual dataset sample sizes listed in Table 1 (range 5–685). No power calculation described for the curation study itself. GroupsCuration study: datasets retained vs. excluded at each filter step; validation: GXB fold-change/trend vs. originally published values Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Qualitative trend and fold-change concordance check (informal visual/manual comparison, not a formal statistical test) Validation of GXB-displayed gene expression against values reported in associated publications (Table 1, 'Trend validation' and 'FC validation' columns) not stated
Approaches that could also have been used
  • Validation of GXB data against published results was performed qualitatively, categorizing concordance as 'Strong', 'Good', or similar ordinal labels for a small set of manually selected genes per dataset.
    Could also: A quantitative validation metric such as Spearman rank correlation or Pearson correlation of log2 fold-changes between GXB-derived values and published values could also have been computed across all reported genes. — A numeric correlation coefficient would provide a continuous, reproducible measure of concordance and would allow readers to compare validation quality across datasets on a common scale rather than relying on subjective category assignment.
  • The dataset selection workflow was described narratively with step-by-step counts of datasets retained or excluded at each filter stage.
    Could also: A PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) or similar structured reporting framework could also have been used to document the search and selection process. — PRISMA-style reporting provides a standardized, externally reproducible audit trail for systematic searches and is increasingly expected for literature- or database-curation studies to facilitate replication and updating of the collection.
  • Seven independent search strategies were merged and duplicates removed, but inter-rater agreement for the manual relevance-filtering step is not described.
    Could also: A Cohen's kappa or percent-agreement statistic computed between two independent reviewers could also have been reported for the manual inclusion/exclusion step. — Inter-rater reliability metrics make explicit how reproducible the subjective curation decisions are, which is a recognized quality indicator for systematic curation studies.
  • The 33 curated datasets are described individually with sample-size information (Table 1) but no aggregate summary statistics across the collection are provided (e.g., distribution of sample sizes, platform types).
    Could also: Descriptive statistics (median and IQR of sample sizes, counts and proportions by platform, disease category, and experimental design type) could also have been presented in a summary table or figure. — Aggregate distributional summaries help readers quickly assess the coverage and potential biases of the collection (e.g., platform over-representation) without inspecting all 33 rows of Table 1.
  • Fold-change values used for validation are taken directly from GXB output without a stated normalization pipeline specific to this paper.
    Could also: A standardized re-normalization pipeline (e.g., RMA for Affymetrix, limma-voom or DESeq2 with a common normalization for RNA-seq) applied uniformly across all 33 datasets could also have been used prior to cross-dataset comparison. — Applying a consistent normalization method across heterogeneous datasets reduces platform- and study-specific technical variation and is a common approach in multi-study meta-analyses, potentially improving the comparability of fold-change estimates across datasets.
  • The paper describes the collection thematically by disease category and platform but does not assess cross-dataset consistency for any shared gene signatures.
    Could also: A cross-study meta-analysis (e.g., using the metaMA or MetaDE R package, or a random-effects model on log fold-changes) for genes that appear differentially expressed in multiple datasets could also have been performed. — Formal cross-study synthesis would allow the collection to support pooled effect-size estimates and heterogeneity assessment, going beyond visualization to quantitative evidence aggregation — a natural next step given the stated goal of biomarker discovery.
Software: Gene Expression Browser (GXB) · R (scripts available on GitHub at BenaroyaResearch/gxrscripts)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
4
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GPL11154 GEO in Methods (http://purl.org/orb/Methods)
also used by 2 papers:
GPL201 GEO in Methods (http://purl.org/orb/Methods)
also used by 2 papers:
GPL16699 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GPL16791 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GPL8300 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
10.6084/m9.figshare.7176851 DOI in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
GPL14550 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GPL20171 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GPL22841 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GPL2700 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE13619 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE13785 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE14842 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE15174 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE15823 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE19190 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE19301 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE22528 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE26952 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE36842 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE37157 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE41649 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE41861 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE44956 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE51587 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE52742 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE54522 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE56681 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE64639 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE70050 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE70760 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE70898 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE70900 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE72324 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE72542 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE8507 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE87399 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE88796 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE92866 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE99948 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-31290545

Paper: Huang et al. 2019, Database (Oxford) baz066 — "A curated collection of transcriptome datasets to investigate the molecular mechanisms of IgE-mediated atopic diseases." A data-curation / web-tool paper.

Artifacts:

  • Code: https://github.com/BenaroyaResearch/gxbrowser — the GXB web application (Grails/Tomcat/MySQL/Mongo/R viewer stack). It is infrastructure to host & display curated expression data, NOT an analysis pipeline that produces the paper's reported numbers. Standing it up reproduces nothing about the claims → out of scope.
  • Data: 33 GEO datasets (GSE...), incl. GSE87399 (brief's nominal accession), GSE8507, GSE88796 (the two used in the paper's own FC validation).

What is pipeline-derived (in scope)

The paper's only computational result is the fold-change (FC) ranking GXB performs: for a user-defined two-group comparison (experimental vs control), each gene's mean fold-change is computed in linear scale, and genes are ranked by FC. The authors validate this against published literature FCs with two concrete data points (Figure 2 / text):

# dataset comparison gene GXB FC (reported) literature FC lit source
C1 GSE8507 (PBMC) Job's syndrome vs healthy controls CD151 1.7 2.0 Holland et al. 2007
C2 GSE88796 egg allergic vs tolerant controls CEACAM1 1.69 1.6 Kosoy et al. 2016

In-scope reproduction (P16 = reimplement the described method on the paper's own data, equally valid): download the GEO series matrices, reconstruct the two-group comparison, compute the mean linear fold-change for CD151 / CEACAM1, and compare to the GXB-reported 1.7 / 1.69. This directly tests whether the paper's headline validation numbers are derivable from the public data.

Secondary (descriptive curation counts)

The collection-level numbers (33 datasets, 1,860 transcriptome profiles, 3 RNA-seq

  • 12 microarray platforms, 20 in-vitro + 13 ex-vivo, 7 disease categories) are manual-curation tallies, not pipeline outputs. Verifiable in principle by summing Table 1, but not a bioinformatic pipeline result → noted, lightly checked if Table 1 is machine-readable, not the focus.

Out of scope (not attempted)

  • Standing up the GXB Grails web app (no claim depends on it; obsolete stack: Tomcat 6 / Java 7 / Grails 2.1 / Mongo 2 / MySQL 5.1).
  • Re-curating all 33 datasets.
  • The interactive browser features (URL sharing, KEGG filtering, export).

Known reproduction risk (the hard 20%)

Both validation datasets are multi-group / multi-condition (GSE8507: 141 samples, PMN+PBMC × stim × timepoint; GSE88796: 132 samples, 3 phenotype groups × ±egg-stim). The paper does not pin the exact sample subset (which timepoint / stimulation state / which "tolerant control" group) behind each FC. Group definition is therefore the main uncertainty; computed FC is reported per the most natural grouping with the assumption flagged.

Figures / tables: Fig 2
C1
Reported
CD151 mean fold-change 1.7 (GXB), Job's syndrome PBMC vs healthy, GSE8507 (lit 2.0, Holland 2007)
Reproduced
1.81 (baseline PBMC control 0min, Job's n=7 vs healthy n=10; range 1.45-1.81 across grouping, all up in Job's)
within tolerance
C2
Reported
CEACAM1 mean fold-change 1.69 (GXB), egg allergic vs tolerant, GSE88796 (lit 1.6, Kosoy 2016)
Reproduced
single probe ILMN_1716815 = 1.68-1.74 in egg-stimulated allergic-vs-control contrasts (brackets 1.69); mean of 3 probes = 1.06-1.47
partial
S1
Reported
33 curated datasets / 1,860 transcriptome profiles
Reproduced
not independently recomputed (manual-curation tally, out of pipeline scope)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 62/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

A data-curation/web-tool paper whose only computational outputs are two literature-validation fold-changes; the linked repo is just the viewer app and produces no claim value. C1 (CD151/GSE8507) reproduces within tolerance (1.81 vs 1.7, direction up-in-Job's holds). C2 (CEACAM1/GSE88796) is bracketed by one of three probes (1.68–1.74 vs 1.69) but the paper pins neither the sample subset nor probe selection, so it is derivable yet not uniquely reproducible. Deviations are small and sit on the input/grouping side (our self-chosen cohort + the paper's underspecification), not in core computation; no fabrication, overall a solid but not clean-1:1 reproduction.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

120.8 k
tokens (I/O) · 8.4 M incl. cache
13 min
runtime · 0.01 CPU-h
2.5 GB
peak RAM
1
HPC jobs
hummel
machine