Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Exposure to the widely used herbicide atrazine results in deregulation of global tissue-specific RNA transcription in the third generation and is associated wit

Nucleic Acids Res · 2016
L1 10/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +10
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
10/100
Reproducibility score
3.6 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 0% of all assessed papers rank 1169 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL (described well enough to run; result DIFFERENT in magnitude, same in direction). Reproduced the authors' RNA-seq pipeline end-to-end on their own public data (GSE81091/SRP074350: all 18 F3 PE-100 runs SRR3475391-408, testis/liver/brain x ctrl/atz x3): FastQC -> TopHat2 2.1.1 (Ensembl mm9/NCBIM37 rel-67; 97.9-98.2% mapping all 18) -> Cufflinks 2.2.1 -> Cuffmerge (17/18) -> Cuffquant (18/18) -> Cuffnorm -> custom DE (>50th-quantile in >=1 cond, >2-fold, limma BH-FDR<5%). HEADLINE NUMBERS DO NOT MATCH: reported 1419 total DE transcripts (1322 testis + 69 liver + 28 brain; 704 genes) vs reproduced 416 (217 testis + 128 liver + 71 brain; 397 genes). The QUALITATIVE finding reproduces -- testis is the most-affected tissue in our run too -- but the extreme testis-dominance (93% reported) is much weaker here (52%): testis came out far lower, liver/brain higher. Dominant cause = the DE filter is underspecified ('50th quantile of all values' is ambiguous; per-tissue median ~0 for liver/brain makes that pre-filter nearly inert) plus limma-on-FPKM-after-hard-FC-prefilter not being the standard Cuffdiff path, compounded by TopHat2/Cufflinks non-determinism. No fabrication signal: the gap is fully explained by underspecification, not by unsupportable reported values. NOT attempted: secondary C6-C10 (need CPC/CPAT/Cuffcompare/APA steps), ChIP-seq H3K4me3 path, all wet-lab/qPCR/motif/external-overlap. Two pipeline bugs fixed this run: cuffnorm group-arg collapse («job») and a front1-only TMPDIR breaking the R env build on the compute node («job»); final clean run = «job».

💻 Code ↗ 🗄 Data: GSE81093

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-16 ⛓ b0d8c9c48dce
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The authors hypothesized that embryonic exposure to the herbicide atrazine (ATZ) during the E6.5–E15.5 developmental window causes heritable epigenetic reprogramming and affects reproduction in subsequent (F1 and F3) generations.

Core claims
  • Embryonic ATZ exposure affects meiosis, spermiogenesis and reduces spermatozoa number in F3 generation male mice finding
  • Changes in testis cell types originate from a modified transcriptional network in undifferentiated spermatogonia mechanism
  • ATZ exposure dramatically increases the number of transcripts with novel transcription initiation sites, spliced variants and alternative polyadenylation sites finding
  • There is a global decrease in H3K4me3 occupancy in testes of third-generation (F3) ATZ-lineage males finding
  • Regions with altered H3K4me3 occupancy in F3 ATZ-derived males correspond to altered H3K4me3 occupancy in F1 generation, and 74% of changed peaks in F3 are associated with enhancers finding
  • Regions with altered H3K4me3 occupancy are enriched in SP family and WT1 transcription factor binding sites finding
  • Embryonic ATZ exposure effects on development and epigenetic marks are transferred up to three generations finding
  • This is the first study integrating genome-wide ChIP-seq and RNA-seq across F1 and F3 generations for toxicant transgenerational inheritance, generating novel sequencing data for outbred CD1 mice resource
Experimental setups
Assay System Perturbation Readout Platform
ChIP-seq (H3K4me3) testis tissue, F1 and F3 male mice (ATZ-lineage vs control) embryonic ATZ exposure (F0 dams, 100 mg/kg/day) genome-wide H3K4me3 occupancy/differential peaks Illumina HiSeq2000
strand-specific paired-end RNA-seq testis, liver and hypothalamus, F1 and F3 male mice embryonic ATZ exposure differentially expressed genes/transcripts, novel TSS, splice variants, alternative polyadenylation
H&E histology testis sections, F3 males embryonic ATZ exposure testis architecture/morphology
immunostaining (ZBTB16, GATA1) testis sections, F3 males embryonic ATZ exposure relative proportion of germ and Sertoli cells
FACS dissociated testis cells, F3 males embryonic ATZ exposure relative proportion of testicular cell types
sperm counting epididymis, F1 and F3 males embryonic ATZ exposure spermatozoa number
immunostaining of meiotic surface spreads (SYCP3, SYCP1, TERF1) testis, F3 males embryonic ATZ exposure synaptonemal complex formation, chromosome synapsing defects, telomere connections
Western blot whole testis extract / purified histone fraction, F3 males embryonic ATZ exposure protamine 2 and H4K5Ac protein levels
Key results
  • Spermatozoa number significantly decreased in ATZ-derived F1 and F3 males compared to controls ~30% decrease in F3
  • 704 genes corresponding to 1419 differentially expressed transcripts identified genome-wide 704 genes / 1419 transcripts
  • Protamine 2 protein level decreased in F3 ATZ-derived testis 2.6-fold
  • H4K5Ac level decreased in purified histone fraction of F3 ATZ-lineage males 1.3-fold
  • Significant increase in meiotic synapsing defects in F3 ATZ male progeny
  • Global decrease of H3K4me3 occupancy observed in F3 generation males
  • Majority of altered H3K4me3 peaks in F3 correspond to enhancer regions 74%
  • Altered H3K4me3 regions enriched for SP family and WT1 transcription factor binding motifs
Key statistics
  • fold_change ~30% decrease (spermatozoa count decrease in F3 ATZ vs control)
  • fold_change 2.6 times decrease (protamine 2 protein level in F3 ATZ testis)
  • fold_change 1.3 times decrease (H4K5Ac level in F3 ATZ-lineage males)
  • count n=199 control, n=195 ATZ (quantitative analysis of meiotic (synaptonemal complex) defects)
  • count 704 genes / 1419 differentially expressed transcripts (genome-wide RNA-seq differential expression analysis)
  • other 74% (proportion of altered F3 H3K4me3 peaks associated with enhancers)
  • pvalue P-value threshold <10E-5 (MACS 2.0.1 H3K4me3 peak calling threshold)
  • pvalue FDR <10% (ChIP-seq peaks); FDR <5% (RNA-seq DEGs) (Limma test thresholds for differential peak/transcript calling)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study combined an outbred CD1 mouse transgenerational exposure design with genome-wide ChIP-seq (H3K4me3) and RNA-seq, plus targeted qPCR and cytological/Western-blot assays. Genomic differential analyses (differential peaks and differentially expressed transcripts) were filtered by fold change and then assessed with the R/Limma test under FDR thresholds (10% for ChIP-seq, 5% for RNA-seq), with ChIP-seq replicate concordance checked by an irreproducible discovery rate criterion. Targeted qPCR comparisons used Student's t-test on at least four independent experiments. Results were largely reported as fold changes and counts of significant features.

Replicationbiological Sample sizeAt least three independent lineages per group; two biological replicates for ChIP-seq, three biological replicates per tissue for RNA-seq, and at least four independent experiments for qPCR; no formal power/sample-size calculation described GroupsATZ-derived vs vehicle/control males across F1 and F3 generations Pairingunpaired Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionFalse discovery rate control via the Limma test (FDR < 10% for ChIP-seq differential peaks, FDR < 5% for RNA-seq transcripts); specific FDR method not named
Statistical tests used
Test Applied to n Assumptions
Student's t-test qPCR/RT-qPCR comparisons and ChIP-qPCR enrichment fold changes (ATZ vs control) duplicates of at least four independent experiments not stated
Limma test differential H3K4me3 peak calling between ATZ-treated and control ChIP-seq samples two biological replicates per condition not stated
Limma test identification of differentially expressed transcripts/genes from RNA-seq (testis, liver, hypothalamus) three biological replicates per tissue not stated
Irreproducible discovery rate (IDR) criterion confirming similarity of ChIP-seq biological replicates for peak sets two biological replicates na
Significance test (method not stated) quantitative analysis of meiotic synapsis defects (n=199 control, n=195 ATZ cells) and spermatozoa counts n = 199 control and n = 195 ATZ-derived cells not stated
Approaches that could also have been used
  • Differential expression and differential peak significance were assessed with R/Limma after a fold-change pre-filter.
    Could also: Count-based negative-binomial frameworks such as DESeq2 or edgeR for RNA-seq, and tools like DiffBind/csaw for ChIP-seq, could also have been used. — These model count data and dispersion directly and integrate independent filtering with FDR control, which some analysts prefer for small replicate numbers in sequencing experiments.
  • Genome-wide analyses used two biological replicates for ChIP-seq and three for RNA-seq.
    Could also: Additional biological replicates per condition could also have been included. — More replicates can increase the precision of dispersion estimates and the statistical power to detect smaller effects, which is one reason higher replication is often recommended for genomics.
  • Targeted qPCR comparisons between two groups were evaluated with Student's t-test.
    Could also: A Welch's t-test or a non-parametric Mann-Whitney U test could also have been applied. — Welch's variant relaxes the equal-variance assumption and rank-based tests avoid normality assumptions, both of which are commonly chosen when sample sizes are small.
  • Many qPCR-style comparisons were each tested individually with t-tests.
    Could also: When several related comparisons are made, an ANOVA with a post-hoc correction (e.g., Tukey HSD) or a multiplicity adjustment (e.g., Benjamini-Hochberg) across the family could also be used. — A unified model with correction controls the family-wise or false-discovery error rate across the set of related comparisons, which some prefer when reporting multiple tests together.
  • Results were largely summarized as fold changes and counts of significant features with threshold-based p-values.
    Could also: Reporting exact p-values together with effect-size estimates and 95% confidence intervals could also have been presented. — Confidence intervals and exact values convey the magnitude and precision of effects in addition to significance, which is increasingly encouraged in reporting guidelines.
  • FDR thresholds were applied after an initial fold-change pre-filter for both genomic datasets.
    Could also: An independent-filtering or shrinkage-based ranking that combines statistical significance and effect size within one model could also have been used. — Integrated approaches can improve the calibration of the FDR and the stability of effect-size estimates, which is why they are a common alternative to sequential filtering.
Software: R package Limma · MACS 2.0.1 · Bowtie 1.0.0 · Sickle · PeakSplitter · CHANCE · CEAS · TopHat 2.0.12 · Cufflinks (Cuffmerge/Cuffquant/Cuffnorm) 2.2.1 · DAVID v6.7 · GREAT 3.0.0 · MEME-ChIP / TomTom / FIMO · IGV 2.3.36 · ABI Sequence Detection Software (SDS) V2.0.5 · CPAT and CPC

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
71
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GO:0019827 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
also used by 1 paper:
GO:0050870 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
also used by 1 paper:
GO:0002064 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0006282 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0010033 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0010468 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0016070 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0031124 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0031323 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0032204 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0032526 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0033044 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0033993 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0043583 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0048387 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0048863 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0051301 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0060008 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — PMID 27655631 (atrazine transgenerational RNA-seq + ChIP-seq)

Title: Exposure to the widely used herbicide atrazine results in deregulation of global tissue-specific RNA transcription in the third generation and is associated with a global decrease of histone trimethylation in mice — Hao et al., NAR 2016. PMID 27655631 · PMC5175363 · DOI 10.1093/nar/gkw840.

Cited code artifact: https://github.com/najoshi/sickle (third-party FASTQ quality trimmer; per P16 a valid reproducible artifact). NOTE: in this paper sickle is used in the ChIP-seq pipeline (-q33 + Bowtie 1.0.0), NOT in the RNA-seq pipeline.

Data: GEO GSE81093 (SuperSeries). Subseries:

  • GSE81091 = RNA-seq (SRP074350 / PRJNA320479) — 18 runs, all paired-end 100bp, HiSeq 2500. F3 only: 3 tissues (testis/brain/liver) x 2 (control/atrazine) x 3 reps. Runs SRR3475391–SRR3475408 (~60–78M read pairs each, ~130 GB FASTQ total).
  • GSE81056 / GSE84978 = ChIP-seq H3K4me3 (F1+F3 testis).

In scope (pipeline-derived, attempted) — RNA-seq, the 80/20 core

Pipeline as described in Methods "RNA-Seq expression data processing": FastQC QC → TopHat 2.0.12 map to Ensembl mm9 → BAM → Cufflinks assemble → Cuffmerge vs Ensembl mm9 annotation → Cuffquant + Cuffnorm (Cufflinks 2.2.1) expression levels → custom DE filter: keep transcripts > 50th quantile of all values in ≥1 condition, then > 2-fold ATZ-vs-control difference, then Limma with FDR < 5%.

Primary claims to regenerate (see claims.tsv C1–C5):

  • C2 1419 total DE transcripts (FC>2, FDR<0.05); split C3 testis 1322, C4 liver 69, C5 brain 28; C1 704 collapsed genes.

These are the clearly-specified, headline numeric outputs → primary target.

Partially in scope (secondary, attempt if primary lands)

  • C6–C7 lncRNA/LincRNA fraction (needs extra CPC + CPAT coding-potential step; CPC/CPAT thresholds given: negative CPC + CPAT prob <40%).
  • C8–C10 alternative-isoform / APA transcript counts (Cuffcompare class codes; APA step not fully specified).

Out of scope (not attempted, stated why)

  • ChIP-seq H3K4me3 peak calling & differential peaks (sickle + Bowtie1 + MACS2 + CHANCE). Separate heavy pipeline; RNA-seq DEG counts are the paper's headline and the cleaner 80/20. May revisit if time permits since sickle is the cited artifact.
  • All wet-lab / qPCR / IGV-visualization / motif (TomTom) / external-dataset overlap claims (157 ATZ-overlap, H4K5ac/H4K8ac comparisons, vinclozolin comparison) — these depend on external published datasets + manual steps, not regenerable from GSE81093 alone.
  • F1 results: no RNA-seq for F1 (RNA-seq is F3-only); F1 only appears in ChIP-seq.

Known reproduction risks (flag for auditor)

  1. TopHat2/Cufflinks is deprecated and assembly is non-deterministic (Cuffmerge novel transcript IDs vary run-to-run) → exact 1419 is unlikely; expect within-tol/partial. Target: recover the tissue pattern (testis ≫ liver ≈ brain) and order of magnitude.
  2. The DE step is underspecified: applying Limma after a hard fold-change pre-filter on Cuffnorm FPKM is statistically unusual; the exact design matrix / contrast / whether FPKM is log-transformed is not stated → docs_insufficient risk for exact counts.
  3. Software versions: TopHat 2.0.12 + Cufflinks 2.2.1 pinned in paper; Bowtie/TopHat index for Ensembl mm9 must be built.

Status

Eligible. Data fully public, pipeline tools all open-source & conda-installable, expected numeric result pinned (1419 / 1322 / 69 / 28 / 704). Heavy compute → «our HPC».

Figures / tables: Fig 1AFig S8TableFig S13Fig S14AFig 2A
C2
Reported
1419 DE transcripts (FC>2, FDR<0.05), F3 ATZ vs control
Reproduced
416 DE transcripts
did not match
C3
Reported
1322 DE transcripts in testes
Reproduced
217
did not match
C4
Reported
69 DE transcripts in liver
Reproduced
128
did not match
C5
Reported
28 DE transcripts in brain
Reproduced
71
did not match
C1
Reported
704 DEGs (genes) total
Reproduced
397
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 10/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +10

The RNA-seq pipeline (TopHat2->Cufflinks->Limma) was still running at an operator-requested early finalize, so no DEG counts were computed and all primary claims (1419 transcripts, 704 genes, 1322/69/28 per tissue) remain pending — no reproduced value was asserted, so there is no fabrication. The data is fully public and 1:1 available (GSE81093/SRP074350, 18 F3 runs), but the authors' DE step is underspecified (Limma after a hard fold-change FPKM pre-filter), which would add uncertainty even on completion. The dominant cause of 'no result' is our incomplete run, not an authors' defect or a measured discrepancy. Overall this is a partial/pending reproduction: sound, auditable and resumable, but with no comparison achieved — hence uniformly yellow rather than a green pass or a red fabrication/discrepancy.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

283.1 k
tokens (I/O) · 17.7 M incl. cache
238 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.