Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Diapause vs. reproductive programs: transcriptional phenotypes in a keystone copepod.

Commun Biol · 2021
L1 57/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
57/100
Reproducibility score
1.0 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 17% of all assessed papers rank 965 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced PMID 33782539 (Lenz et al. 2021, Commun Biol; copepod diapause-vs-reproductive transcriptomes) on «our HPC» SLURM compute nodes via two routes. R1 (downstream pipeline on the authors' Dryad count matrix, the recommended P16 route): (1) expressed-transcript filter reproduces within 0.07% (27,851 vs 27,870); (2) t-SNE+DBSCAN ordination reproduces EXACTLY (3 clusters: field EF+LF together, EC, LC; stable over 5 seeds); (3) dominant WGCNA turquoise module reproduces EXACTLY (5689) but secondary modules diverge AND the authors' own deposited module map is internally inconsistent with the paper; (4) DEG identification does NOT reproduce as described -- literal Methods (glmLRT+BH<=0.05) give 23,290 (~2x reported 11,503); the deposited DEGs are 99.6% a subset of mine but no standard edgeR variant (QLF, logFC cutoffs, FDR to 1e-4, global BH) recovers their specific set => under-specified DE step (reproducibility flag, NOT fabrication). R2 (Bowtie2 re-map of 3 raw-read libraries): per-transcript counts correlate Spearman 0.989-0.991 with exactly one deposited column each, confirming the count matrix is faithful to the reads (so the DEG gap is purely analytic) and de-anonymizing those samples. Incidental: the TSA reference GAXK01 holds 206,012 transcripts (paper's 96K is an undescribed subset); Dryad downloads are behind an Anubis bot-wall (bypassed via the presigned version-bulk endpoint). Overall: described-well-enough for the data deposit and for ordination/network/mapping, but the headline DEG numbers are not reproducible as written -> PARTIAL (mix of exact/within-tol and mismatch). Not attempted: GO-enrichment narrative, mtCOI check, biomarker interpretation (out of scope).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 41
    assessed: 2026-06-18 ⛓ de4dd5495d9f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether Calanus finmarchicus stage CV copepodids on the reproductive vs. diapause developmental program can be distinguished by their gene expression (transcriptomic) phenotypes, and whether such differences reveal new insights into the physiological basis of diapause preparation.

Core claims
  • t-SNE clustering of all-gene expression data groups field-collected (diapause program) samples into one cluster while early and late culture (reproductive program) samples separate into two distinct phenotypes finding
  • Gene expression changes substantially with maturation in individuals on the reproductive program, whereas individuals on the diapause program show little change over time finding
  • A GO filter comprising all genes annotated to RNA metabolism effectively separates copepods by developmental program, a result confirmed by differential gene expression analysis finding
  • 54 differentially expressed genes annotated to oogenesis, RNA metabolism, and fatty acid biosynthesis are consistently up-regulated in individuals on the diapause program relative to the reproductive program finding
  • A combined computational pipeline (t-SNE/DBSCAN clustering, GLM-based DEG analysis, WGCNA, and GO-term filtering) provides a protocol to reliably distinguish reproductive- from diapause-program individuals without prior knowledge of origin method
  • WGCNA identifies two major co-expression modules: a blue module (glycerophospholipid biosynthesis) positively correlated with the diapause program and a turquoise module (positive regulation of RNA metabolic process) positively correlated with the reproductive program finding
  • Oogenesis-associated gene expression reflects gonad development occurring in reproductive-program but not diapause-program stage CVs finding
  • Only two of six tested GO-based filters effectively separated samples by developmental program finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (read mapping/quantification) Calanus finmarchicus, stage CV copepodids; lab-cultured (reproductive program) and field-collected from Trondheimsfjord (diapause program) none (natural developmental program comparison; early vs. late timepoints) mapped transcript counts / gene expression levels Bowtie2 mapping to Gulf of Maine C. finmarchicus reference transcriptome (96K transcripts)
dimensionality reduction (t-SNE) with DBSCAN clustering same RNA-seq dataset, 16 samples across EF/LF/EC/LC groups none sample clustering by transcriptional phenotype t-SNE (perplexity=5, 50,000 iterations), DBSCAN, Dunn index
differential gene expression analysis (GLM, pairwise likelihood ratio tests) same RNA-seq dataset, four groups (EF, LF, EC, LC) none number of differentially expressed genes (DEGs), up/down regulation EdgeR
weighted gene co-expression network analysis (WGCNA) 11K DEGs from same RNA-seq dataset none co-expression modules and module-trait correlation (eigengene expression) WGCNA
gene ontology enrichment and functional annotation DEGs and reference transcriptome annotations none enriched/over-represented biological processes and GO term distribution TopGO, ReviGO, SwissProt-based annotation
GO-term-filtered t-SNE/DBSCAN analysis subsets of reference transcriptome genes annotated to oogenesis, lipid metabolic process, fatty acid biosynthesis, RNA metabolic process none clustering of samples by filtered gene-set expression t-SNE, DBSCAN
microscopic examination of gonad development/oogenesis cultured and field-collected C. finmarchicus stage CV individuals none presence/absence of gonad development
Key results
  • 16 samples formed 3 t-SNE clusters: all field samples (EF, LF) merged into a single cluster, while early and late culture samples separated into two distinct phenotypes
  • GLM identified 11,503 DEGs among the four treatment groups, with 5509 annotated by GO terms enriched for very long-chain fatty acid and RNA metabolic processes 11,503 DEGs; 5509 annotated
  • Fewest DEGs (1739) found between early and late field samples; most DEGs (10,077) found between late culture and early field samples 1739 to 10,077 DEGs
  • Among four GO-filtered gene sets tested (oogenesis, lipid metabolic process, fatty acid biosynthesis, RNA metabolic process), only the RNA metabolic process filter separated samples into distinct field vs. culture clusters
  • 54 DEGs annotated to oogenesis, RNA metabolism, and fatty acid biosynthesis were consistently up-regulated in diapause-program individuals 54 genes
  • WGCNA blue module (3827 genes) was positively correlated with the diapause program and enriched for glycerophospholipid biosynthesis; turquoise module (5689 genes) showed the opposite pattern and was enriched for positive regulation of RNA metabolic process n=3827 and n=5689 genes
  • Oogenesis GO filter (584 genes) separated the 16 samples into two t-SNE clusters, consistent with gonad development occurring in reproductive- but not diapause-program CVs 584 genes
Key statistics
  • count 11,503 DEGs (GLM comparison across four groups (EF, LF, EC, LC))
  • count EF vs LF: 1739 DEGs (982 up, 757 down) (pairwise likelihood ratio test, early vs. late field)
  • count EC vs LC: 6908 DEGs (3090 up, 3818 down) (pairwise test, early vs. late culture; comparable to prior reported 7470)
  • count 5509 DEGs with GO annotation (out of 11,503 total DEGs)
  • count 584 genes annotated to oogenesis (GO:0048477) (reference transcriptome annotation used for filter analysis)
  • count 54 DEGs (consistently up-regulated in diapause-program individuals (filter 2))
  • count WGCNA module sizes: blue n=3827, turquoise n=5689, yellow n=745, brown n=1133, gray (unassigned) n=109 (DEG module assignment from WGCNA)
  • other t-SNE perplexity=5, 50,000 iterations (dimensionality reduction parameters used across analyses)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study applied a three-strategy workflow to RNA-Seq data from 16 stage-CV copepod samples (four groups: early/late field on diapause program, early/late culture on reproductive program). Strategy 1 used t-SNE dimensionality reduction followed by DBSCAN clustering to separate samples agnostically by transcriptional phenotype. Strategy 2 used an EdgeR generalized linear model (GLM) with downstream pairwise likelihood ratio tests and Benjamini–Hochberg FDR correction to identify differentially expressed genes (DEGs), followed by WGCNA correlation network analysis and TopGO enrichment. Strategy 3 applied GO-term-based filters to retrieve gene subsets for additional t-SNE/DBSCAN analyses, with results visualized as heatmaps and box-and-whiskers plots of module eigengene expression.

Replicationbiological Sample size16 samples total; 4 biological replicates per group (four groups: early field, late field, early culture, late culture); no formal power analysis described GroupsDiapause program (early field [EF], late field [LF]) vs. reproductive program (early culture [EC], late culture [LC]) Pairingunpaired Randomization/blindingnot stated DispersionIQR Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionBenjamini–Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
Generalized linear model (GLM) via EdgeR (negative binomial) Overall identification of DEGs across all four groups (EF, LF, EC, LC); yielded 11,503 DEGs 16 samples (4 groups × 4 biological replicates each) not stated
Pairwise likelihood ratio test Six pairwise group comparisons reported in Table 1 (EF vs LF, EF vs EC, EF vs LC, LF vs EC, LF vs LC, EC vs LC) 16 samples total; 4 replicates per group per pairwise comparison not stated
Benjamini–Hochberg FDR adjustment Applied to p-values from the pairwise likelihood ratio tests; significance threshold p ≤ 0.05 post-adjustment null na
Weighted Gene Correlation Network Analysis (WGCNA) Correlation network analysis on 11,503 DEGs to identify co-expression modules; four major modules identified (blue, turquoise, brown, yellow) 16 samples; 11,503 DEGs not stated
TopGO enrichment analysis GO-term over-representation testing within WGCNA modules and overall DEG set; identified enriched processes per module null not stated
t-SNE + DBSCAN with Dunn index Unsupervised dimensionality reduction and cluster identification applied to full expression data and to GO-filtered gene subsets (Strategies 1 and 3); perplexity = 5, 50,000 iterations; MinPts = 3, Eps maximizing Dunn index 16 samples na
Approaches that could also have been used
  • t-SNE was used for all dimensionality reduction steps (full gene set and GO-filtered subsets), with perplexity fixed at 5 across all analyses
    Could also: UMAP (Uniform Manifold Approximation and Projection) could also be used for dimensionality reduction of RNA-Seq expression data — UMAP tends to better preserve global structure (distances between clusters) in addition to local neighborhood relationships, can be faster on large gene matrices, and is increasingly common in transcriptomic analyses; comparing both can help confirm cluster robustness
  • Unsupervised DBSCAN clustering with Dunn-index-based Eps selection was used to define sample groupings after t-SNE embedding
    Could also: Hierarchical clustering (e.g., Ward linkage on Euclidean or Pearson distances) or k-means clustering could also be used to define and visualize sample groupings — Hierarchical clustering produces a dendrogram that makes inter-sample distances explicit and is commonly paired with heatmaps in transcriptomic studies; k-means with stability analysis (e.g., silhouette scores) provides an alternative data-driven cluster count selection
  • EdgeR GLM with pairwise likelihood ratio tests was used to identify DEGs across four groups
    Could also: DESeq2 with Wald tests (or likelihood ratio tests for multi-factor designs) could also be used for DEG identification from count data — DESeq2 uses a different empirical Bayes shrinkage estimator for dispersion and log-fold-change, and the two pipelines often yield complementary gene lists; reporting the overlap between EdgeR and DESeq2 results is a common way to increase confidence in DEG calls, especially with small n per group
  • GO-filter strategy 3 relied on a priori selection of specific GO terms (oogenesis, lipid metabolic process, fatty acid biosynthesis, RNA metabolic process) drawn from prior biological knowledge
    Could also: Gene Set Enrichment Analysis (GSEA) on a ranked gene list could also be used to test GO term enrichment in a data-driven, threshold-free manner — GSEA considers the full ranked distribution of expression changes rather than a binary DEG cutoff, which can surface enrichment signals in gene sets where no single gene reaches significance; this could complement the a priori filter approach by identifying additional processes without pre-specifying GO terms
  • Module eigengene expression was displayed as box-and-whiskers plots with n = 4 per group, showing median and IQR
    Could also: With n = 4 per group, overlaying individual data points (strip or dot plots) on the box summary could also be used — At n = 4, each individual value is informative; showing all points alongside the summary statistic makes the full data visible and is increasingly recommended by journals for small-sample visualizations
  • FDR correction was applied within the EdgeR GLM framework covering six pairwise comparisons simultaneously
    Could also: A contrast-based approach within a single GLM (e.g., specifying program × time interaction contrasts) could also be used to directly test program-by-time interaction effects rather than deriving them post-hoc from pairwise differences — Explicit interaction contrasts within the GLM would directly estimate and test whether the time effect differs between programs, which is a central biological question of the study, while also providing a single unified model for FDR control
Software: Bowtie2 · EdgeR · WGCNA · TopGO · ReviGO · t-SNE (Rtsne or equivalent) · DBSCAN

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33782539

Paper: Lenz, Roncalli, Cieslak, Tarrant, Castelfranco, Hartline (2021) "Diapause vs. reproductive programs: transcriptional phenotypes in a keystone copepod." Communications Biology 4:392. DOI 10.1038/s42003-021-01946-0. PMCID PMC8007741.

Organism / design: Calanus finmarchicus, copepodid stage CV. 16 pooled RNA-seq libraries, 4 groups × 4 reps: EF (early field / pre-diapause, 28 May 2013), LF (late field, 10 Jun 2013), EC (early culture / reproductive, 3 d post-molt), LC (late culture, 10 d post-molt).

Datasets the paper relies on

Accession Role Content
PRJNA231164 (SRA) RNA-seq reads of this study 16 expression libraries (EF/LF/EC/LC ×4) + 1 "C5 transcriptome" run (SRR1144929, not an expression sample). 50 bp PE HiSeq2000.
PRJNA236528 (SRA) → TSA GAXK00000000.1 reference transcriptome source de-novo developmental-series assembly. Paper maps to a 96,090-transcript reference derived from it.
Dryad 10.5061/dryad.12jm63xw7 processed results Bowtie counts, RPKM, log2[RPKM+1], z-scores (96,089 transcripts × 16), DEG+module map (11,503 DEGs), SwissProt+GO annotation.

Pipeline (Methods)

Bowtie2 (default, v2.1.0) map reads → 96,090-transcript reference → counts → RPKM, log2[RPKM+1] → edgeR GLM + pairwise LRT (BH-adjusted p ≤ 0.05) for DE → t-SNE (Rtsne, perplexity = 5, 50,000 iterations) + DBSCAN (MinPts = 3, eps maximizing Dunn index, clusterCrit) for ordination → WGCNA (signed, soft-power 14, minModuleSize 100) for modules → topGO/ReviGO/AmiGO GOOSE for function.

IN SCOPE (pipeline-derived; we attempt these)

  1. Expressed-transcript filter — ≥1 cpm in ≥4 of 16 samples → reported 27,870 of 96,090.
  2. edgeR pairwise DEGs (BH p ≤ 0.05): EF/LF 1739, EF/EC 7197, EF/LC 10077, LF/EC 7675, LF/LC 9939, EC/LC 6908; union 11,503.
  3. t-SNE + DBSCAN → reported 3 clusters (the two field groups together; EC and LC separate).
  4. WGCNA modules → paper reports 4 major: turquoise 5689, blue 3827, brown 1133, yellow 745.
  5. Cross-check reproduced DEGs / modules against the authors' own deposited cfin_DEGs+moduleMap_log2RPKM.csv (ground-truth per-gene calls).

Two complementary reproduction routes

  • (R1) Downstream from the authors' deposited count matrix (Dryad cfin_96k_BowtieCounts.csv): run edgeR/t-SNE/WGCNA exactly as described → tests whether the headline numbers follow from the deposited data. Primary; executed.
  • (R2) From raw reads (re-map PRJNA231164 with Bowtie2 → counts) → tests the mapping/counting step itself. Harder; gated by the 96K reference subset + «our HPC» storage quota — see AUDIT.

OUT OF SCOPE (not pipeline-reproducible here)

  • Wet-lab culturing, field sampling, RNA extraction, library prep.
  • GO enrichment narrative (topGO/ReviGO/AmiGO GOOSE) — manual/curated steps.
  • mtCOI species-contamination check (manual).
  • Biological interpretation of the 54 "biomarker" diapause transcripts.

Code pointer

Brief lists github.com/jkrijthe/Rtsne (the third-party t-SNE package used for ordination), not an authors' analysis repo. Per P16, reproducing by applying the described tools (edgeR/Rtsne/dbscan/WGCNA) to the paper's own deposited data is the valid route.


Corrections after running (2026-06-24, this room)

  • WGCNA is unsigned (verbatim Methods), not signed as written above.
  • Reference subset: the deposited TSA assembly GAXK01 has 206,012 transcripts; the paper maps to a 96,089/96,090-transcript subset ("96K"; Methods text also says 96,060). The subset-selection criterion is not described.
  • DE step does not reproduce: literal Methods (glmLRT + BH ≤ 0.05, no logFC cutoff) give 23,290 DEGs (~2× the reported 11,503). Authors' 11,503 are 99.6% a subset of mine, but no standard edgeR variant (glmQLF, |logFC|≥1/≥2, FDR down to 1e-4, global BH) recovers their set.
  • Reproduced cleanly: expressed filter (27,8
Figures / tables: Table
expressed_filter
Reported
27,870 transcripts (>=1 cpm in >=4 of 16)
Reproduced
27,851
within tolerance
DEG_union
Reported
11,503 DEGs (union of 6 pairwise glmLRT, BH FDR<=0.05, no logFC cutoff)
Reproduced
23,290 (literal Methods, ~2.0x). Authors' 11,503 are 99.6% a SUBSET; NOT recovered by glmQLF(23214)/|logFC|>=1(18426)/>=2(9820)/stricter FDR to 1e-4/global BH(23218) -> DE step under-specified
did not match
DEG_pairwise
Reported
EF/LF 1739; EF/EC 7197; EF/LC 10077; LF/EC 7675; LF/LC 9939; EC/LC 6908
Reproduced
2465/11445/18730/11735/18443/9586 (same rank order, 1.39-1.86x higher)
did not match
tsne_dbscan_clusters
Reported
3 clusters (EF+LF together, EC and LC separate)
Reproduced
3 clusters, identical membership, stable across 5 seeds
exact
wgcna_modules
Reported
turquoise 5689/blue 3827/brown 1133/yellow 745
Reproduced
turquoise 5689 (EXACT)/blue 2532/brown 1295/yellow 1133. Authors' own deposited map (5648/3846/1009/525) also differs from paper
partial
R2_remap_counts
Reported
deposited Bowtie counts from Bowtie2 (default) mapping to 96K reference
Reproduced
Bowtie2 re-map of 3 libs -> Spearman 0.989-0.991 vs the single matching deposited column each (2nd best 0.86-0.93); self-IDs SRR1138709=E-field-4, SRR1139733=L-field-4, SRR1141100=E-culture-4
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 57/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

Run on the authors' own deposited count matrix, the expressed-transcript filter reproduces near-exactly (27,851 vs 27,870) and the rank order of all six pairwise DEG comparisons matches exactly, so the central claim of distinct diapause-vs-reproductive transcriptional programs holds qualitatively. However, the DEG counts do not reproduce 1:1: the literally-described glmLRT+BH p<=0.05 gives ~2x the reported DEGs (union 23,290 vs 11,503), with the authors' set being a 99.6% subset — pointing to an under-specified stricter threshold on the authors' side, not fabrication. A further authors'-side inconsistency: the deposited WGCNA moduleMap (5648/3846/1009/525) differs from the paper's reported modules (5689/3827/1133/745). Severity is moderate (magnitude/direction preserved) and the reproduction is still partial (t-SNE/WGCNA pending).

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

770.4 k
tokens (I/O) · 53 M incl. cache
126 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.