Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genomic regions and candidate genes selected during the breeding of rice in Vietnam.

Evol Appl · 2022
L1 75/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
75/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 45% of all assessed papers rank 612 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction via strategy P16 (third-party deposited tool applied to the paper's own pipeline + data). Three SLURM jobs ran on «our HPC» compute nodes. REPRODUCED: (M1, exact) the deposited code XP-CLR v1.1.2 (commit dc47063) executes the selection scan with the paper's EXACT parameters (100kb window / 10kb step / max 200 SNPs/window) and emits well-formed XP-CLR scores; (M2, exact) the deposit PRJEB36631 contains exactly 616 Illumina WGS paired Oryza sativa samples = the paper's stated 616 (the prior room's '1024 runs' was wrong), ~0.81 TB, fastq download + parse cleanly; (R3/R4, method-demo) the ENTIRE named upstream toolchain (BWA-MEM -> IRGSP-1.0 -> FreeBayes min-cov10 -> VCFtools biallelic/mac3/Q30/<50%miss -> XP-CLR -> Weir-Cockerham Fst) runs end-to-end on 6 real PRJEB36631 samples (chr1), producing 680,207 raw variants -> 122,780 biallelic SNPs -> 4327 XP-CLR windows -> Fst. NOT REPRODUCED (honest blockers, NOT fabrication): the headline full-scale numbers (C1: 2.03M/1.13M SNPs; C2: Table 4 region/gene counts; C3: Fst 0.185/0.305, Table 7 candidate genes). These need (a) the full 672-genome pipeline at scale and (b) the per-sample I1-I5/J1-J4 subpopulation labels, which are an UNDEPOSITED derived product -- so every selection comparison and the headline Fst are not regenerable from the deposit alone. The demo's arbitrary grouping yields low XP-CLR (max 0.568) and Fst (0.014/0.035), which correctly illustrates that the published signals depend on the real, missing subpop labels. KEY AUDITABILITY FINDING: method fully described + tool public + raw data complete, but the grouping metadata linking raw reads to the published analysis is the missing reproducibility link (flagged for human audit).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ 4da3fb43d61a
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-26
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

By leveraging the strong population structure among 672 native Vietnamese rice genomes, genomic regions and candidate genes that were selected during rice breeding in Vietnam can be identified, particularly in the outlying Indica-5 (I5) subpopulation, which represents an untapped gene pool for breeding.

Core claims
  • XP-CLR and FST scans identify genomic regions with distorted allele frequency/differentiation patterns resulting from differential selective pressures between Vietnamese rice subpopulations finding
  • The I5 subpopulation shows the highest proportion of the genome under selection (52 regions, 8.1% of the genome, 4576 genes) finding
  • 65 candidate genes in I5 selected regions were identified as promising breeding targets, several harboring nonsynonymous substitutions finding
  • XP-CLR (cross-population composite likelihood ratio test) models multilocus allele frequency differentiation to detect selective sweeps, including both hard and soft sweeps method
  • Selected regions overlap with previously identified QTLs from GWAS studies on the same rice diversity panel (root development, panicle architecture, drought tolerance, leaf development, jasmonate regulation, phosphate starvation) finding
  • GO term enrichment (topGO, weight01 algorithm) identifies over-represented functional categories among genes in selected regions method
  • Genome scanning approach is applicable for identifying novel loci and alleles to breed a new generation of sustainable and resilient rice finding
  • Japonica subpopulations have fewer, longer selected regions concentrated in specific genome areas (e.g., chr2 long arm, chr4 centromeric flanks), while Indica selected regions are more spread and variable in length finding
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome resequencing and SNP calling 672 Vietnamese native rice accessions (616 newly sequenced + 56 from 3K Rice Genomes Project), mapped to Nipponbare IRGSP-1.0 reference none bi-allelic SNPs (6.3M and 6.8M sets merged to 4.4M, filtered to 3.8M; MAF-filtered to 2,027,294 SNPs in 426 Indica and 1,125,716 SNPs in 211 Japonica samples) BWA-MEM, Picard, SAMtools v1.5, FreeBayes v1.0.2, VCFtools v0.1.13, Beagle v4.1
XP-CLR selective sweep scan 5 Indica and 4 Japonica Vietnamese rice subpopulations none (comparative population genomics) XP-CLR scores per 100kbp sliding window (10kbp step) used to define selected genomic regions XP-CLR (updated version of Chen et al. 2010 code, https://github.com/hardingnj/xpclr); BEDTools v2.26.0
FST calculation (Weir & Cockerham method) I5 subpopulation (43 samples) vs I2, I3, I4 subpopulations (190 samples) none Per-SNP and per-100kbp sliding window (10kbp step) FST values, mean FST per gene/region VCFtools (weir-fst-pop option)
SNP functional effect annotation Whole-genome SNP set, Rice Genome Annotation Project release 7.0 none Low/medium/high putative functional effect classification of SNPs, nonsynonymous substitutions SnpEff
GO term enrichment analysis Genes within selected genomic regions (Rice MSU7.0 annotation) none Over-represented GO terms (Fisher test, FDR<0.05) topGO (R package, weight01 algorithm), agriGO
Protein domain/pathway enrichment analysis Genes within selected regions none Enriched protein domains and pathways (Bonferroni FDR<0.05) PhytoMine (Phytozome)
QTL overlap annotation 672-accession diversity panel and 182-accession GWAS subset none Overlap between selected regions and previously mapped QTLs (root development, panicle architecture, drought tolerance, leaf development, jasmonate regulation, phosphate starvation) BEDTools map
Key results
  • I5 subpopulation had the highest whole-genome mean XP-CLR selection scores among Indica subpopulations except versus I3 63.6 (I5 vs I1)
  • I5 subpopulation had 52 selected regions covering 8.1% of the rice genome, the highest among all nine subpopulations 8.1% genome, 4576 genes
  • Indica subpopulations had a higher proportion of genome under selection than Japonica subpopulations Indica 5.3%-8.1% vs Japonica 3.7%-4.9%
  • J4 subpopulation had the highest selection scores among Japonica, especially against J1 46.1 (J4 vs J1)
  • I1 subpopulation (elite cultivars) had consistently the lowest selection scores among Indica subpopulations scores 4.0-9.8
  • 65 candidate genes selected as promising breeding targets from I5 selected regions, several with nonsynonymous substitutions
  • Selected regions in I5 overlapped with other landrace subpopulations (I2, I3, I4) on chromosome 1 short arm and chromosome 9 long arm, but were absent on chromosome 4 long arm where other landraces overlapped
Key statistics
  • count 672 rice genomes (616 newly sequenced + 56 from 3K RGP) (total sequenced sample set)
  • count 3.8 million SNPs after heterozygosity filtering (final filtered SNP set before imputation)
  • count 2,027,294 SNPs (Indica) and 1,125,716 SNPs (Japonica) (MAF 5%-filtered SNP sets per subtype)
  • fold_change mean XP-CLR score I5 vs I1 = 63.6 (highest pairwise Indica selection score)
  • other 52 selected regions, mean length 583,706 bp, total length 30,352,734 bp (I5 subpopulation selected regions)
  • other 8.1% of genome (373,245,519 bp reference) (proportion of genome under selection in I5)
  • count 4576 genes in I5 selected regions (genes annotated within I5 selected regions)
  • count 65 candidate genes selected as breeding targets (final candidate gene set from I5 selected regions)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a population-genomics selection-scan design comparing nine Vietnamese rice subpopulations (five Indica, four Japonica) pairwise across the genome, rather than a classical experimental replicate design. Selective sweeps were identified with XP-CLR applied to pairwise subpopulation comparisons, using an empirical cut-off (mean of the 99th percentile of scores) to define candidate regions, and cross-checked with FST (Weir & Cockerham estimator) between the I5 subpopulation and the pooled I2-I4 landraces. Gene sets within selected regions were further tested for Gene Ontology and protein-domain/pathway enrichment using Fisher's exact tests with FDR (topGO) or Bonferroni-FDR (PhytoMine) correction. Results were reported mainly as mean scores, percentile-based cut-offs, and percentages of the genome under selection, rather than as classical hypothesis-test p-values.

Replicationunclear Sample sizeSample sizes given as accession counts per subpopulation (Table 1), e.g., 672 total samples (426 Indica, 211 Japonica), with I5 n=43; no formal power calculation described GroupsNine rice subpopulations (Indica I1-I5, Japonica J1-J4) compared pairwise for selection signatures, with I5 as the focal outlier group Pairingna Randomization/blindingna Dispersionmixed Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionFDR (topGO) and Bonferroni-FDR (PhytoMine) for enrichment tests; empirical 99th-percentile threshold (not a p-value correction) used for XP-CLR/FST selection cut-offs
Statistical tests used
Test Applied to n Assumptions
XP-CLR (cross-population composite likelihood ratio test) Pairwise comparisons among the five Indica subpopulations and among the four Japonica subpopulations, genome-wide in 100 kbp windows (Tables 3-4, Figure 1) 672 total samples (426 Indica, 211 Japonica); subpopulation sizes given in Table 1 (e.g., I5 n=43) not stated
FST (Weir and Cockerham 1984 estimator, via VCFtools weir-fst-pop) I5 subpopulation vs. pooled I2, I3, I4 subpopulations, per-SNP and over 100 kbp sliding windows 43 samples in I5 vs. 190 samples in I2+I3+I4 not stated
Fisher's exact test (topGO 'weight01' algorithm), FDR<0.05 GO term enrichment among genes within each selected region vs. whole transcriptome gene lists per selected region (exact n not stated) not stated
Protein domain/pathway enrichment test, Bonferroni FDR<0.05 Genes within selected regions, analyzed in PhytoMine/Phytozome not stated
Approaches that could also have been used
  • Selective sweeps were detected using XP-CLR, a likelihood-based method modeling multilocus allele-frequency differentiation between population pairs.
    Could also: Haplotype-based statistics such as iHS or XP-EHH, or additional allele-frequency methods like Tajima's D or Bayesian outlier detection (e.g., BayeScan) — These would provide complementary evidence on sweep type (e.g., ongoing vs. completed sweeps) and could cross-validate regions identified by XP-CLR, similar to how the paper notes other studies (e.g., He et al. 2017) combined several selection-detection approaches.
  • Cut-off thresholds for putative selected regions were defined using the mean of the 99th percentile of XP-CLR/FST scores across pairwise comparisons, an empirical percentile-based approach.
    Could also: A coalescent- or demographic-simulation-based null distribution reflecting the population's history (e.g., bottlenecks, migration) — Simulation-based nulls can help distinguish selection signals from demographic effects that also distort allele frequencies, complementing the percentile-based cut-off used here.
  • FST was calculated using the Weir and Cockerham (1984) estimator between I5 and the pooled I2+I3+I4 subpopulations.
    Could also: The Hudson FST estimator, or separate pairwise comparisons (I5 vs. I2, I5 vs. I3, I5 vs. I4) instead of pooling — The Hudson estimator can behave differently under unequal sample sizes, and unpooled pairwise comparisons could show whether the differentiation signal is driven by a specific subpopulation contrast.
  • GO term enrichment in selected regions was tested using Fisher's exact test (topGO 'weight01') with an FDR cut-off of 0.05.
    Could also: A rank-based gene set enrichment method (e.g., GSEA) using continuous selection scores rather than a binary in/out-of-region gene list — Rank-based enrichment does not depend on a hard region boundary and can also capture signal from genes just below the significance cut-off.
  • Sample sizes per subpopulation comparison (e.g., 43 I5 vs. 190 I2/I3/I4 samples) reflected available sequenced accessions rather than a predetermined target.
    Could also: A post hoc power assessment or simulation study of detection power given the observed sample sizes — This could clarify how well-powered site-frequency-based methods like XP-CLR and FST were to detect selection in smaller subpopulations (e.g., J3, n=17).
  • Selected regions required support from at least three (Indica) or two (Japonica) pairwise comparisons before being merged into a final subpopulation-level set.
    Could also: A formal combination of per-comparison evidence, such as Fisher's combined probability test or Stouffer's method, across the pairwise comparisons — Such approaches yield a single combined statistic and significance level across comparisons, offering a more quantitative alternative to a vote-count threshold for requiring convergent support.
Software: XP-CLR (github.com/hardingnj/xpclr) · VCFtools 0.1.13 · BWA-MEM · Picard tools 1.128 · SAMtools 1.5 · FreeBayes 1.0.2 · BCFtools isec 1.3.1 · Beagle 4.1 · SnpEff · BEDTools 2.26.0 · topGO (R) · ggplot2 (R) · PhytoMine/Phytozome

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 35899250 (Vietnamese rice breeding selection scan)

Title: Genomic regions and candidate genes selected during the breeding of rice in Vietnam. Higgins et al., Evolutionary Applications 2022. DOI 10.1111/eva.13433 · PMCID PMC9309459. Code: https://github.com/hardingnj/xpclr (XP-CLR; pinned v1.1.2 = dc470638161a01593da3d545ddec3b7ab6c7baf9). Data: ENA study PRJEB36631 (raw Illumina WGS reads only).

The paper's pipeline (from Methods)

  1. Align: 616 Vietnamese WGS samples → BWA-MEM (default) → IRGSP-1.0 (Nipponbare) reference. Picard v1.128 dedup; SAMtools v1.5 merge.
  2. Call SNPs: FreeBayes v1.0.2 --min-coverage 10. VCFtools v0.1.13 filter (bi-allelic, min allele count 3, QUAL>30, <50% missing) → 6.3 M SNPs.
  3. Merge with 56 3K-RGP Vietnamese samples → 4.4 M SNPs (present in both, ≥70% samples) → heterozygosity cutoff 0.591 → 3.8 M SNPs → Beagle v4.1 imputation. MAF filter → 2,027,294 SNPs (426 Indica), 1,125,716 SNPs (211 Japonica).
  4. Subpopulations: 9 clusters — Indica I1–I5, Japonica J1–J4 (Table 1).
  5. XP-CLR selection scan (the deposited code): 100 kb windows, 10 kb step, max 200 SNPs/window; per-subpop cutoff = mean 99th percentile of pairwise XP-CLR; regions <80 kb dropped; ±200 kb of centromeres removed; merged with BEDTools v2.26.0 (max gap 100 kb); kept if in ≥2 (Japonica) / ≥3 (Indica) comparisons → Table 4.
  6. F_ST: VCFtools --weir-fst-pop (Weir & Cockerham), per-SNP + 100 kb/10 kb windows.
  7. Downstream: SnpEff (RGAP r7.0) effects; topGO GO enrichment; QTL overlap (external sets).

IN SCOPE (pipeline-derived, attempted)

id result pipeline feasibility
R1 XP-CLR tool executes the selection scan with the paper's exact parameters and produces XP-CLR scores xpclr v1.1.2 (deposited code) ✅ feasible — fixture + small VCF + scaled real data
R2 Dataset profiling of PRJEB36631 (N runs/samples, data type, QC) ENA + samtools/seqkit ✅ feasible
R3 Scaled end-to-end pipeline: subset of ENA samples → BWA-MEM (IRGSP-1.0, one chr) → FreeBayes → VCFtools filter → XP-CLR full named stack ⚠️ method-demo only (not full-scale)
R4 F_ST method: VCFtools --weir-fst-pop on the scaled VCF VCFtools ⚠️ method-demo only

PARTIAL / NOT FULLY REPRODUCIBLE AT SCALE (honest blockers)

id reported value blocker
C1 2,027,294 (Indica) / 1,125,716 (Japonica) final SNPs needs full 672-genome BWA+FreeBayes+Beagle on ~1.3 TB raw reads — hundreds of CPU-h, not run at full scale
C2 Table 4 selected regions: I5 = 52 regions, 8.1% genome, 4576 genes (and 8 other subpops) needs full genotype VCF AND the I1–I5/J1–J4 subpopulation labels — labels are NOT deposited and ENA sample aliases are opaque LIBxxxxx IDs → grouping is unrecoverable
C3 F_ST 0.185→0.305; 65 candidate genes (Table 7); 14 QTL overlaps same blocker as C2 (needs labelled full VCF + external QTL sets)

OUT OF SCOPE (not a reproducible pipeline output from the deposit)

  • Subpopulation clustering itself (ADMIXTURE/PCA-derived I1–I5/J1–J4) — authors' derived labels, not deposited.
  • topGO GO enrichment, QTL-overlap curation — downstream of C2/C3 candidate genes + external QTL databases.
  • Manual candidate-gene prioritisation / phenotype interpretation — wet-lab / expert curation.

KEY AUDITABILITY FINDING

The deposit PRJEB36631 contains only raw sequencing reads — not the processed genotype VCFs, the per-sample subpopulation assignments, or the selected-region BED files. The headline numerical results (Table 4 region/gene counts, Table 7 candidate genes, F_ST 0.185/0.305) therefore cannot be regenerated from the deposited data alone, even with unlimited compute, because the I1–I5/J1–J4 subpopulation labels — on which every selection comparison depends — are an undeposited derived product. This is recorded as a rep

Figures / tables: Table
M1
Reported
XP-CLR runs the selection scan with paper params (size 100kb/step 10kb/maxsnps 200) and produces XP-CLR scores
Reproduced
xpclr v1.1.2 (commit dc47063) ran with those EXACT params -> 9 windows, 13-col output incl xpclr & xpclr_norm; nSNPs<=155<=200 cap honoured
exact
M2
Reported
616 Vietnamese WGS samples deposited (PRJEB36631)
Reproduced
616 ENA runs = 616 distinct samples (all ILLUMINA/WGS/PAIRED/Oryza sativa), ~0.81 TB, 5.15 B reads; 2 runs downloaded+parsed cleanly
exact
R3
Reported
full named pipeline (BWA->FreeBayes->VCFtools->XP-CLR/Fst) on real reads (method demo)
Reproduced
6 real samples -> IRGSP-1.0 chr1 -> 680,207 raw variants -> 122,780 biallelic SNPs -> XP-CLR 4327 windows -> Fst 0.0136/0.0350
partial
R4
Reported
Weir & Cockerham Fst (VCFtools --weir-fst-pop)
Reproduced
mean Fst 0.0136 / weighted 0.0350 over 99,483 SNPs on the demo VCF (arbitrary groups, NOT comparable to paper's 0.185/0.305)
partial
C1
Reported
2,027,294 Indica / 1,125,716 Japonica final MAF-filtered SNPs
Reproduced
not attempted at full scale (~0.81 TB raw, hundreds of CPU-h, full 672-genome BWA+FreeBayes+Beagle)
m.public.grade.not-attempted
C2
Reported
Table 4 selected regions/genes (e.g. I5: 52 regions, 8.1% genome, 4576 genes)
Reproduced
not recoverable: requires the per-sample I1-I5/J1-J4 subpopulation labels, which are NOT deposited (undeposited derived product)
m.public.grade.not-attempted
C3
Reported
Fst 0.185->0.305; 65 candidate genes (Table 7); 14 QTL overlaps
Reproduced
not recoverable: needs labelled full VCF + selected-region BEDs + external QTL sets
m.public.grade.not-attempted

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 75/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is an in-progress / partial reproduction of a whole-genome selection-scan paper. The deposit (PRJEB36631) contains only raw reads, not the processed VCFs or the I1–I5/J1–J4 subpopulation labels on which every Table 4 count (e.g. I5: 52 regions, 8.1%, 4576 genes), Table 7 candidate genes, and F_ST value (0.185/0.305) depend, so those headline numbers cannot be regenerated from the deposit alone. The deviation sits on the input/data-availability side (an undeposited derived product — authors'/deposit-completeness gap), and the full 672-genome pipeline was never run at scale, so the central conclusion was neither confirmed nor contradicted. There is no fabrication signal and no demonstrated numeric discrepancy — the limitation is honest reproducibility incompleteness, which keeps overall criticality at yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

84 k
tokens (I/O) · 4.8 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.