Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genomic insight into the influence of selection, crossbreeding, and geography on population structure in poultry.

Genet Sel Evol · 2023
L1 100/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> EXACT 1:1. The paper's core methodological contribution is the rIBD (relative IBD) introgression-scan tool (github.com/wzuhou/rIBD_WUR @386b8a24, the authors' own code), which ships its Example as the paper's Drenthe-Fowl-bantam chr6 case study (DrFwB/DB/DrFw) WITH the expected output. Ran the tool with the documented command on «our HPC» (SLURM «job», conda env py3.10/pandas2.3.3/pybedtools0.12/bedtools2.31 on a compute node) over the shipped 1,121,918 IBD segments; produced output is BYTE-IDENTICAL to the shipped expected file (same md5, 3635/3635 windows, zero numeric difference). Computation is deterministic (getopt+pandas, no RNG). NOT attempted (heavy 80%, intermediates not shipped): the upstream WGS pipeline behind paper claims A-E -- 2.30M variants from 136 birds/37 breeds (PRJEB34245), PLINK PCA, ADMIXTURE K=4/6, 387 FLK signals in 299 genes, Beagle5.0/refined-IBD calling; these require TB-scale raw-read alignment + variant calling and the called/phased genotype matrices are not deposited, only raw SRA + the downstream chr6 IBD file. See scope.md. No completeness claim.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 100
    assessed: 2026-06-15 ⛓ a66e70401b89
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

What multiple factors—selection by management type, geographic distribution, phenotypic selection, and crossbreeding (bantamization)—shape the complex population structure of 37 traditional Dutch chicken breeds, and how are these reflected in their genomes?

Core claims
  • Dutch traditional chicken breeds display a complex, admixed, subdivided population structure that broadly matches historical management-based clustering (past-productive, ornamental, country fowl, Lakenvelder). finding
  • Geographic distance contributes to genetic differentiation within historical clusters, while differing management purposes act as a genetic 'barrier' limiting gene flow between clusters. finding
  • Signatures of genetic differentiation (FLK/hapFLK) reveal genomic regions associated with diversifying phenotypic selection between breeds, including dwarf (bantam) size and feather color. finding
  • Crossbreeding to create neo-bantams (bantamization) leaves only a few introgressed bantam-donor genomic segments after backcrossing, traceable via relative identity-by-descent (rIBD). mechanism
  • Combining single-variant FLK and haplotype-based hapFLK while accounting for hierarchical population structure detects signatures of selection between breeds. method
  • Relative IBD (rIBD) comparison of a neo-bantam to its bantam and normal-sized donors provides a genomic perspective on crossbreeding effects. method
Experimental setups
Assay System Perturbation Readout Platform
whole-genome sequencing (PE125) 136 individuals from 37 traditional Dutch chicken breeds none SNP and InDel genotypes after filtering (MAF >1%, call-rate >80%) Illumina HiSeq 3000, 350 bp insert; BWA-MEM v0.7.17, GRCg6a (GCA_000002315.5), Freebayes
principal component analysis (PCA) 37 Dutch chicken breeds (autosomal variants) none variance explained / breed clustering across PC1-PC3 PLINK v1.9; R plotly
phylogenetic analysis (neighbor-joining tree on Reynolds' distance) 37 Dutch chicken breeds none pairwise genetic distance / branch lengths between breeds FigTree v1.4.4
ADMIXTURE ancestry analysis 37 Dutch chicken breeds none ancestry coefficients at K=4 and K=6 ADMIXTURE
isolation-by-distance / Mantel test 28 Dutch breeds with known geographic origin (all breeds and within clusters CL1, CL2, CL3) none correlation between genetic (1-ibs) and geographic (haversine) distance PLINK v1.9; R geosphere; 9999 permutations
identity-by-descent (IBD) haplotype sharing detection phased haplotypes of Dutch breed individuals, within vs between clusters none count and length of shared IBD segments Beagle v5.0; Refined-IBD (LOD >3)
FLK and hapFLK selection scan 37 Dutch breeds, 2.30 million pruned autosomal variants none signatures of selection / genetic differentiation (significant signals at FDR 5%) FLK/hapFLK (K=15 haplotype groups); Ensembl v95 annotation; PANTHER v.11; clusterProfiler
relative IBD (rIBD) crossbreeding case study Drenthe Fowl bantam (DrFwB), Dutch Bantam (DB), Drenthe Fowl (DrFw); plus Groningen Mew bantam comparison (GrMwB/GrMw) crossbreeding/backcrossing (bantamization) normalized IBD sharing (10 kb bins) between neo-bantam and donors Refined-IBD; method of Bosse et al.
Key results
  • PC1 separated past-productive breeds (WelSummer, Barnevelder, North Hollands Blue) and their bantams from other breeds PC1 = 12.63% variance
  • PC2 separated the two true bantam breeds (Dutch Bantam and Eikenburger bantam) from other breeds PC2 = 7.79% variance
  • No significant isolation-by-distance across all Dutch breeds combined mantel r = -0.09, P = 0.99
  • Strong positive isolation-by-distance within past-productive cluster CL1 mantel r = 0.73, P = 1×10^-4
  • Positive isolation-by-distance within country fowl cluster CL3 mantel r = 0.33, P = 1×10^-4
  • Weak positive isolation-by-distance within ornamental cluster CL2 mantel r = 0.13, P = 0.034
  • Haplotypes (IBD segments) were more extensively shared within clusters than between clusters, supporting a management-based genetic barrier
  • FLK test detected significant signals indicating genetic differentiation associated with selected traits (e.g., bantam size, feather color) 387 significant signals in 299 genes (FDR 5%)
Key statistics
  • count 136 individuals from 37 breeds (whole-genome sequenced chickens analyzed)
  • correlation mantel r = 0.73, P = 1×10^-4 (isolation-by-distance within CL1 (past-productive))
  • correlation mantel r = 0.33, P = 1×10^-4 (isolation-by-distance within CL3 (country fowl))
  • correlation mantel r = 0.13, P = 0.034 (isolation-by-distance within CL2 (ornamental))
  • correlation mantel r = -0.09, P = 0.99 (isolation-by-distance across all Dutch breeds)
  • count 387 significant signals in 299 genes (FDR 5%) (FLK genome-wide genetic differentiation)
  • count 2.30 million variants (pruned autosomal variants used for FLK/hapFLK)
  • other PC1 12.63%, PC2 7.79%, PC3 7.35% (variance explained by first three principal components)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This observational population genomics study characterized population structure in 136 whole-genome-sequenced chickens from 37 Dutch traditional breeds. Population structure was assessed through PCA, neighbor-joining phylogeny on Reynolds' pairwise genetic distances, and ADMIXTURE ancestry analysis at selected K values. Isolation-by-distance was evaluated with Mantel tests (9999 permutations) globally and within historical management clusters. Genome-wide signatures of genetic differentiation were identified using FLK and hapFLK tests at 5% FDR, and identity-by-descent segment sharing was used to characterize within- and between-cluster relatedness and to reconstruct crossbreeding history in a neo-bantam case study.

Replicationbiological Sample size136 individuals from 37 breeds stated descriptively; no formal power analysis or sample size justification reported Groups37 Dutch traditional chicken breeds, historically assigned to 4 clusters: CL1 past-productive (6 breeds, 24 individuals), CL2 ornamental (10 breeds, 38 individuals), CL3 country fowl (19 breeds, 65 individuals), CL4 Lakenvelder (2 breeds, 9 individuals) Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionFalse discovery rate (FDR) at 5%; specific algorithm (e.g., Benjamini-Hochberg) not named in text
Statistical tests used
Test Applied to n Assumptions
Principal component analysis (PCA) Population structure visualization across 37 breeds 136 individuals not stated
Neighbor-joining tree on Reynolds' pairwise genetic distances Phylogenetic tree of 37 breeds; distances also used as hierarchical structure input for FLK/hapFLK 136 individuals (breed-level pairwise distances) not stated
ADMIXTURE ancestry estimation Ancestry proportions across 37 breeds at K=4 and K=6 136 individuals not stated
Mantel test with 9999 permutations Isolation-by-distance: genetic (1-IBS) vs. haversine geographic distance; tested across all breeds and within CL1, CL2, and CL3 28 breeds for all-breed test (9 excluded for unknown location); breed subsets within each cluster not stated
FLK test (extended Lewontin-Krakauer, single-variant) Genome-wide scan for signatures of genetic differentiation across 37 breeds 136 individuals; 2.30 million variants after MAF and LD filtering not stated
hapFLK test (haplotype-based, chi-square scaled statistic, K=15 haplotype clusters) Genome-wide scan for haplotype frequency differences across 37 breeds 136 individuals; 2.30 million variants after MAF and LD filtering not stated
Refined-IBD identity-by-descent segment detection and relative IBD frequency (rIBD) Haplotype sharing within and between historical clusters; Drenthe Fowl bantam crossbreeding case study comparing three breeds 136 individuals within-cluster analysis; one individual per breed in case study comparison not stated
Gene ontology enrichment analysis (PANTHER v11, clusterProfiler) Functional annotation of genes overlapping FLK/hapFLK significant signals not stated
Approaches that could also have been used
  • ADMIXTURE was run across several K values and two representative outputs (K=4, K=6) were presented
    Could also: Cross-validation (CV) error could be computed and plotted across a range of K, with the K minimizing CV error formally selected and reported alongside the chosen visualizations — Reporting the CV curve allows readers to assess whether the presented K reflects the statistically preferred solution or was chosen primarily for interpretability; it also conveys uncertainty when the CV curve is flat across K values
  • Four Mantel tests (all breeds, CL1, CL2, CL3) were conducted as separate tests without adjustment for multiplicity across the four comparisons
    Could also: A multiple matrix regression with randomization (MMRR) could incorporate geographic distance and cluster membership as simultaneous predictors in a single permutation model; alternatively, a Bonferroni or sequential Bonferroni correction could be applied across the four Mantel tests — MMRR partitions the unique contributions of geography and management cluster to genetic distance in one model, avoiding the collinearity inherent in running separate within-cluster tests; a multiplicity correction would control the family-wise error rate across the four separate tests
  • Selection signatures were detected using FLK and hapFLK, which identify differentiation relative to the population's neutral phylogenetic structure
    Could also: Cross-population extended haplotype homozygosity (XP-EHH) or the integrated haplotype score (iHS) could complement FLK/hapFLK as they are specifically sensitive to incomplete or recent selective sweeps within particular lineages rather than to among-population allele frequency divergence — Different selection-scan methods have different power profiles depending on sweep age, completeness, and whether selection is shared or lineage-specific; corroboration across methods with distinct assumptions strengthens confidence in identified regions
  • A neighbor-joining tree was fitted to Reynolds' pairwise genetic distances to represent breed relationships
    Could also: A graph-based admixture model such as TreeMix could be fitted, explicitly estimating migration edges alongside a population tree topology; alternatively, IQ-TREE or a similar maximum-likelihood framework could be applied to genome-wide variant data with bootstrap branch support — NJ is computationally efficient but does not model gene flow, which the paper demonstrates is a major feature of these breeds; TreeMix would represent admixture events within the same phylogenetic framework rather than treating them as inference artifacts
  • PCA was applied to all autosomal variants for unsupervised visualization of population structure
    Could also: Discriminant analysis of principal components (DAPC) could also be applied using the historical cluster assignments as a supervised grouping, maximizing between-cluster separation — DAPC may better resolve clusters that overlap in unsupervised PCA space (as observed here for CL2 and CL3) and provides posterior cluster membership probabilities, complementing ADMIXTURE's ancestry proportions
  • Haplotype sharing in the neo-bantam case study was summarized using the rIBD metric comparing two pairwise IBD counts in 10-kb bins
    Could also: Local ancestry inference methods such as RFMix or ELAI could be applied to probabilistically assign each genomic window in the admixed breed to one of the parental populations, producing posterior ancestry proportions across the genome — Local ancestry deconvolution provides position-wise posterior probabilities rather than a relative comparison of segment counts, and can distinguish ancestry tracts from multiple donors simultaneously, which may be informative when more than two parental populations contributed
Software: PLINK V1.9 · Beagle 5.0 · Refined-IBD · ADMIXTURE · R 3.6.1 · R/geosphere · R/plotly · R/clusterProfiler · PANTHER v11 · FigTree V1.4.4 · BWA-MEM V0.7.17 · Freebayes · sambamba V0.6.3 · Sickle

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
13
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GCA_000002315.5 GCA in Methods (http://purl.org/orb/Methods)
also used by 1 paper:

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36670351

Paper: Wu, Bosse, Rochus, Groenen, Crooijmans (2023). Genomic insight into the influence of selection, crossbreeding, and geography on population structure in poultry. Genet Sel Evol. PMID 36670351 · PMCID PMC9854048 · DOI 10.1186/s12711-022-00775-x

Authors' code: https://github.com/wzuhou/rIBD_WUR (commit 386b8a2429d16c73bf3d365053b9be35755fca4c, 2023-09-02) — this IS the authors' own tool for the rIBD analysis (P16 not needed; it is the paper's code).

Raw data: ENA PRJEB34245 (+ PRJEB39725) — whole-genome resequencing, 136 individuals from 37 breeds.


Pipeline-derived results in the paper

# Reported result Pipeline In scope?
A 2.30M variants after LD pruning, from 136 birds / 37 breeds WGS read mapping → variant calling → PLINK LD prune OUT — needs full WGS variant-calling on raw SRA (TB-scale, the heavy 80%); phased VCFs / called variants not shipped
B PCA & genetic-distance structure PLINK v1.9 OUT — requires the called genotype matrix (not shipped)
C ADMIXTURE ancestry at K=4 and K=6 ADMIXTURE OUT — requires the called genotype matrix (not shipped)
D 387 significant FLK signals in 299 genes (FDR 5%) FLK/hapFLK OUT — requires the called/phased genotypes (not shipped)
E Beagle v5.0 phasing (0.02 cM window) + Refined-IBD (window 0.06 cM, length 0.03 cM) → IBD segments Beagle 5.0 + refined-IBD OUT as a full run (needs the called VCF). The IBD-segment OUTPUT for chr6 of the case-study breeds IS shipped as Example/test_chr6.ibd.gz and is used as the input to F.
F rIBD (relative IBD) introgression scan — Drenthe Fowl Bantam case study: normalize IBD sharing Dutch-Bantam→DrFwB vs Drenthe-Fowl→DrFwB to find bantam-introgressed haplotype blocks. Windowed rIBD = nIBD_AB − nIBD_AC. rIBD_pd.py (the paper's tool), windows 20 kb / step 10 kb / method 1 IN — primary target

What we reproduce (IN scope)

Claim F1 (primary): Running the paper's own rIBD_pd.py with the documented example command on the shipped chr6 IBD-segment file reproduces the shipped expected output Example/rIBD_DrFwB_DB_DrFw (144 091 bytes), windowed rIBD + weighted-rIBD per 20 kb/10 kb window for the DrFwB / DB / DrFw trio.

This is a deterministic computation (no RNG; pandas/getopt) on the paper's own shipped data — a faithful 1:1 reproduction of the paper's core methodological contribution (the rIBD introgression-scan metric and tool), for the exact Drenthe-Fowl-bantam case study highlighted in the paper.

Command (per repo README):

python rIBD_pd.py -i Example/test_chr6.ibd -A Example/List.DrFwB \
  -B Example/List.DB -C Example/List.DrFw -o rIBD_DrFwB_DB_DrFw \
  -W 20000 -S 10000 -M 1

(test_chr6.ibd.gz must be gunzip-ed first; the script reads plain text.)

Out of scope (NOT attempted) — and why

A–E above: reproducing the upstream WGS→variant-calling→phasing→IBD pipeline requires downloading and aligning whole-genome resequencing of 136 birds (PRJEB34245), calling/pruning to 2.30M variants, ADMIXTURE, FLK. This is the hard ~80%: the called/phased genotype matrices are not shipped, only the raw SRA, so a 1:1 match of those intermediate numbers is not feasible at proportionate cost. We record them as out-of-scope rather than fabricate. The shipped chr6 IBD file lets us reproduce the downstream rIBD step (F) exactly.

F1
Reported
shipped expected rIBD output Example/rIBD_DrFwB_DB_DrFw (Drenthe-Fowl-bantam chr6 case study, trio DrFwB/DB/DrFw, 20kb window/10kb step, method 1; 3635 windows, md5 3d4a372370d428fdbf4456f98b11c38c)
Reproduced
byte-identical: md5 3d4a372370d428fdbf4456f98b11c38c, 3635/3635 rows, max abs diff rIBD=0.0 and weighted_rIBD=0.0 (cmp -s = IDENTICAL); strongest positive peak chr6:23.05-23.14Mb rIBD~1.83-1.92
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

The single attempted claim (F1, the rIBD introgression scan for the Drenthe-Fowl-bantam chr6 trio) reproduces byte-identically to the authors' shipped expected output (md5 3d4a3723…, 3635/3635 windows, max abs diff 0.0), confirming the chr6:23.05-23.14Mb bantam-introgression peak — clean, deterministic, no fabrication concern. However, this is essentially regenerating the authors' own tool's shipped example, not a test of the paper's substantive data claims (2.30M variants, ADMIXTURE, 387 FLK signals), which were not attempted because the phased/called genotype matrices for PRJEB34245 were never deposited. The deviation on what was tested is zero (our side, technical/expected determinism), but coverage is narrow, so the central conclusion is only partially confirmed and overall quality is solid-but-limited rather than a full 1:1 of the paper.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

98.1 k
tokens (I/O) · 8 M incl. cache
12 min
runtime · 0.05 CPU-h
2.5 GB
peak RAM
1
HPC jobs
hummel
machine