The impact of haplotypes derived from Chinese pigs on genetic variation and economic traits in the Duroc breed.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
- Nothing in this column.
- 🔴Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🔴Overall, the reproduction showed a material discrepancy
▸Reproduction agent’s raw note
DROP. This is a re-analysis paper (Genet Sel Evol 2025) combining FIVE previously-published public datasets with ~13 third-party popgen/quant-gen tools; it ships NO analysis code. The room's auto-harvested code link (github.com/ENCODE-DCC/chip-seq-pipeline2) is a LINK-EXTRACTION FALSE POSITIVE - ENCODE ChIP-seq pipeline is unrelated to this pig population-genomics study and produces none of its results. Described-well-enough? Tool names+versions are given, but with no shipped code, no sample->group sheets, no reference-panel definitions and incomplete per-step parameters, the specific reported numbers are not reproducible 1:1 within the 80/20 budget. We selected the cleanest low-hanging target - the directional FST claim (Fig S15) on the small open SNP-chip dataset (Dryad doi:10.5061/dryad.30tk6, 990 ind/50,705 SNP) using the paper's own named tool VCFtools - scripted it, and VERIFIED the toolchain on «our HPC» (PLINK 1.9, VCFtools 0.1.17 = exact paper match). It could NOT run because Dryad now serves downloads behind an AWS WAF JavaScript challenge (HTTP 401 on API routes; HTTP 202 + awsWafCookieDomainList/gokuProps interstitial on file_stream) that non-browser clients cannot solve; the data is openly licensed but practically un-fetchable by automation, and cannot be routed to «infra». NOT attempted (out of 80/20 scope): all crisp numbers from the large Chinese-server resequencing sets (CNCB GVM000479 578-pig; GigaDB 100894 3,056-pig) requiring RFMix local-ancestry, Selscan, GCTA fastGWA, FastQTL, coloc and SMR. NO positive evidence of fabrication - reported values are plausible/field-consistent; this is a reproducibility & data-access gap. Human reviewer must confirm grades (all provisional).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-14 ⛓ c2a509d5c0ef
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe paper investigates the genetic and biological implications of historical introgression of Chinese pig haplotypes into the European Duroc breed, asking what genomic components derive from Chinese ancestry, how strongly they were selected, and how they affect economically important traits and gene expression.
- ★ Significant genetic introgression from Chinese pigs into commercial European lines was confirmed, with introgressed segments predominantly deriving from Southern Chinese domestic pigs (CSDP) with additional contributions from Eastern Chinese domestic pigs (CEDP). finding
- ★ Selection pressure for Chinese pig introgression was stronger in Duroc pigs than in Large White and Landrace breeds. finding
- ★ GWAS based on ancestral CEDP/CSDP haplotypes identified 10 QTLs, five of which were not detected in previous studies or by SNP-based analyses. finding
- ★ eGWAS based on introgressed haplotypes in duodenum, liver, and muscle revealed signals enriched near transcript start sites. finding
- ★ A region ~300 Kb from TAF11, enriched with open chromatin and containing a super-enhancer in the same TAD as TAF11, is associated with both TAF11 expression and loin muscle depth. mechanism
- ★ An integrative framework combining GWAS, eGWAS, Hi-C, ATAC-seq, and ChIP-seq with co-localization was used to dissect how introgressed loci influence muscle traits. method
- Haplotype blocks were recoded by ancestral origin (European vs CEDP/CSDP) via local ancestry inference to enable haplotype-based association testing. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| SNP chip genotyping / population genetics (D-statistic, FST, admixture) | 990 pigs (Chinese indigenous, ECP, EDP, EWB, outgroup species) | none | introgression signals, allele sharing, genetic distance | SNP chip (50,705 SNPs); Dsuite v0.5, VCFtools v0.1.17, ALDER v1.03 |
| Whole-genome resequencing / global & local ancestry inference | 937 pigs worldwide + 14 warthogs; HD dataset of 578 pigs (16,549,697 variants) | none | introgressed segment ancestry (CEDP/CSDP origin) and frequency | GTX FPGA accelerator, Beagle v5.4, bcftools v1.21, RFMix v2.03 |
| Low-coverage genome sequencing GWAS | 3,056 Duroc pigs (LCS dataset, 7,436,569 SNPs) | none | QTLs for BF, TN, LMD, LMA, litter size | Beagle v5.4 imputation |
| Expression GWAS (eGWAS) with RNA-seq | 100 Duroc pigs; muscle, liver, duodenum tissues | none | gene expression (TPM) vs introgressed haplotype associations | Trimmomatic v0.39, Bowtie2, HISAT2 v2.2.1, StringTie v2.1.2 |
| ChIP-seq (H3K27ac) / super-enhancer identification | muscle of 2-month-old Duroc, Large White, Meishan, Enshi pigs | none | H3K27ac peaks and super-enhancer regions | ENCODE ChIP-seq pipeline v2.2.2, BWA mem v0.7.17, Homer v4.11 |
| ATAC-seq | muscle tissue of Meishan, Duroc, Large White, Enshi pigs | none | open chromatin regions / narrow peaks | GEO GSE143288 processed bigwig/bed files |
| Hi-C / chromatin architecture (TAD, loop calling) | muscle tissue of a Large White pig | none | topologically associating domains and chromatin loops | Fastp, BWA mem v0.7.17, pairtools v1.0.2, cooler v0.9.1, juicertools v1.22.01 |
| Selective sweep analysis (XP-EHH, iHS, FST) | HD dataset of 578 pigs (Duroc/Large White/Landrace vs EWB) | none | positive selection signals in introgressed regions | Selscan v1.2.0a, VCFtools v0.1.17 |
- – Introgressed segments predominantly derive from Southern Chinese domestic pigs (CSDP) with additional CEDP contributions
- ▲ Selection pressure for Chinese introgression stronger in Duroc than Large White and Landrace
- – GWAS on ancestral haplotypes identified 10 QTLs, 5 novel relative to prior/SNP-based studies 10 QTLs (5 novel)
- – eGWAS signals from introgressed haplotypes enriched near transcript start sites
- – Region ~300 Kb from TAF11 (with super-enhancer in same TAD) associated with both TAF11 expression and loin muscle depth ~300 Kb
- count 10 QTLs identified (5 novel) (GWAS based on ancestral CEDP/CSDP haplotypes)
- count 143 CSDP, 71 CEDP (SNP chip Chinese samples) (Chinese pig populations in SNP chip dataset)
- count 990 individuals with 50,705 SNPs (SNP chip dataset retained for analysis)
- count 578 pigs with 16,549,697 autosomal variants (final HD resequencing dataset)
- count 3,056 pigs with 7,436,569 SNPs (final LCS dataset for GWAS (incl. 2802 Duroc))
- count 97 muscle, 97 duodenum, 92 liver samples (12,143 / 14,586 / 13,065 TPMs) (transcriptomic matrices for eGWAS)
- other Z-score > 2 and p-value < 0.05 (threshold for significant gene flow in D-statistic test)
- count 7,437,797 SNPs (100 Duroc pigs) (Duroc resequencing for eGWAS after QC)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This population genomics study used SNP chip data (~990 individuals, 50,705 SNPs) and whole-genome resequencing data (up to 3,056 pigs, 7–16 million SNPs) to trace Chinese pig introgression into European commercial pig breeds, with emphasis on Duroc. Global introgression was quantified with the D-statistic (ABBA-BABA); local ancestry was inferred with RFMix; selective sweeps were characterised with XP-EHH, iHS, and windowed FST. GWAS was performed on low-coverage Duroc resequencing data paired with phenotypes, and eGWAS on RNA-seq from duodenum, liver, and muscle; integrative Hi-C, ATAC-seq, and ChIP-seq analyses supported mechanistic interpretation of a key locus. The provided text excerpt ends before the GWAS and eGWAS statistical methods are fully described.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| D-statistic (ABBA-BABA test); Dsuite v0.5 r52; threshold Z-score > 2 and p < 0.05 | Global introgression detection between each European commercial pig breed (P2) and each Chinese pig population (P3), with EWB or EDP as P1 and warthog/outgroup species as O | 990 individuals (SNP chip dataset); 578 pigs (HD resequencing dataset) | not stated |
| Weir and Cockerham weighted FST in 50 Kb sliding windows (step 25 Kb); VCFtools v0.1.17 | Pairwise genetic distance between Chinese (CEDP, CSDP) and European (EDP, ECP, EWB) groups; also used to validate selective sweep signals in Duroc, Large White, and Landrace vs. EWB | not stated per pairwise comparison | not stated |
| Duncan's multiple range test | Post-hoc comparison of average windowed FST values across all Chinese-European group pairs to test whether genetic distances differed significantly among pairs | — | not stated |
| XP-EHH (cross-population extended haplotype homozygosity); Selscan v1.2.0a; normalised in 50 Kb bins; threshold: minimum XP-EHH > 2 within bin | Selective sweep signals in three pairwise comparisons: Duroc vs. EWB, Large White vs. EWB, Landrace vs. EWB | 578 pigs (HD dataset) | not stated |
| iHS (integrated haplotype score); Selscan v1.2.0a; normalised in 50 Kb bins; threshold: top 1% normalised iHS | Within-breed selective sweep signals for Duroc, Large White, and Landrace; used jointly with XP-EHH to define positively selected regions | 578 pigs (HD dataset) | not stated |
| Chi-square test | Whether the enrichment of positively selected bins within CEDP- or CSDP-introgressed bins differed significantly among Duroc, Landrace, and Large White | — | not stated |
| Pearson correlation coefficient | Correlation of Chinese introgression haplotype frequencies between European commercial pig breeds and European crossbreed lines (text excerpt ends mid-sentence; full application not described) | — | not stated |
| Kullback-Leibler divergence (KLD) | Distributional comparison of Chinese introgression haplotype frequencies (text excerpt cut off; full application not described) | — | na |
| GLM smooth (ggplot2 'glm' method); R/ggplot2 v3.5.2 | Fitting introgression frequency data from European commercial pigs (DLY and WDU crossbreed lines) | — | not stated |
| ALDER admixture time inference (LD-decay method); ALDER v1.03 | Estimation of admixture time between ECP and Chinese pig groups (CEDP, CMDP, CSWDP, CSDP); one generation = 5 years | — | not stated |
| GWAS (specific statistical model not described in provided text excerpt) | Association of Chinese-introgressed haplotypes with backfat thickness, teat number, loin muscle depth, loin muscle area, and litter size in Duroc pigs; LCS dataset | 2802 Duroc pigs (LCS dataset) | not stated |
| eGWAS (specific statistical model not described in provided text excerpt) | Association of Chinese-introgressed haplotypes with gene expression in muscle, duodenum, and liver tissue in Duroc pigs; HD resequencing + RNA-seq | 97 muscle samples, 97 duodenum samples, 92 liver samples (100 Duroc pigs) | not stated |
-
Global introgression was assessed pairwise with the D-statistic (ABBA-BABA) in a four-taxon framework using Dsuite↳ Could also: f-branch statistics (fbranch, available in Dsuite), TreeMix, or ADMIXTURE could also quantify and visualise admixture proportions from multiple donor populations simultaneously — f-branch explicitly partitions admixture across a species tree and can resolve competing donor contributions in a single model, which may be informative given that the paper identifies contributions from both CSDP and CEDP; TreeMix additionally produces a graphical representation of migration edges across populations
-
Selective sweeps were identified using empirical genome-wide top-1% and XP-EHH > 2 thresholds for iHS and XP-EHH respectively↳ Could also: Composite likelihood ratio tests (e.g., SweepFinder2, SweeD) or the nSL (number of segregating sites by length) statistic could also be applied — CLR methods model the expected site-frequency spectrum under a hard sweep and can distinguish selection from demographic effects; nSL is less sensitive to demographic history than iHS and can detect incomplete sweeps; using complementary statistics helps assess whether sweep signals are consistent across methods
-
Pairwise genetic distances among multiple Chinese-European group combinations were compared using Duncan's multiple range test on average windowed FST values↳ Could also: Tukey's HSD or a permutation-based multiple comparison procedure could also be used following ANOVA of windowed FST values — Tukey's HSD controls the family-wise error rate for all pairwise contrasts at the same nominal alpha level and is more widely reported in the genetics literature, facilitating comparison across studies; permutation approaches avoid the assumption of normality for windowed FST distributions
-
Haplotype block ancestry was assigned by thresholding RFMix posterior probabilities at 0.5 and encoding each block as European (0) or Chinese (1)↳ Could also: A continuous posterior probability or a higher confidence threshold (e.g., > 0.9) could also be used, or ELAI (efficient local ancestry inference) as an alternative tool — Using a continuous ancestry dosage rather than a binary assignment retains uncertainty information and can increase power in downstream GWAS; a higher threshold reduces misclassification at the cost of excluding ambiguous segments; ELAI does not require a predefined reference panel structure and can handle multi-way admixture natively
-
Admixture timing between European commercial pigs and Chinese pig groups was estimated with ALDER using decay of admixture-induced linkage disequilibrium↳ Could also: MALDER (multi-wave ALDER) or demographic inference with fastsimcoal2 / SMC++ could also be used — MALDER extends ALDER to detect and date multiple admixture pulses within a single analysis, which may be relevant given centuries of repeated importation of Chinese pigs into Europe; fastsimcoal2 and SMC++ can jointly infer population size histories and admixture timing under an explicit coalescent model
-
Pearson correlation coefficients were used to relate Chinese introgression haplotype frequencies across European commercial breeds and crossbreed lines↳ Could also: Spearman rank correlation could also be used, or a linear mixed model that accounts for population structure among the breeds being compared — Spearman rank correlation makes no distributional assumptions and is more robust to skewed genome-wide frequency distributions and outlier windows; a linear mixed model with a genetic relatedness matrix as a random effect can account for the non-independence of breeds sharing recent common ancestry
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
7 downstream papers · 2 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Loss of Monoallelic Expression of <i>IGF2</i> in the... 2022 · 7 cites
- DeepSATA: A Deep Learning-Based Sequence Analyzer In... 2023 · 5 cites
- Constructing eRNA-mediated gene regulatory networks... 2024 · 5 cites
- DPImpute: A Genotype Imputation Framework for Ultra-... 2025 · 4 cites
- Lewis x-carrying O-glycans are candidate modulators... 2023 · 0 cites
- Identification of transcriptional regulatory variant... 2022 · 16 cites
- Mutations on a conserved distal enhancer in the porc... 2023 · 3 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
scope.md — pmid-41131451
Paper: Fang S et al. (2025) The impact of haplotypes derived from Chinese pigs on genetic variation and economic traits in the Duroc breed. Genet Sel Evol. PMID 41131451 · PMCID PMC12551222 · DOI 10.1186/s12711-025-01010-z.
This is a re-analysis / meta-genomics study: it combines five previously published public datasets and analyses them with ~13 standard population- and quantitative-genomics tools. It introduces no new wet-lab data of its own ("All data and materials were collected from previously published studies").
Code artifact — IMPORTANT (screening false positive)
- The auto-harvested code link in the room brief is
github.com/ENCODE-DCC/chip-seq-pipeline2. This is a link-extraction false positive. ENCODE's ChIP-seq pipeline has nothing to do with a pig population-genomics / GWAS study. The repo is real and active, but is not the paper's code and is not used to derive any result here. - The PMC full text has no Code-availability statement and no GitHub link. → The paper ships no analysis code. (Per P16 a third-party tool on the paper's data is still a valid reproduction, so this alone is not a drop.)
Datasets the paper reuses
| # | data | accession / host | access | size |
|---|---|---|---|---|
| D1 | SNP-chip (Yang 2017 global pigs), lifted to Sscrofa11.1 by Wang et al.; 990 ind / 50,705 SNP after QC | Dryad doi:10.5061/dryad.30tk6 (PLINK ped/map) | open license, but downloads now behind AWS WAF JS-challenge → not programmatically fetchable | ped.zip 80 MB |
| D2 | HD resequencing, 578 pigs / 16,549,697 variants | CNCB GVM000479 (China) | public, large | very large |
| D3 | low-coverage resequencing, 3,056 pigs / 7,436,569 SNP | GigaDB 100894 | public, large | very large |
| D4 | transcriptomics (97 muscle/92 liver/96 duodenum) | ENA PRJEB58030 | open (FTP) | large |
| D5 | ATAC-seq / ChIP-seq | GEO GSE143288 | open | medium |
In-scope (pipeline-derived) results, by feasibility
| result | tool(s) named | dataset | feasibility |
|---|---|---|---|
| FST between pop groups, 50 kb windows; directional claim Chinese–EDP > Chinese–EWB (Fig S15) | VCFtools (Weir & Cockerham) | D1 | cleanest target, small data — BUT D1 is AWS-WAF-gated |
| genome-wide CSDP/CEDP haplotype freq 17.9–20.7% / 2.9–3.1% (Fig 2a) | RFMix local ancestry | D2 | needs large data + reference panels (under-specified) |
| introgression timing 27.8–48.9 / 26.1–38.0 gen (Fig S2) | ALDER | D1/D2 | needs D1 (gated) + LD decay params |
| selective-sweep region counts (100/49/130; Table S9) | Selscan + windowing | D2 | large data, threshold params under-specified |
| 10 GWAS QTLs, FDR<0.05 (Table S10–11) | GCTA fastGWA mixed model | D3 | large data + phenotypes + haplotype dosages |
| eGWAS bin counts (Table S12) | FastQTL | D2+D4 | needs genotypes (gated) + expression |
| coloc/SMR 73 candidate SNPs (Table S13–15) | coloc, SMR | D2+D3+D4 | end of a long chain |
Out of scope
- Functional/biological interpretation, GO enrichment narrative, manual curation.
- Anything depending on the Chinese-server HD/low-coverage data at full scale (D2, D3) — large data + multi-tool pipelines whose reference panels, sample→ group lists and exact thresholds are not given → exceeds the 80/20 budget.
Chosen attempt and outcome
Target = the directional FST claim on D1 (smallest, openly licensed, named
tool VCFtools/PLINK). Toolchain (PLINK 1.9, VCFtools 0.1.17, Python/pandas) was
built and verified on «our HPC». The attempt was blocked at data fetch: the
Dryad file endpoints return HTTP 401 (API) / an AWS WAF JavaScript challenge
(awsWafCookieDomainList / gokuProps) on the public file_stream route, which a
non-browser client cannot solve. See AUDIT.md for the full trace.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.