Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genomic prediction based on selective linkage disequilibrium pruning of low-coverage whole-genome sequence variants in a pure Duroc population.

Genet Sel Evol · 2023
L1 No data access 2/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
No data access Data access not granted

This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.

Reproduction agent’s raw note

DROP (data_unavailable). The paper (Zhu et al. 2023, GSE, Duroc SLDP genomic prediction) is described well enough in prose to understand, but is NOT reproducible at bounded cost. (1) The processed genotypes + phenotypes needed to reproduce ANY reported value are in GigaDB doi:10.5524/100894, whose CNGB single-page-app endpoints do not serve files: /api/dataset/?identifier=100894 -> {"data":{}}, /api/file/ -> HTTP500, FTP mirrors (parrot.genomics.cn, ftp.cngb.org) time out / 404 -- verified from BOTH «host» and a «our HPC» compute node with confirmed internet («job», 2179398). (2) The public SRA accessions resolve but hold only the ~97-sample high-coverage REFERENCE panel (PRJNA681437=58 runs ~2.08TB; PRJNA712489=39 runs ~1.0TB), not the 3,549-animal ~0.68x lcWGS genotypes or the AGE/BF/TTN phenotypes. (3) The authors' analysis/SLDP/prediction pipeline code is NOT shipped -- the only linked repo (github.com/xiaolei-lab/SIMER) is a trait-SIMULATION R package, so the headline accuracies would require re-implementing a multi-TB lcWGS->STITCH-imputation->GCTA-GWAS->PLINK-LD-pruning(180 combos)->GBLUP/BayesR cross-validation pipeline from the prose (HARD RULE 3 out-of-scope tail). NOT attempted: the end-to-end pipeline (out of scope) and the cheap panel-SNP/LD-pruning counts (blocked on data access). SIMER is runnable but produces only simulated phenotypes, no reported value to grade 1:1. No value was fabricated to avoid the drop.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-15 ⛓ 460c4fac31bb
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can selectively pruning low-coverage whole-genome sequence SNPs by linkage disequilibrium using GWAS prior information (SLDP) improve the accuracy of genomic prediction for complex traits in a pure Duroc pig population, and for which trait genetic architectures does this benefit hold?

Core claims
  • Selective linkage disequilibrium pruning (SLDP) refines whole-genome SNP sets using GWAS prior information to improve genomic prediction accuracy. method
  • SLDP improves prediction accuracy mainly for traits controlled by major QTL or a small number of QTN, but offers no significant advantage for traits without major QTL. finding
  • SLDP improved real-trait prediction accuracy by 0.84 to 3.22% (traits with major/moderate QTL) over commercial SNP chips, GBS data, and an unselected whole-genome SNP panel. finding
  • SLDP improved simulated-trait prediction accuracy by 1.23 to 11.47% for traits controlled by a small number of QTN. finding
  • SLDP performance depends on the genetic architecture of the trait and the reliability of the GWAS prior information. finding
  • SLDP can be incorporated into mainstream prediction models (GBLUP and BayesR). method
  • Low-coverage sequencing genotyping yields millions of highly accurate SNPs in pigs usable for genomic prediction. resource
  • Massive noisy markers not in LD with causal variants in WGS data may explain why WGS data did not improve prediction accuracy in prior studies. mechanism
Experimental setups
Assay System Perturbation Readout Platform
Low-coverage whole-genome sequencing (LCS) with imputation 3579 Duroc boars (pigs), one nucleus farm none autosomal SNP genotypes (15,689,585 called; 10,109,688 after QC) Illumina paired-end 150 / BGI paired-end 100; BaseVar v1.01, STITCH v2.0; ~0.68x depth
Genome-wide association study (GWAS) 1000-individual Duroc discovery population none significant SNPs / QTL association P-values for marker selection GCTA mlma model v1.92.0
Genomic prediction (GBLUP) 2108 Duroc training, 441 validation SLDP/LDP marker selection GEBV accuracy (Pearson correlation) and bias
Genomic prediction (BayesR) 2108 Duroc training, 441 validation SLDP/LDP marker selection GEBV accuracy and bias BayesR
Phenotype/trait simulation real LCS 10.1M genotypes (Duroc) varying QTN number (100 or 10,000) and heritability (0.15/0.30/0.50) simulated additive genetic + residual phenotypes / TBV R package SIMER
In silico SNP panel extraction (commercial chip and GBS simulation) Duroc LCS clean panel none panel SNP counts (CC 80k, CC 50k, GBS) for genotyping technology comparison VCFtools v1.17
Key results
  • SLDP improved prediction accuracy for real traits controlled by major or moderate QTL versus SNP chips, GBS, and unselected WGS panel 0.84-3.22%
  • SLDP improved prediction accuracy for simulated traits controlled by a small number of QTN 1.23-11.47%
  • SLDP performed better for traits controlled by major QTL or few QTN, with no significant advantage for traits lacking major QTL (e.g., infinitesimal trait AGE)
  • TTN is mainly affected by several major QTL on SSC7; BF is a transitional trait with major loci plus minor genes; AGE is infinitesimal with no major QTL
  • After quality control, 3549 of 3579 pigs and 10,109,688 SNPs were retained
Key statistics
  • count 15,689,585 autosomal SNPs identified; 10,109,688 SNPs retained after QC (SNP calling and QC in study population)
  • count 3579 Duroc boars sampled; 3549 retained after QC (population size)
  • count 180 combinations of two core parameters (GWAS P-value thresholds and LD r2) (SLDP parameter testing)
  • other P-value threshold gradient 0.0001 to 0.01 (GWAS marker selection thresholds)
  • other average sequencing depth ~0.68x (low-coverage sequencing depth)
  • count 55,216 (CC 80k), 36,851 (CC 50k), 94,832 (GBS) SNPs (simulated genotyping panel sizes)
  • other r2 >= 0.50 informative SNP / QTL-covered region of 50 kb either side (informative SNP and false-positive definitions)
  • count Discovery 1000, Training 2108, Validation 441 individuals (population split for AGE/TTN)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This genomic prediction study in 3579 Duroc boars proposes SLDP (selective linkage disequilibrium pruning), a marker pre-selection method integrating GWAS prior information with LD-based variant filtering to improve prediction accuracy from low-coverage whole-genome sequence data. Animals were partitioned into non-overlapping discovery (n=1000), training (n=2108), and validation (n=441, born ≥2014) populations; GWAS was conducted in the discovery set via a mixed linear model (GCTA mlma), and 180 parameter combinations (9 GWAS P-value thresholds × 20 LD r² thresholds) were optimized by five-fold cross-validation in the training set. Two genomic prediction models (GBLUP and BayesR) were then evaluated in the held-out validation population across three real traits and six simulated traits, with accuracy measured as Pearson correlation between GEBV and phenotype or true breeding value, and bias as the regression coefficient of phenotype/TBV on GEBV.

Replicationbiological Sample size3579 total Duroc boars partitioned into discovery (n=1000, randomly selected from non-validation animals), training (n=2108, remaining non-validation animals), and validation (n=441, born ≥2014); no formal a priori power analysis stated GroupsSix SNP panels (SLDP, LDP, CC 80k, CC 50k, GBS, full WGS 10.1 M panel) × two prediction models (GBLUP, BayesR) × three real traits and six simulated traits (2 QTN counts × 3 heritability levels) Pairingna Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Additive mixed linear model GWAS (GCTA mlma) with genomic relationship matrix as polygenic background control Discovery population: SNP association analysis for three real traits (AGE, BF, TTN) and six simulated traits 1000 (discovery population) not stated
Five-fold cross-validation with Pearson correlation as the selection criterion Training population: optimization of 180 SLDP parameter combinations (GWAS P-value threshold × LD r²) and LDP parameters 2108 (training population) not stated
GBLUP (genomic best linear unbiased prediction) Genomic prediction for all real and simulated traits across all SNP panels 441 (validation population) not stated
BayesR Genomic prediction for all real and simulated traits across all SNP panels 441 (validation population) not stated
Pearson correlation (GEBV vs. phenotype or GEBV vs. true breeding value) Primary accuracy metric in validation population for all models and marker panels 441 (validation population) na
Linear regression coefficient (phenotype or TBV regressed on GEBV) Bias assessment of all genomic prediction models 441 (validation population) na
Approaches that could also have been used
  • Panel and model comparisons were based on point-estimate Pearson correlations in a single temporally defined holdout set (n=441), reported as percentage improvements
    Could also: Bootstrap confidence intervals or permutation tests for differences in Pearson correlation between panels/models could also be computed — A single held-out set yields one realization of accuracy without uncertainty quantification; intervals or tests on accuracy differences would indicate whether observed improvements are distinguishable from sampling variation, especially for differences in the 0.84–3.22% range
  • Prediction accuracy was measured as Pearson correlation between GEBV and observed phenotype or true breeding value
    Could also: Rank-based (Spearman) correlation or concordance correlation coefficient (CCC) could also quantify predictive performance — Spearman correlation is less sensitive to outliers and violations of normality, and CCC jointly captures systematic bias and random error; both are used in the animal genetics literature as complementary accuracy metrics
  • GWAS was conducted with a standard additive mixed linear model (GCTA mlma) using a single GRM constructed from all LCS SNPs as the polygenic background
    Could also: Leave-one-chromosome-out (LOCO) GWAS (as implemented in BOLT-LMM or fastGWA-LOCO) could also be used — LOCO excludes the tested chromosome from the background GRM, avoiding proximal contamination that can reduce power for large-effect loci; this is potentially relevant given the major QTL reported for BF and TTN on specific chromosomes
  • GWAS P-value thresholds were used as marker-selection hyperparameters without a stated genome-wide multiple-testing correction
    Could also: FDR-based thresholds (e.g., Benjamini-Hochberg q-value) or stepwise conditional analysis to identify independent signals could also frame the selection step — FDR control provides a statistical framework for the expected proportion of false discoveries among selected markers, which could help interpret differences in selected marker counts and false-positive rates across traits with contrasting genetic architectures
  • Simulated traits modeled only additive genetic and residual effects, with QTN effects sampled from a standard normal distribution
    Could also: Simulations could also include gamma-distributed or Laplace-distributed effect sizes to represent oligogenic architectures with heavy tails, as well as dominance or epistatic terms — Real quantitative traits often exhibit skewed effect-size distributions and non-additive components; varying the effect-size distribution would allow broader generalization of the simulation conclusions to the real traits analyzed
  • Validation used a single temporal split (animals born before vs. ≥2014), approximating one generation of forward prediction
    Could also: Rolling-origin or multiple sequential temporal splits could also be used to estimate accuracy across several generation intervals — A single temporal cutoff provides one realization of forward-validation accuracy; multiple splits at different birth-year boundaries would yield a more stable accuracy estimate and reveal whether performance is consistent across different generational gaps between training and validation
Software: GCTA 1.92.0 · BaseVar 1.01 · STITCH 2.0 · VCFtools 1.17 · R/SIMER

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
19
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37853325

Paper: Zhu et al. 2023, Genet Sel Evol 55:72. "Genomic prediction based on selective linkage disequilibrium pruning (SLDP) of low-coverage whole-genome sequence variants in a pure Duroc population." DOI 10.1186/s12711-023-00843-w. PMCID PMC10583454.

Linked code: https://github.com/xiaolei-lab/SIMER (R package, trait simulation only). Linked data: SRA PRJNA681437 + PRJNA712489; processed data GigaDB 10.5524/100894.

The reported computational results (candidate claims)

id result reported location
panel-snp SNP counts per panel after QC: CC50k=36,851; CC80k=55,216; GBS=94,832; LCS=10,109,688 counts Methods / Table
acc-real GBLUP prediction accuracy: AGE≈0.229 (CC80k), BF≈0.377 (GBS), TTN≈0.356 (GBS); SLDP gains +1.19–3.22% accuracies Tables 5–6, Fig 5
acc-sim SLDP gain in simulated traits: +6.93–11.47% (100-QTN, GBLUP) vs LCS; ~0% for 10,000-QTN gains Fig 7

In scope vs out of scope (pipeline-derived)

Pipeline behind the results (narrative Methods, no shipped script): lcWGS reads (~0.68×, 3,549 boars) → BaseVar v1.01 + STITCH v2.0 variant calling/imputation → ~10.1M SNPs → VCFtools QC → GCTA v1.92 GWAS → PLINK sliding-window LD pruning (win 500, step 200; 180 r²×P combos for SLDP) → GBLUP (MTG2) / BayesR cross-validation (discovery 1,000 / training 2,108 / validation 441) → prediction accuracy.

  • OUT OF SCOPE (the "hard 80%", not attempted): the end-to-end prediction pipeline. Reasons: (a) the authors' own analysis/SLDP pipeline code is not shipped — only SIMER (trait simulator) is linked, so the prediction results would have to be re-implemented from the prose; (b) inputs are multi-TB lcWGS + a high-coverage reference panel; STITCH imputation of 3,549 samples to ~10M SNPs is weeks of compute; (c) phenotypes for AGE/BF/TTN are not in the public SRA records.
  • The only runnable shipped artifact: SIMER (third-party trait simulator), used in the paper to make the 6 simulated-trait scenarios. SIMER can be run, but it produces simulated phenotypes, not any number the paper reports as a standalone value — the paper's simulated-trait claims are prediction accuracies that still need the full out-of-scope pipeline.
  • Cheapest in-scope target if data is obtainable: panel-snp — count SNPs in a published genotype panel, and/or run PLINK LD pruning to reproduce a pruned-SNP count. This needs the processed genotype files (GigaDB). Whether those are downloadable at bounded cost is being checked on a «our HPC» compute node («job»).

Data-availability findings (control-plane checks, 2026-06-15)

  • ENA PRJNA681437 = 58 high-coverage runs (~2.08 TB) — NOT the lcWGS set.
  • ENA PRJNA712489 = 39 high-coverage runs (~1.0 TB) — also high-coverage reference samples.
  • => The public SRA holds only the ~97-sample high-coverage reference panel, not the 3,549-animal lcWGS genotypes or the phenotypes.
  • GigaDB 10.5524/100894 is a CNGB single-page app; its /api/dataset/ returns empty and /api/file/ 500s from «host»; FTP mirrors 404/550 from «host». Re-checking from a compute node before concluding.
Figures / tables: TablesFig 5Fig 7
panel-snp
Reported
CC50k=36,851; CC80k=55,216; GBS=94,832; LCS=10,109,688
Reproduced
not attempted (genotype files not obtainable)
partial
acc-real
Reported
AGE=0.229; BF=0.377; TTN=0.356; SLDP +1.19-3.22%
Reproduced
not attempted (data unavailable + pipeline not shipped + out-of-scope scale)
partial
acc-sim
Reported
SLDP +6.93-11.47% (100-QTN); ~0% (10,000-QTN)
Reproduced
not attempted (data unavailable + pipeline not shipped)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 31/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

Honest drop (data_unavailable), not a mismatch. Every reported value — panel SNP counts (LCS=10,109,688), real-trait accuracies (AGE=0.229, BF=0.377, TTN=0.356) and SLDP gains (+1.19–3.22% real, +6.93–11.47% simulated) — depends on processed genotypes/phenotypes deposited at GigaDB doi:10.5524/100894, whose endpoints serve no files, while public SRA holds only the ~97-sample reference panel. The blocker is on the data-availability and code-sharing side (the authors' prediction pipeline is not shipped; only the SIMER simulator is), verified independently from a «our HPC» node. No deviation, discrepancy, or fabrication was observed — the claims are simply untested, so this is criticality-yellow rather than a critical red.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

113.1 k
tokens (I/O) · 5.2 M incl. cache
15 min
runtime · 0 CPU-h
0 GB
peak RAM
2
HPC jobs
hummel
machine