Cost-effectively dissecting the genetic architecture of complex wool traits in rabbits by low-coverage sequencing.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH, but the public deposit does not support 1:1 reproduction of any reported number. Methods are clear and the code artifact (BaseVar) is real and works: I built it from source at a pinned commit and ran it successfully on «our HPC» (C1 self-test = 82 variants on 100 ultra-low-pass NIPT BAMs), and launched it on the paper's own rabbit data PRJNA810279 (C2, BWA-mem->OryCun2.0->BaseVar over the FGF10 window; still indexing the genome at finalization). However NONE of the paper's reported quantitative results are reproducible from public data: (a) every result is COHORT-SCALE -- PRJNA810279 is 642 runs / 7.06 Tbases / ~5.0 TB (verified, matches the paper's '7305 Gb'), and BaseVar+STITCH pool all ~627 samples, so SNP counts (18.5M; 1.74M on Chr11), imputation accuracy and LD/selection cannot be matched from a subset and downloading+aligning the full 5 TB is the hard last 20% we are told not to chase; (b) the GWAS/QTL/heritability headline results (incl. FGF10) ADDITIONALLY require per-animal wool-trait phenotypes (DFW/CVDFW/BW/LFW/LCW/RCW) that are NOT deposited -- the availability statement deposits raw reads only -- making them unverifiable from public data. NOT ATTEMPTED: full-cohort BaseVar/STITCH variant set, imputation-accuracy benchmark, GMAT GWAS, GCTA heritabilities, popgen -- all out of scope for feasibility (cohort/5TB) and/or missing public phenotypes. Verdict = partial: the paper's variant-calling CODE ARTIFACT was reproduced; no reported VALUE could be (this is a transparency/availability gap, flagged for the human auditor, not on this evidence a fabrication finding).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-14 ⛓ 6baa784b29ef
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether low-coverage whole-genome sequencing (LCS) combined with genotype imputation can serve as a cost-effective alternative to array-based or high-depth sequencing genotyping for dissecting the genetic architecture of complex wool traits in Angora rabbits via GWAS.
- ★ BaseVar + STITCH at 1.0X sequencing depth with a sample size >300 achieves the highest genotyping accuracy among tested imputation strategies (genotype concordance >98.8%, genotype accuracy >0.97). finding
- ★ LCS followed by imputation is a cost-effective alternative to genotyping arrays and high-depth sequencing for assessing common variants. finding
- ★ Six QTL were detected for wool traits, explaining 0.4 to 7.5% of phenotypic variation. finding
- ★ FGF10 is identified as a candidate gene associated with fiber growth and fiber diameter in Angora rabbit wool. finding
- ★ GWAS combined with LCS can identify new QTL and candidate genes for quantitative traits. finding
- BaseVar was developed/used to call variants for large-scale low-pass WGS data, avoiding reference-allele bias seen with standard tools (SAMtools/Bcftools, GATK) on low-coverage data. method
- Beagle v5.1 provides significantly faster genotype phasing than Beagle v4.1 with similar imputation accuracy. method
- A conditional GWAS with drop-in-log-P confidence interval estimation was developed to refine QTL localization in a multivariate linear mixed model. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Low-coverage whole-genome sequencing with genotype imputation (BaseVar+STITCH, Bcftools+Beagle4, GATK+Beagle5) | Angora rabbit (Oryctolagus cuniculus), 629 individuals (15 deep-sequenced at 10X for validation) | none (down-sampled coverage: 0.1X-2.0X; sample sizes 100-600) | genotype concordance (GC) and genotype accuracy (GA) vs high-coverage genotypes | DNBSEQ-T7; BaseVar, STITCH, Bcftools, Beagle v4.1/v5.1, GATK4 |
| Multivariate and conditional genome-wide association study (GWAS) | Angora rabbit population (n=629) | none (natural phenotypic variation in wool traits) | QTL detection, phenotypic variance explained, candidate gene mapping | — |
| Principal component analysis (PCA) | Angora rabbit population (n=629) | none | population structure (first five PCs) | GCTA |
| Linkage disequilibrium (LD) decay analysis | Angora rabbit whole genomes | none | r2 decay across genome | PopLDdecay |
| Selection scan (composite likelihood ratio, nucleotide diversity) | Angora rabbit population | domestication/selective breeding (implicit) | CLR and Pi values to detect signatures of selection | SweeD, Vcftools |
| Phylogenetic and ADMIXTURE population structure analysis | 14 domesticated vs 14 wild rabbits | domestication status | maximum-likelihood tree, PCA, Fst, Pi, ancestry proportions (K=1-5) | IQ-TREE2, ADMIXTURE, Vcftools |
| SNP-based heritability estimation (multi-trait mixed model) | Angora rabbit population | none | heritability of wool traits (LFW, DFW, CVDFW, LCW, RCW) and body weight | — |
| Functional enrichment analysis (GO/KEGG) | Candidate genes from selection/GWAS regions | none | enriched gene ontology terms and KEGG pathways (FDR-corrected) | DAVID |
- ▲ BaseVar+STITCH at 1.0X depth with sample size >300 gave the highest genotyping accuracy GC>98.8%, GA>0.97
- – Six QTL detected for wool traits 0.4 to 7.5% of phenotypic variation explained
- – FGF10 gene associated with fiber growth and fiber diameter
- – Average genomic coverage achieved across 627 low-coverage samples 3.84X (range 1.51X-8.03X)
- – Total sequence data generated for the study 7305 gigabases
- – ADMIXTURE analysis of domesticated vs wild rabbits found lowest cross-validation error K=2
- other genotype concordance > 98.8% (BaseVar+STITCH imputation accuracy at 1.0X, sample size>300)
- correlation genotype accuracy > 0.97 (BaseVar+STITCH imputation accuracy at 1.0X, sample size>300)
- other 0.4 to 7.5% phenotypic variation explained (six QTL detected for wool traits)
- mean 3.84X average coverage (range 1.51X-8.03X) (low-coverage sequencing of Angora rabbit samples)
- count 629 rabbits (298 males, 331 females) (study population)
- count 7305 gigabases of genomic sequence data generated (total sequencing output)
- other imputation info score > 0.4; MAF > 0.05; missing rate < 0.1; HWE p-value > 1e-6 (SNP filtering criteria after STITCH imputation)
- count 15 samples deep-sequenced at 10X coverage (used for genotype validation)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study generated low-coverage whole-genome sequencing (LCS) data on 629 Angora rabbits, compared three genotype imputation pipelines (BaseVar+STITCH, Bcftools+Beagle4, GATK+Beagle5) across five sequencing depths and six sample sizes, and validated imputation accuracy against 15 high-coverage (10X) samples using genotype concordance and Pearson-r-based genotype accuracy. Dense imputed SNPs were then used in multivariate linear mixed model GWAS of six wool traits across three time points, with conditional GWAS and drop-in-log-P QTL confidence intervals for fine-mapping. Population genetics analyses (PCA, LD decay, CLR selection scan, FST, nucleotide diversity, ADMIXTURE, and maximum-likelihood phylogeny) characterized the domesticated vs. wild rabbit comparison, and GO/KEGG enrichment of candidate genes was corrected by Benjamini-Hochberg FDR.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Genotype concordance (proportion of imputed genotypes matching high-coverage calls) and genotype accuracy (Pearson r between imputed genotypes/dosages and high-coverage genotypes) | Benchmarking of three imputation pipelines across five sequencing depths (0.1X–2.0X) and six sample sizes (100–600), relative to 15 high-coverage (10X) validation samples | 15 high-coverage validation samples; imputation reference panels ranged from 100 to 600 individuals | not stated |
| Multivariate linear mixed model GWAS (SNP-level association test with genomic relationship matrix K as random effect) | GWAS of five wool traits (LFW, DFW, CVDFW, LCW, RCW) and body weight at up to four time points; three-trait model for wool traits, four-trait model for body weight | 629 Angora rabbits | not stated |
| Conditional GWAS | Independent signal identification within GWAS-significant regions to resolve multiple QTL | 629 Angora rabbits | not stated |
| Drop (Δ) in log-transformed P values method for QTL confidence interval estimation | Defining genomic confidence intervals around the six identified QTL | 629 Angora rabbits | not stated |
| SNP-based heritability estimation via multivariate REML-based linear mixed model | Heritability of wool traits (3-trait model across three time points) and body weight (4-trait model across four time points); fixed effects included sex and rabbit house | 629 Angora rabbits | not stated |
| Principal Component Analysis (PCA, GCTA; first five PCs extracted) | Population structure assessment of all 629 rabbits; PCs used as covariates | 629 rabbits | not stated |
| Composite likelihood ratio (CLR) test (SweeD) and nucleotide diversity (Pi) sliding-window analysis (50 kb window, 10 kb step; top 1% threshold) | Genome-wide selection signature scan in the full domesticated population | 629 rabbits (full population scan) | not stated |
| FST sliding-window analysis and Pi comparison (Vcftools; 50 kb window, 10 kb step; top 5% threshold for highly divergent windows) | Population differentiation between 14 domesticated and 14 wild rabbits | 14 domesticated, 14 wild rabbits | not stated |
| ADMIXTURE with cross-validation (CV) error for K selection (K = 1–5); K = 2 selected as minimum CV error | Ancestry and population structure analysis of 14 domesticated vs. 14 wild rabbits | 28 rabbits | not stated |
| Maximum-likelihood phylogenetic tree (IQ-TREE2) | Phylogenetic relationships among 14 domesticated and 14 wild rabbits | 28 rabbits | not stated |
| Benjamini-Hochberg FDR correction | GO and KEGG pathway gene set enrichment analysis (DAVID) for candidate genes in selection and GWAS regions | — | not stated |
-
Imputation accuracy was summarized as overall genotype concordance (GC) and overall genotype accuracy (GA, Pearson r) collapsed across all sites and all 15 validation samples↳ Could also: Dosage r-squared (DR2) or squared Pearson correlation between imputed dosages and true dosages stratified by minor-allele-frequency (MAF) bin could also be used — MAF-stratified metrics reveal differential accuracy at rare vs. common variants, which are expected to differ substantially under low-coverage imputation; site-aggregated metrics can mask poor performance at low-MAF loci that drive discovery power
-
Selection signatures were identified using empirical top-percentile cutoffs (top 1% CLR and Pi; top 5% FST) across non-overlapping sliding windows↳ Could also: Neutral-simulation-derived or permutation-based genome-wide significance thresholds could also be applied to calibrate the false-positive rate — Percentile thresholds flag a predetermined proportion of windows by design regardless of effect magnitude; simulation- or permutation-derived thresholds provide a calibrated error rate and allow explicit statements about detection power
-
QTL confidence intervals were estimated by the Δ log P drop method around each lead SNP↳ Could also: Bootstrap resampling, Bayesian credible sets derived from approximate Bayes factors, or LD-based credible intervals (r² ≥ 0.1 around lead SNP) could also be used — Bootstrap and Bayesian credible sets incorporate sampling uncertainty more formally; LD-based credible sets are more robust when the p-value landscape around a peak is shallow, which is common in populations where LD decays rapidly
-
Population structure was characterized with ADMIXTURE using minimum CV error to select K (K = 1–5 tested; K = 2 chosen)↳ Could also: The Evanno ΔK criterion alongside CV error, or Bayesian model-selection approaches (e.g., fastSTRUCTURE), could also be used for K selection — Multiple complementary K-selection criteria can corroborate the inferred number of ancestral components; ΔK often emphasizes hierarchical structure differently from CV error, which is informative when domestic-wild divergence is moderate
-
Wool traits at three time points were analyzed jointly in a three-trait multivariate linear mixed model, and body weight at four time points in a four-trait model↳ Could also: A random-regression (repeated-measures) model treating time as a continuous covariate could also be used, or single-trait GWAS per trait-time combination followed by multi-trait meta-analysis (e.g., MTAG) — Random-regression models can capture the genetic trajectory of trait change across ages as a continuous function, potentially identifying loci with age-dependent effects that a fixed-time multivariate model averages across; single-trait approaches also simplify interpretation of individual time-point associations
-
HWE filtering was applied with a uniform p-value threshold (> 1 × 10⁻⁶) across the entire mixed cohort of 629 rabbits of both sexes housed across five facilities↳ Could also: HWE testing could also be stratified by sex or housing group, or omitted in favor of relying on the genomic relationship matrix in the mixed model to handle stratification-induced departures — In structured or recently selected populations, HWE departures can reflect true population stratification rather than genotyping error; stratified or model-based approaches reduce the risk of inadvertently discarding variants carrying genuine biological signal
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36401180
Paper: Wang et al. 2022, Genet Sel Evol 54:75. "Cost-effectively dissecting the genetic architecture of complex wool traits in rabbits by low-coverage sequencing." DOI 10.1186/s12711-022-00766-y. Primary code artifact (per BRIEF): BaseVar — https://github.com/ShujiaHuang/basevar (C++ variant caller for ultra-low-pass WGS; the paper's primary variant caller). Data: SRA PRJNA810279.
Pipeline reported in the paper (Methods)
FastQC → Trimmomatic → BWA-mem (ref GCF_000003625.3 / OryCun2.0) → Picard dedup → BaseVar variant calling (vs bcftools/GATK4 for benchmarking) → STITCH imputation (+ Beagle v5.1) → GMAT GWAS (multivariate LMM) + GCTA SNP-heritability.
Reported pipeline-derived results (candidate claims)
| # | Result | Pipeline | In scope? |
|---|---|---|---|
| R1 | 18,577,154 high-quality imputed SNPs (genome-wide) | BWA+BaseVar+STITCH, 627 samples | NO — cohort-scale |
| R2 | 1,737,601 polymorphic sites on Chr11 (BaseVar) | BaseVar, 627 samples | NO — cohort-scale |
| R3 | SNP functional split 57.78% intergenic / 35.50% intronic / 0.52% exonic; 72,552 syn / 23,328 nonsyn | annotate R1 | NO — needs R1 |
| R4 | BaseVar+STITCH @2.0X: genotype concordance 99.08%, accuracy 0.98 | BaseVar+STITCH downsampling vs 10X truth | NO — cohort + 15×10X truth |
| R5 | 6 QTLs (FGF10 Chr11, FAM184B Chr2, NCAPG/DCAF16 Chr4 …); top-SNP p-values; QTL var 0.4–7.5% | GMAT GWAS | NO — needs phenotypes (not deposited) |
| R6 | SNP-heritabilities 7.5–39.1% (mean 19.0%) | GCTA | NO — needs phenotypes (not deposited) |
| R7 | LD decay r²=0.50 @6.5 kb; 151 selection regions / 309 genes | popgen on R1 | NO — cohort-scale |
| C1 | BaseVar builds (v1.2.0) and runs deterministically on real ultra-low-pass WGS, producing variant calls + per-sample allele-frequency/coverage output | BaseVar | YES |
| C2 | BaseVar applied to the paper's own rabbit data (PRJNA810279): a few SRR runs aligned to OryCun2.0 → BaseVar emits variant calls over a target region | BWA+BaseVar | ATTEMPT (best-effort) |
Why the headline numbers are out of scope (80/20 + feasibility)
- Cohort scale. PRJNA810279 = 642 runs, 7.06 Tbases, ~5.0 TB fastq (verified via ENA filereport — matches paper's "7305 Gb / 627 samples"). BaseVar and STITCH are cohort methods: every reported SNP count, AF, concordance and the imputed panel depend on pooling all ~627 low-pass samples. No reported number is scale-invariant, so none can be matched 1:1 without downloading + aligning the full 5 TB cohort — explicitly the hard last 20% we are told not to chase, and beyond a single room's feasible budget.
- Phenotypes not deposited. The GWAS/QTL/heritability results (R5, R6 — the paper's
headline "genetic architecture", incl. the FGF10 finding) require per-animal wool-trait
measurements (DFW, CVDFW, BW, LFW, LCW, RCW). The data-availability statement deposits
only raw reads (SRA); no phenotype file is deposited. Without phenotypes these
results are not reproducible from public data at all (a
data_restricted-class gap for the phenotype side). Flagged for the human auditor.
What we DO attempt (feasible, deterministic, on «our HPC»)
- C1 (primary): Build BaseVar @ pinned commit from source (htslib + cmake, C++17) on «our HPC»; run its bundled self-test — 100 real ultra-low-pass human NIPT BAMs over chr11:5,246,595-5,248,428 (HBB) + chr17:41,197,764-41,276,135 (BRCA1), hg19 ref subset → BaseVar VCF + coverage. Records that the paper's variant-calling code artifact compiles and produces calls deterministically. Uses BaseVar's bundled human test data, not rabbit data — stated honestly.
- C2 (best-effort): Download a few SRR runs from PRJNA810279, BWA-mem to OryCun2.0, BaseVar over a target region → confirm the tool runs on the paper's own data. Cannot match any reported count (cohort-scale); demonstration only.
Honest expected outcome
partial. We reproduc
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The paper's primary variant-calling code artifact (BaseVar) was genuinely reproduced — built from source at pinned commit 19482655 and run on «our HPC» (C1: 82 variants on 100 ultra-low-pass NIPT BAMs; C2: launched on the paper's own rabbit data) — but none of the reported quantitative values (18.5M SNPs, 1.74M Chr11 sites, GC 99.08%, FGF10 p=7.12e-23, h2 mean 19.0%) could be put against an output. The blockers split two ways: SNP/imputation counts are cohort-scale (~5 TB, our feasibility choice → our_method_choice), while the GWAS/heritability headline results are authors-side because the per-animal wool-trait phenotypes were never deposited (data_unavailable / value_not_derivable). Severity is indeterminate rather than severe — there is no measured discrepancy and no fabrication evidence, just a transparency/availability gap — so overall this is a yellow: a real, working tool with explainable, unverifiable numbers.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.