Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Assessment of genotyping array performance for genome-wide association studies and imputation in African cattle.

Genet Sel Evol · 2022
L1 90/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
90/100
Reproducibility score
0.9 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 79% of all assessed papers rank 211 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH for the part that ships data: 1:1 reproduction of the paper's headline per-breed WGS SNP counts. The shipped Zenodo deposit (10.5281/zenodo.6855979) contains exactly one artifact: the final VQSR tranche-99.0 autosomal biallelic SNP joint-call VCF for 40 Boran + 40 N'Dama + 40 Holstein (120 samples, 35,842,537 SNPs). Recomputing per-breed segregating-SNP counts with bcftools 1.11 on «our HPC» reproduced all three reported values EXACTLY to printed precision: Boran 32,010,959 (paper 32.0M), N'Dama 17,841,567 (17.8M), Holstein 13,330,181 (13.3M); diversity ordering Boran>N'Dama>Holstein confirmed. Overall status = partial because this is one in-scope result out of a large multi-part study. NOT ATTEMPTED (the hard ~20%, inputs never published): Table 1 array-overlap counts and Fig 1/6 LD r2>0.8 tagging (need proprietary manifests of 23 commercial bovine arrays on ARS-UCD1.2); Figs 2/4/7/9/10 imputation ER2/dosage-R2 (need Minimac4+BEAGLE phasing + the 289-individual global reference panel + masked per-array target sets, none in the deposit). The paper's only code link is the third-party SNPchiMp strand-conversion tool, not the analysis pipeline; the authors' analysis scripts were not released. No fabrication concern for the in-scope results: the three counts are fully derivable from the shipped data and match exactly.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.6855979

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 90
    assessed: 2026-06-15 ⛓ ffbe15c3e7a5
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Which currently available bovine genotyping arrays best capture genetic diversity across African cattle breeds and perform best for genome-wide imputation, in order to determine the optimal strategy for performing GWAS in African cattle (a taurine/indicine admixture).

Core claims
  • Commercially available bovine arrays are ineffective at capturing variants segregating among African indicine animals, with only 6% of high-LD (r2>0.8) variants captured by the best arrays versus 17% in African taurine and 25% in European taurine. finding
  • Imputation from HD arrays can capture most WGS variants (accuracies up to 0.93), performing best with a global rather than continent-specific reference panel, partly reflecting high admixture on the continent. finding
  • The GGPF250 array performs best for tagging functional WGS variants and for imputation of functional variants. finding
  • A two-stage imputation approach (low-density to HD, then HD to WGS) allows low-density arrays to perform almost as well as HD arrays, potentially reducing GWAS costs. method
  • There is no advantage to using indicus-specific arrays for indicus breeds, regardless of study objective; HD and BOS1 are best for GWAS in both taurine and indicine breeds, and GGPF250 is preferable for fine-mapping. finding
  • Using a reference panel representing global bovine diversity improves imputation accuracy, particularly for non-European taurine populations. finding
  • Approach demonstrated and validated on a real cohort of 2481 genotyped African cattle samples. resource
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome sequencing (Illumina WGS) with alignment and variant calling 409 global cattle WGS including 40 NDama (W. African taurine), 40 Boran (E. African indicine Zebu), 40 Holstein-Friesian (European taurine) none SNP variant calls used to evaluate array variant capture Illumina sequencing; BWA-mem v0.1.17; GATK best practices with VQSR; BCFtools v1.11
Array variant capture / LD analysis (in silico overlap of WGS with array positions) NDama, Boran, Holstein-Friesian WGS vs 23 commercial bovine arrays (~3000-777,962 SNPs) none proportion of WGS variants in high LD (r2>0.8) with array SNPs; allele frequencies; LD decay VCFtools v0.1.15; PLINK v1.90b4
Genotype imputation (masked WGS analysis) Target sets derived from WGS overlapping each of 23 arrays; Global Reference Panel (289 individuals, 55 populations) and African (87), Asian (106), European (77) subset panels reference panel choice (global vs continent-specific) imputation accuracy to WGS level
Two-stage imputation (50k > HD > WGS) Illumina HD array intermediate reference panel; cattle genotype data low-density to HD to WGS imputation strategy imputation accuracy vs direct HD imputation Illumina bovine HD array (777,962 SNPs)
SNP array genotyping (real cohort) 3092 cattle from Tanzania, Ghana, Nigeria, Burkina Faso (combined to 3852; 2481 retained after QC) none genotypes for variant capture/imputation validation Illumina bovine HD array (777,962 SNPs)
SNP array genotyping (50k subset) 668 of the 3092 African cattle (602 retained after QC) none 50k genotypes for two-stage imputation GeneSeek bovine 50k array (BOVG50V1, 47,844 SNPs)
Key results
  • Only 6% of high-LD (r2>0.8) WGS variants in African indicine animals are captured on the best-performing arrays 6%
  • 17% of high-LD variants captured in African taurine cattle 17%
  • 25% of high-LD variants captured in European taurine cattle 25%
  • Imputation from HD arrays achieves accuracies up to 0.93, best with a global reference panel up to 0.93
  • GGPF250 array best for tagging functional WGS variants and functional-variant imputation
  • Two-stage imputation (50k>HD>WGS) performs almost as well as direct HD imputation
  • Of 23 arrays tested, the 16 highest-density arrays imputed successfully while the 7 lowest-density (<~20,500 SNPs) failed 16 of 23
Key statistics
  • correlation r2>0.8 (LD threshold for high correlation) (threshold for WGS-array variant capture)
  • correlation accuracies up to 0.93 (imputation accuracy from HD arrays with global reference panel)
  • count 6% (indicine high-LD variants on best arrays)
  • count 17% (African taurine high-LD variants captured)
  • count 25% (European taurine high-LD variants captured)
  • count 2481 (African cattle samples retained after QC for HD genotyping)
  • count 289 individuals, 55 populations (Global Reference Panel (13 European, 12 African, 28 Asian, 2 Middle Eastern))
  • count 409 whole-genome sequences (WGS spanning global cattle breeds used in study)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study benchmarked 23 commercial bovine genotyping arrays by estimating pairwise linkage disequilibrium (r²) between array SNPs and whole-genome sequence (WGS) variants using PLINK across 500-kb windows, with r² > 0.8 as the high-LD tagging threshold, across three breeds (Holstein-Friesian, NDama, Boran; 40 WGS individuals each). Imputation performance was assessed using masked WGS target sets imputed back to full WGS coverage via a global reference panel of 289 animals and three continent-specific subpanels, including a two-stage (low-density → HD → WGS) strategy. Results were reported as percentages of WGS variants tagged at r² > 0.8 and as correlation-based imputation accuracy coefficients, validated on a real cohort of 2481 HD-genotyped African cattle.

Replicationbiological Sample sizeSample sizes stated throughout (120 WGS for array evaluation; 289 in global WGS reference panel; 2481 HD-genotyped animals after QC; 602 animals with 50k array after QC); no power calculation or a priori sample-size justification described GroupsThree breeds (Holstein-Friesian, NDama, Boran); 23 commercial arrays varying in density (~3,000–777,000 SNPs); four reference panels (global, African, Asian, European subsets) Pairingna Randomization/blindingnot stated Dispersionnone Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Pairwise LD (r²) estimation to assess array tagging coverage of WGS variants (PLINK, 500-kb window, r² > 0.8 threshold) Comparison of 23 arrays for three breeds — Holstein-Friesian, NDama (W. African taurine), Boran (E. African indicine) 120 WGS individuals (40 per breed) not stated
Imputation accuracy (correlation-based metric; specific statistic not named in available text) Masked imputation from array-overlapping WGS subsets to full WGS, across 15 arrays and four reference panels (global n=289; African n=87; Asian n=106; European n=77) 289 individuals in the Global Reference Panel and derived continental subsets not stated
GATK Variant Quality Score Recalibration (VQSR) at 99% tranche for SNP quality filtering Quality control of WGS SNP calls prior to all downstream analyses 120 WGS individuals for breed-specific calls; 289-individual global panel not stated
LD decay comparison across breeds (described as a comparison; statistical formalism not stated in available text) Comparison of LD decay curves for Holstein-Friesian, NDama, and Boran 40 WGS individuals per breed not stated
Approaches that could also have been used
  • Array tagging coverage was summarized using a single LD threshold of r² > 0.8
    Could also: Multiple r² thresholds (e.g., > 0.6 and > 0.5) could be evaluated and reported alongside r² > 0.8 — A single threshold may not capture how tagging performance degrades at lower LD, and reporting a range of thresholds—as is common in human GWAS array benchmarking—would allow readers to judge coverage across different study designs that tolerate lower LD
  • Imputation accuracy was summarized as a single aggregate correlation coefficient per array/panel combination
    Could also: Accuracy could additionally be stratified by minor allele frequency (MAF) bin (e.g., < 0.05, 0.05–0.20, > 0.20) — Imputation accuracy is known to degrade at lower MAF; MAF-stratified metrics would reveal which array–panel combinations are most effective for rare-variant GWAS versus common-variant discovery, making the benchmarking more actionable for different study objectives
  • Global and continent-specific reference panels were compared as discrete pre-defined groupings
    Could also: Panel size could be held constant while varying composition, or incremental addition of African-breed WGS to the global panel could be simulated — Because African breeds are numerically underrepresented in the global panel, disentangling the effect of panel size from panel composition would clarify whether diversity or quantity drives the observed accuracy gains, informing sequencing investment priorities
  • Quality control applied fixed thresholds across all breeds (MAF > 0.01, call rate ≥ 0.90 or CR ≥ 75%, GQ ≥ 25, missingness < 25%)
    Could also: Sensitivity analyses varying QC thresholds, or breed-specific adaptive thresholds guided by empirical quality distributions, could also be reported — Fixed thresholds are a well-established standard, but may differentially exclude low-frequency variants that are informative specifically for underrepresented African populations; reporting QC sensitivity would help readers assess robustness of conclusions to threshold choice
  • The LD decay comparison across breeds appears to be presented visually or descriptively
    Could also: Bootstrap confidence intervals or permutation-based tests around mean r² per distance bin could be computed to formally quantify uncertainty in the breed differences — Confidence intervals on LD decay curves (e.g., 95% bootstrap CI on mean r² at each kb interval) would support quantitative claims about differences in LD structure between taurine and indicine populations and facilitate comparison with published LD decay estimates in other studies
  • Two-stage imputation (low-density → HD → WGS) was compared to direct imputation as a binary design choice
    Could also: Imputation accuracy could be modeled as a continuous function of input array density (e.g., regression or interpolation across all 15 arrays that achieved successful imputation) — A cost-accuracy curve relating array density to imputation accuracy would allow practitioners to identify the minimum-cost array sufficient to achieve a target accuracy, extending the practical value of the benchmarking beyond the specific arrays tested
Software: BWA-mem 0.1.17 · GATK (with VQSR) · BCFtools 1.11 · VCFtools 0.1.15 · PLINK 1.90b4 · Imputation software (not named in available text excerpt)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
22
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36057548

Paper: Riggio et al. 2022, Assessment of genotyping array performance for genome-wide association studies and imputation in African cattle. Genet Sel Evol 14:54. PMID 36057548 / PMC9441065 / DOI 10.1186/s12711-022-00751-5.

Artifacts available (resolved)

  • Code link in paper = github.com/nicolazzie/SNPchimpRepo — this is the SNPchiMp v.3 toolbox (iConvert.py, strand TOP→FOR conversion). It is a third-party utility, NOT the paper's analysis pipeline. Last push 2021-03-26, commit 31ed929, public, not archived, no license file. The paper's actual analysis (array overlap, LD tagging, imputation) was done with custom R/Python scripts that were NOT released.
  • Data = Zenodo 10.5281/zenodo.6855979 (CC-BY-4.0, open). Contains exactly ONE genotype artifact: BOR-NDM-HOL_sample40_finalAutoSNPs_tranche99.0_recode.vcf.gz (~1.19 GB, +.tbi, +.md5). This is the breed-reference WGS variant set for Boran / N'Dama / Holstein — the final, VQSR tranche-99.0, autosomal, biallelic SNP joint call set used as the imputation reference / for the diversity comparison.

Pipeline-derived results — in / out of scope

IN SCOPE (low-hanging, derivable directly from the shipped Zenodo VCF)

The paper reports per-breed WGS variant counts in Results ("Genetic variation in the three breeds"): Boran 32.0 M, N'Dama 17.8 M, Holstein 13.3 M SNPs, and states the indicine Boran is the most variable. From the shipped final autosomal SNP VCF we can faithfully recompute, with standard tools (bcftools/PLINK, the same families the paper used):

  • total autosomal biallelic SNPs in the final set;
  • sample/breed composition (from VCF header);
  • per-breed segregating-SNP counts (sites carrying a non-reference allele in each breed) → compare pattern + magnitude to 32.0/17.8/13.3 M;
  • genome-wide MAF distribution + per-breed mean MAF;
  • Ts/Tv as a QC sanity metric.

This is the honest 1:1 core: the relative diversity ordering (Boran ≫ N'Dama > Holstein) is the central, checkable, reproducible claim. NB: the paper's 32/17.8/ 13.3 M are from the per-breed pre-merge discovery stage, whereas the Zenodo file is the final joint, autosomal, tranche-99 filtered set — so absolute counts are expected to be LOWER; we report both and grade honestly.

OUT OF SCOPE (the hard ~20% — required inputs NOT in the deposit; not attempted)

  • Array overlap counts (Table 1) + LD r²>0.8 tagging (Fig 1, 6) — need the manifests (ARS-UCD1.2 positions) of 23 commercial bovine arrays (HD, BOS1, GGP*, SNP50V3, …). These manifests are proprietary/manufacturer files, not shipped with the paper or Zenodo.
  • Imputation accuracy (Figs 2,4,7,9,10): ER2 / dosage-R² — need Minimac4 + BEAGLE phasing + the 289-individual global reference panel and per-array masked target sets, none of which are in the single Zenodo deposit; the driving scripts were not released.
  • Real-data cohort (HD: 2,481 samples; 50k: 602 samples) — genotype data not public in this deposit.

Rationale for the cut: per HARD RULE 3 (80/20) and the operator directive ("a few CLEAR data points, honest 1:1; do NOT chase the last 20%"), we reproduce the diversity result that the shipped data fully supports, and explicitly do not attempt the array/imputation results whose inputs were never published.

C1
Reported
32.0 million (Boran WGS SNPs)
Reproduced
32,010,959 (32.0 M)
exact
C2
Reported
17.8 million (N'Dama WGS SNPs)
Reproduced
17,841,567 (17.8 M)
exact
C3
Reported
13.3 million (Holstein WGS SNPs)
Reproduced
13,330,181 (13.3 M)
exact
C4
Reported
Boran > N'Dama > Holstein (indicine most variable)
Reproduced
32.0M > 17.8M > 13.3M
exact
C5
Reported
(no single in-text number) final autosomal SNP set
Reproduced
35,842,537 autosomal biallelic SNPs; 120 samples (40/breed); mean MAF 0.128
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 90/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

For the only result the deposited data supports — the per-breed WGS SNP counts — this is a clean 1:1 reproduction: Boran 32,010,959, N'Dama 17,841,567, Holstein 13,330,181 all round exactly to the paper's 32.0M/17.8M/13.3M, and the diversity ordering (indicine most variable) is confirmed. There is no discrepancy on our side or theirs and no fabrication concern; the only difference is rounding. The major caveat is scope/availability, not quality: ~80% of the study (Table 1 array overlaps, Figs 1/6 LD tagging, Figs 2/4/7/9/10 imputation) was untestable because the analysis scripts were never released and only one of many required artifacts was deposited — an authors'-side transparency gap that does not bear on the correctness of the reproduced numbers.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

145.6 k
tokens (I/O) · 9.5 M incl. cache
26 min
runtime · 0.42 CPU-h
1.3 GB
peak RAM
2
HPC jobs
hummel
machine