Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Accurate sequence variant genotyping in cattle using variation-aware genome graphs.

Genet Sel Evol · 2019
L1 45/100 PQI 81
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
45/100
Reproducibility score
1.7 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 6% of all assessed papers rank 1103 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL - honest reproduction with one clean code result and one own-data demonstration. C1 (grade=exact): the paper's exact tool GraphTyper v1.3 (commit 04ab5ee) builds and runs on «our HPC» (6 modern-toolchain build fixes, no algorithm change; recipe in kartei 'graphtyper'). C2 (grade=partial): the genotyping pipeline runs on the paper's own deposit - aligned an Original-Braunvieh-project bull (ERR2743225/ETH15, 2018, ~contemporaneous) to UMD3.1 (chr25, 10.76x) and ran GraphTyper v1.3 + bcftools(samtools model) on chr25:1-6Mb; biallelic-SNP overlap shared/GraphTyper=99.4%, shared/bcftools=51.2% (GraphTyper conservative on a single low-cov sample). This shows the callers substantially agree but is NOT the genome-wide 49-bull 85.5-90.4% figure (different N/coverage/tool-set). KEY FINDING: released GraphTyper v1.3 has a hardcoded human GRCh37 contig map and mislabels cattle variants as CHROM='Un' (vcf_break_down crashes) - the authors must have patched it for cattle; not fabrication, a reproducibility gap. NOT ATTEMPTED with reasons: Table 1 genome-wide counts (192-2792 CPU-h cohort-scale); Tables 3-4 accuracy/Mendelian (BovineHD/SNP50 array truth NOT deposited in PRJEB28191 -> not reproducible from public data). DATASET: PRJEB28191 is an umbrella accession now holding 450 WGS runs (2018-2024), far more than the paper's 49; the exact 49 are not labelled as a subset. Infra notes: shared «infra» account quota repeatedly exhausted (worked around with chr25-only BAM + retry-submit); no fabrication.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 44
    assessed: 2026-06-14 ⛓ bfb6adb84cca
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator headless) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Sequence variant genotyping using a variation-aware genome graph (Graphtyper) is more accurate and sensitive than conventional linear-reference-based methods (GATK, SAMtools) in cattle.

Core claims
  • Graphtyper outperformed GATK and SAMtools in genotype concordance, non-reference sensitivity, and non-reference discrepancy compared to microarray genotypes finding
  • Graphtyper produced the fewest Mendelian inconsistencies among sire-son pairs compared to GATK and SAMtools finding
  • Genotype phasing and imputation with Beagle improved sequence variant genotype quality for all three tools, especially for low-coverage animals finding
  • After imputation, genotype concordance with microarray data was nearly identical across the three methods (99.32%, 99.46%, 99.24% for GATK, Graphtyper, SAMtools) finding
  • Variant filtering using commonly used criteria slightly improved genotype concordance but decreased sensitivity finding
  • Graphtyper required more computing resources than SAMtools but less than GATK finding
  • Graph-based genotyping enables use of variation-aware reference genomes that can incorporate cohort-specific sequence variants, unlike current linear-reference-based methods mechanism
  • The Graphtyper source code was modified to enable analysis of the cattle chromosome complement, since the original implementation was limited to human chromosomes method
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome sequencing (150 bp paired-end) 49 Original Braunvieh cattle (blood/semen-derived DNA) none aligned sequencing reads, coverage depth Illumina HiSeq 2500 and HiSeq 4000
Multi-sample variant discovery/genotyping 49 Original Braunvieh cattle, aligned BAM files none polymorphic sites (SNPs, indels), genotype likelihoods GATK (version 4.0.6), HaplotypeCaller/GenomicsDBImport/GenotypeGVCFs
Graph-based multi-sample variant discovery/genotyping 49 Original Braunvieh cattle, aligned BAM files none polymorphic sites (SNPs, indels) from variation-aware genome graph Graphtyper (version 1.3, modified for cattle genome)
Multi-sample variant discovery/genotyping 49 Original Braunvieh cattle, aligned BAM files none polymorphic sites (SNPs, indels) SAMtools mpileup (version 1.8) / BCFtools call
Genotype phasing and imputation 49 Original Braunvieh cattle sequence variant genotypes none imputed/refined genotype quality Beagle (version 4.1), genotype likelihood mode
SNP microarray genotyping (comparison/validation) 49 Original Braunvieh cattle none genotype concordance, non-reference sensitivity, non-reference discrepancy vs sequence-derived genotypes Illumina BovineHD (N=29) and BovineSNP50 (N=20) BeadChip
Mendelian inconsistency analysis 9 sire-son pairs among the 49 sequenced animals none proportion of opposing homozygous genotypes
Key results
  • GATK, Graphtyper, and SAMtools discovered 21,140,196, 20,262,913, and 20,668,459 polymorphic sites, respectively
  • After imputation, genotype concordance with array data was 99.32% (GATK), 99.46% (Graphtyper), 99.24% (SAMtools) 99.24-99.46%
  • GATK detected four times more multiallelic SNPs (246,220) than SAMtools or Graphtyper (~63,000-67,000) 4-fold
  • SAMtools identified the largest number and highest proportion (14.9%, 3,076,421) of indels among the three tools 3,076,421 indels
  • Of SNPs detected, 7.46% (GATK), 8.31% (Graphtyper), and 5.02% (SAMtools) were novel (not in dbSNP 150) 5.02-8.31%
  • An intersection of 15,901,526 biallelic SNPs (85.51-90.39% of each tool's calls) was common to all three tools 15,901,526 SNPs
  • 1,299,467 biallelic indels were common to all three tools, with a smaller intersection proportion than for SNPs 1,299,467 indels
  • Ti/Tv ratio of detected SNPs was 2.09 (GATK), 2.07 (Graphtyper), 2.05 (SAMtools) 2.05-2.09
Key statistics
  • count 21,140,196 / 20,262,913 / 20,668,459 variants (GATK/Graphtyper/SAMtools) (total autosomal polymorphic sites detected per tool)
  • other 99.32% / 99.46% / 99.24% (genotype concordance with array data after imputation, GATK/Graphtyper/SAMtools)
  • fold_change 4x more multiallelic SNPs (246,220 vs ~63-67k) (GATK vs SAMtools/Graphtyper multiallelic SNP counts)
  • mean 12.75-fold (range 6.00-37.78) (average sequencing depth per animal)
  • mean 98.44% (range 91.06-99.59%) (average percentage of reads mapped to reference genome)
  • count 15,901,526 common biallelic SNPs; 466,029 (2.93%) novel, Ti/Tv 1.81; overall common Ti/Tv 2.22 (SNP intersection across all three tools)
  • other 98.91% (average genotyping rate at autosomal SNPs on the microarray)
  • other Ti/Tv 2.17 / 2.13 / 2.11 (per-animal, GATK/Graphtyper/SAMtools) (average per-animal Ti/Tv ratio for biallelic SNPs)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a methods-comparison study evaluating three sequence variant genotyping pipelines (GATK, Graphtyper, SAMtools) on whole-genome sequencing data from 49 Original Braunvieh cattle, with accuracy and sensitivity assessed by comparing sequence-derived genotypes to microarray-derived genotypes (genotype concordance, non-reference sensitivity, non-reference discrepancy) and by counting Mendelian inconsistencies in nine sire-son pairs. To test whether the three quality metrics differed between tools, non-parametric Kruskal–Wallis tests were used followed by pairwise Wilcoxon signed-rank tests. Results were largely reported descriptively as counts, percentages, and ratios in tables, with ranges given for some per-animal summaries.

Replicationbiological Sample size49 Original Braunvieh bulls selected as frequently used in artificial insemination or explaining a large fraction of genetic diversity; nine sire-son pairs used for Mendelian inconsistency; no formal power/sample-size calculation described Groupsthree genotyping tools (GATK, Graphtyper, SAMtools) on the same 49 animals Pairingpaired Randomization/blindingna Dispersionrange Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Kruskal–Wallis test (non-parametric, omnibus across three tools) comparing genotype concordance, non-reference sensitivity, and non-reference discrepancy among GATK, Graphtyper, and SAMtools 49 sequenced animals (per-animal metric values), as stated not stated
Pairwise Wilcoxon signed-rank test pairwise (post-hoc) comparison of the three metrics between pairs of tools 49 sequenced animals (paired per-animal values), as stated not stated
Approaches that could also have been used
  • Differences between the three tools were assessed with a Kruskal–Wallis test followed by pairwise Wilcoxon signed-rank tests on the per-animal metric values.
    Could also: A Friedman test (with a matched post-hoc such as Nemenyi or pairwise Wilcoxon with adjustment) could also be used since the three tools were evaluated on the same animals. — The Friedman test explicitly models the repeated-measures/blocked structure (same animals measured under three tools), which can increase sensitivity when within-animal correlation is present.
  • The pairwise Wilcoxon comparisons follow an omnibus Kruskal–Wallis test, with no explicit multiple-testing adjustment described.
    Could also: A multiplicity correction such as Holm, Benjamini–Hochberg FDR, or Bonferroni across the pairwise comparisons could also be reported. — An explicit correction documents control of the family-wise error or false discovery rate across the set of pairwise tests and makes the inferential scope transparent.
  • Metric comparisons used non-parametric rank-based tests.
    Could also: Reporting effect sizes (e.g., median differences with 95% confidence intervals, or a rank-based effect size) alongside the tests could also be done. — Effect sizes and intervals convey the magnitude and precision of the tool differences, complementing the significance statements and aiding practical interpretation.
  • Per-animal and aggregate summaries are reported primarily as counts, percentages, and min–max ranges.
    Could also: Summaries could also include a measure of central tendency with SD/IQR or a 95% confidence interval. — Adding a dispersion measure such as IQR or a CL alongside the range conveys the typical spread rather than only the extremes, which is often preferred for summarizing distributions.
  • Significance between tools is described qualitatively (significant/not) without tabulated p-values.
    Could also: Exact p-values for each test could also be tabulated. — Reporting exact p-values lets readers gauge the strength of evidence directly and supports reuse in meta-analyses or re-evaluation under alternative thresholds.
  • Accuracy was benchmarked against microarray-derived genotypes, which the authors note may be subject to ascertainment bias.
    Could also: Stratifying concordance metrics by allele frequency or genomic accessibility, or modeling them jointly (e.g., a mixed-effects model with animal as a random effect), could also be used. — Stratification or a model-based approach can characterize where tools differ (e.g., rare vs. common variants) and account for the repeated-measures structure within animals in a single framework.
Software: R 3.3.3 · ggplot2 (visualization) 3.0.0 · GATK 4.0.6 (also 3.8 for realignment/BQSR) · Graphtyper 1.3 · SAMtools/BCFtools 1.8 · Beagle 4.1 · UpSet R package

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
63
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.
Cited by (assessed papers) (1)

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

PRJEB28191 BioProject in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

Downstream reach in the literature

3 downstream papers · 1 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

This paper is currently under reproducibility review (see the verdict above). The map below shows where the data in question has propagated — so reuse can be traced, not so the downstream work is presumed affected.
PRJEB28191 BioProject reused by 5 papers in the literature
Most-cited downstream papers:

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-31092189

Paper: Crysnanto D, Wurmser C, Pausch H. Accurate sequence variant genotyping in cattle using variation-aware genome graphs. Genet Sel Evol 2019;51:21. DOI 10.1186/s12711-019-0462-x.

One-line: Benchmark of three WGS variant callers — GATK 4.0.6, GraphTyper 1.3, SAMtools 1.8 (mpileup→bcftools) — on 49 Original Braunvieh (OB) bulls (PRJEB28191), aligned to UMD3.1 with BWA-mem 0.7.12, genotypes refined with Beagle 4.1. Accuracy assessed against BovineHD (N=29, 777,962 SNPs) and BovineSNP50 (N=20, 54,001 SNPs) array genotypes and via Mendelian consistency in 9 sire–son pairs.

Pipeline-derived results (candidate in-scope)

Result Pipeline Reproducible here?
Variant counts per tool (Table 1) full 49-sample genome-wide call, 3 tools No — 1000s CPU-h (paper: 1066–2792 CPU-h), exceeds 80/20
Ti/Tv, % novel vs dbSNP (Table 1/text) as above No — same cohort-scale dependency
Genotype concordance / non-ref sensitivity & discrepancy (Table 3) call vs array truth No — array genotypes NOT in PRJEB28191, not deposited
Mendelian inconsistency (Table 4) full cohort + pedigree (9 pairs) No — needs all 49 calls + pedigree
CPU-hours per tool (Fig 4) full cohort run No — cohort-scale

What we DO attempt (clear, bounded, honest)

  • C1 — Code artifact builds & runs. Build the paper's tool GraphTyper v1.3 (commit 04ab5ee460fa36129fb0d8ea5d4b72adc3836f52) on «our HPC» and confirm the binary runs. CLEAR yes/no. This is the central tool of the paper.
  • C2 — Pipeline runs on the paper's OWN data (demonstration). Take one OB bull from PRJEB28191 (smallest, ERR6473928), align with BWA-mem to UMD3.1 (as the paper), run GraphTyper genotype and bcftools (samtools mpileup) on a single chromosome, and report variant counts + cross-tool overlap. This demonstrates the genotyping pipeline operates on the paper's data and that the callers largely agree — it is a provisional demonstration, NOT a reproduction of the genome-wide Table 1/3/4 numbers (different N, no array truth).

Out of scope (not attempted, with reason)

  • Array-based accuracy tables (no truth data deposited).
  • Full 49-sample, 3-caller, genome-wide run + Beagle refinement (compute >> 80/20).
  • Wet-lab / sequencing steps.

Expected outcome class: partial — code artifact + own-data demonstration; headline accuracy not reproducible from the public deposit (array truth absent; cohort-scale compute).

Figures / tables: TableFig 4
C1
Reported
Graphtyper (version 1.3) used as the variation-aware graph genotyper (Methods, Variant calling)
Reproduced
Built from source at tag v1.3 (commit 04ab5ee) on «our HPC» under gcc-10 («job» COMPLETED, 5:14). Binary runs: 'Version: 1.3-dirty (04ab5ee)', SMOKE_OK. 6 documented modern-toolchain build patches, no algorithm change. sha256 4d5317a99b62f20380808bed598e7e99cd0069152d17c9bfa9da91d1b32382f5.
exact
C2
Reported
GraphTyper genotyping on the paper's own WGS data; reported genome-wide inter-tool biallelic-SNP overlap 85.5-90.4% across 49 bulls (Table 1)
Reproduced
Pipeline runs end-to-end on the paper's deposit (PRJEB28191). Aligned bull ERR2743225/ETH15 (2018 deposit, contemporaneous with the paper) to UMD3.1 with BWA-mem -> chr25 BAM at 10.76x mean depth («job»). Ran GraphTyper v1.3 (construct->index->call) and bcftools/samtools-mpileup on chr25:1-6Mb (~9x). Biallelic-SNP site overlap: GraphTyper 5629, bcftools 10940, shared 5597 -> shared/GraphTyper=99.4%, shared/bcftools=51.2%, Jaccard=0.51. GraphTyper's calls are nearly all (99.4%) confirmed by bcftools; bcftools calls ~2x more (GraphTyper is conservative on a single low-cov sample - it is designed for cohort genotyping). This DEMONSTRATES the pipeline + substantial cross-tool agreement but is NOT a 1:1 reproduction of the genome-wide 49-bull 85.5-90.4% figure (different N, coverage, tool-set, and chr25-window only).
partial
C2_caveat_graphtyper_human_contigs
Reported
(implicit) GraphTyper 1.3 applicable to cattle (UMD3.1)
Reproduced
FINDING: the released GraphTyper v1.3 has a HARDCODED human GRCh37 contig map - it wrote all cattle variants with CHROM='Un' under a human ##contig header (chr1 length 249250621 etc) and vcf_break_down crashed (std::out_of_range). Variant POS were correct chr25 coordinates and the graph held only chr25, so calls were valid after relabelling Un->25. Implies the paper's authors must have patched GraphTyper's contig table for cattle; the code as-released does not handle UMD3.1 out of the box. Not fabrication - a reproducibility gap in the shipped tool.
partial
T1_variant_counts
Reported
GraphTyper 20,262,913 / GATK 21,140,196 / SAMtools 20,668,459 total variants; 15,901,526 biallelic SNPs common to all three tools (Table 1)
Reproduced
not attempted - requires calling all 49 bulls genome-wide with three tools (paper reports 192-2792 CPU-h per tool); beyond an 80/20 reproduction.
partial
T3_concordance
Reported
GraphTyper genotype concordance 97.71% raw / 99.52% after Beagle vs array truth (Table 3)
Reproduced
not attempted - the BovineHD (N=29) / BovineSNP50 (N=20) ARRAY truth genotypes are NOT in PRJEB28191 and not deposited anywhere; not reproducible from the public deposit (a deposition gap, not evidence of fabrication).
did not match
T4_mendelian
Reported
GraphTyper Mendelian opposing-hom 0.36% SNP / 0.54% indel (Table 4)
Reproduced
not attempted - needs all 49 bulls called + pedigree of 9 sire-son pairs.
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator headless) · v1.0 L1 45/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

<synthetic>

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

867.3 k
tokens (I/O) · 101.1 M incl. cache
399 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.