A comparative study on recombination activity in cattle.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡Could not use the authors’ exact input data
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH -> 1:1 reproduction of the deposited-data-derived results, run FRESH on «our HPC» compute node n095 (SLURM «job», base R 4.6.0; confirms the earlier «job» on n094). Raw BovineSNP50 genotypes are RESTRICTED (Qualitas AG) and the brief's SRA PRJNA668863 is a mis-harvested unrelated long-read assembly, so the from-genotype map-ESTIMATION pipeline (PLINK/FImpute/LINKPHASE3/hsphase/hsrecombi) cannot be re-run. But the paper deposits its processed outputs openly on Zenodo (maps 10.5281/zenodo.17909700; per-parent phenotypes 10.5281/zenodo.17951753; code github.com/nmelzer/CLARITY = 10.5281/zenodo.18668018), so Table 3 was recomputed directly from those deposits. Result: all 24 genetic-map lengths (HMM male+female, deterministic-male, likelihood-male x6 breeds) reproduce to 3 dp (HMM male EXACT via per-marker cumulative cM; female/det/lik within <=0.003 M); all 24 crossover-count mean+SD values (sires+dams x6 breeds) reproduce exactly; the 34,755 shared-SNP count reproduces from TWO independent files (the naive intersection 38,572 over-counts); Table-2 parent counts (sum 39,714) match the recombination_traits.txt row count exactly; cM:Mbp ranges, named max chromosomes (male BTA19 BrownSwiss 1.671, female BTA28 Angus 1.484), the 17.91% female shortening, and the 67.5%/9.1% method-comparison percentages all match. Status is PARTIAL (not 'reproduced') because this is a faithful deposited-data -> reported-value verification, NOT an independent raw -> map reproduction, and the GWAS/GBLUP variance-component results (ASReml-driven) and the simulation study were not attempted (honestly out of scope). Two independent deposits both matching, including a non-obvious filtered count (34,755) and both mean AND SD of phenotypes, argue strongly against fabrication. Grades provisional pending human PDF check.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 99assessed: 2026-06-16 ⛓ 03fbee47f881
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study aimed to derive breed-specific genetic (linkage) maps for six commercial cattle breeds from medium-density SNP data and determine whether recombination activity (autosomal crossover count and intra-chromosomal allelic shuffling) differs by sex and breed, including a GWAS to identify genome regions associated with recombination activity.
- ★ Genotype data with high systematic missingness across breeds and arrays can be streamlined and analysed with three complementary recombination-estimation approaches (HMM-based LINKPHASE3, deterministic hsphase, likelihood-based hsrecombi) method
- ★ Female recombination activity and female genetic maps are clearly distinct from male maps in cattle finding
- ★ Male genetic maps of Brown Swiss, Simmental and Angus cluster separately from other breeds' male maps finding
- ★ Female genetic maps of dairy breeds are distinct from female maps of dual-purpose breeds, while male maps do not cluster by breed purpose finding
- ★ GWAS on mean autosomal crossover count and intra-chromosomal allelic shuffling identified chromosomal regions on BTA6 and BTA10 (strong evidence in Brown Swiss and Swiss Holstein) near genes known to affect recombination activity finding
- The HMM-based LINKPHASE3 approach, incorporating multi-locus likelihood inference, should deliver similar or better recombination rate estimates than the deterministic (hsphase) or likelihood-based pairwise (hsrecombi) methods mechanism
- The R Shiny app CLARITY was revised to v3.0.0 to provide access to the derived bovine genetic maps as a resource for genomic evaluation and breeding strategies resource
- ★ Estimated male map length varied from 26.97 to 29.83 M and female map length from 23.32 to 26.08 M between the six breeds finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| SNP genotyping (medium-density arrays, ~42-58K 'chip 4' sub-panel) | six cattle breeds: Holstein-CH, Brown Swiss, Simmental, Original Braunvieh, Limousin, Angus | none | SNP genotype calls used for recombination analysis | Illumina Bovine SNP50 array v1/v2 and other commercial arrays |
| Recombination rate / genetic map estimation (HMM-based, haplotype phasing + crossover identification) | paternal and maternal half-sib families across six cattle breeds | none | sex-specific genetic-map coordinates (Morgan units), crossover events | LINKPHASE3 |
| Recombination rate estimation, deterministic multipoint method | paternal half-sib families with ≥30 genotyped progeny | none | male recombination rate estimates | R package hsphase v2.0.2 |
| Recombination rate estimation, likelihood-based pairwise linkage/LD method | paternal half-sib families with ≥30 genotyped progeny | none | male recombination rate estimates | R package hsrecombi 1.0.1 |
| Simulation study, synthetic breeding population | AlphaSimR-simulated population, 1000 founders, single 1M chromosome, 4000 markers, 500x1000 progeny | varied number of half-sib families (10-500), progeny per family (50-1000), marker count (1500-4000), and 50% missingness | accuracy (acc) and mean squared error (mse) of estimated male genetic map | AlphaSimR v1.3.2 |
| Semi-real simulation using empirical haplotypes | Brown Swiss empirical parental haplotypes on BTA4 (chip 4 SNPs, q=2403) | simulated meioses/crossovers with varied N sires (10-500), n progeny (50-500), and 50% missingness | accuracy and mse of estimated male genetic map | FImpute v3 |
| Genomewide association study (GWAS) | six cattle breeds, notably Brown Swiss and Swiss Holstein | none | association of SNPs with mean autosomal crossover count and intra-chromosomal allelic shuffling per parent | — |
| PCA and hierarchical clustering of recombination maps | sex- and breed-specific recombination rate matrices across shared SNPs, six breeds | none | clustering/similarity of breed- and sex-specific genetic maps | R package dimensio v0.14.0; R stats::hclust (L2 norm) |
- – Male genetic map length ranged from 26.97 to 29.83 M across the six breeds
- – Female genetic map length ranged from 23.32 to 26.08 M across the six breeds
- – Female recombination maps were clearly distinct from male maps
- – Male maps of Brown Swiss, Simmental and Angus clustered separately from other male maps
- – Female maps of dairy breeds were distinct from female maps of dual-purpose breeds
- ▲ GWAS revealed two chromosomal regions (BTA6 and BTA10) with strong evidence in Brown Swiss and Swiss Holstein, plus a suggestive signal on BTA10, near genes known to affect recombination activity
- – Mean paternal half-sib family size varied strongly among breeds: 16 (Holstein-CH), 23 (Brown Swiss), 6 (Simmental), 7 (Original Braunvieh), 5 (Limousin), 4 (Angus)
- – Proportion of systematically missing genotypes per breed (averaged over chromosomes) ranged from 0.151 (Limousin) to 0.478 (Holstein-CH)
- count sample sizes ranging from 4,181 (Angus) to 76,875 (Holstein-CH) genotyped individuals (total genotyped animals per breed (Table 1))
- other male map length 26.97-29.83 M (range across six breeds)
- other female map length 23.32-26.08 M (range across six breeds)
- count q (SNPs used) ranged 44,466-51,603 per breed after quality control (Table 1, number of SNPs for recombination rate analysis)
- other missing genotype proportion 0.151-0.478 (Table 1, proportion of missing genotypes averaged over chromosomes, by breed)
- other MAF filter > 0.01; Mendel error exclusion thresholds: families >5%, SNPs >10% (genotype quality control criteria)
- other LINKPHASE3 marker confidence score (MCS) ≥ 0.986; individuals retained if ≤ 58 crossovers (second-run filtering criteria for final map estimates)
- count maximum paternal half-sib family size 2,663 (Brown Swiss) vs 68 (Angus) (variation in largest paternal half-sib family size across breeds)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study derives breed- and sex-specific genetic maps for six cattle breeds from SNP genotype data using a standardized quality-control pipeline followed by LINKPHASE3, an HMM-based approach that simultaneously phases haplotypes and estimates sex-specific recombination rates. Two additional methods (a deterministic multipoint approach via hsphase and a likelihood-based pairwise approach via hsrecombi) were applied for comparison, restricted to paternal half-sib families with at least 30 progeny. PCA and hierarchical clustering (L2 norm) were used to characterize map similarity across breeds and sexes, and a GWAS was performed on mean autosomal crossover count and intra-chromosomal allelic shuffling per parent to identify associated genomic regions. Accuracy and mean squared error were used to evaluate all three map-estimation methods in synthetic and semi-real simulation studies with 10 repetitions per scenario.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| HMM-based recombination rate estimation (LINKPHASE3) | Primary genetic map derivation for all six breeds, yielding sex-specific recombination rates and genetic-map coordinates | 4,181 to 76,875 total genotyped animals per breed (Table 1); two sequential runs with MCS ≥ 0.986 filtering and ≤58-crossover individual filter | not stated |
| Deterministic multipoint recombination estimation (R package hsphase v2.0.2) | Comparison with HMM-based method for male recombination rates in paternal half-sib families with ≥30 progeny | Restricted to n*_t animals per breed in N* qualifying families (Table 1); ranges from 506 to 57,462 | not stated |
| Likelihood-based pairwise recombination estimation (R package hsrecombi v1.0.1) | Comparison with HMM-based method for male recombination rates in paternal half-sib families with ≥30 progeny | Same restriction as hsphase: n*_t per breed (Table 1) | not stated |
| Principal components analysis (R package dimensio v0.14.0) | Comparison of breed- and sex-specific recombination maps across SNPs shared among all breeds | 12 map vectors (6 breeds × 2 sexes); number of shared SNPs not stated in excerpt | not stated |
| Hierarchical clustering (R function hclust, stats v4.4.0; L2 norm of genetic-map coordinate differences) | Grouping of sex- and breed-specific genetic maps by overall coordinate similarity | 12 map vectors (6 breeds × 2 sexes) | not stated |
| Genome-wide association study (GWAS) | Mean autosomal crossover count and intra-chromosomal allelic shuffling per parent; identified signals on BTA6 and BTA10 | null | not stated |
-
Map similarity across breeds and sexes was characterized using PCA on recombination rates and hierarchical clustering with L2 norm distance and an unspecified linkage method↳ Could also: Ward's minimum-variance linkage or complete linkage with Pearson correlation distance could also be used; multidimensional scaling (MDS) offers another standard low-dimensional embedding for the same distance matrix — Different linkage criteria and distance metrics can reveal different aspects of the similarity structure; Ward's linkage tends to produce balanced, compact clusters and is widely used in genomic applications, while correlation distance separates shape from scale differences in recombination profiles
-
Each simulation scenario was repeated 10 times to estimate accuracy (acc) and MSE of map estimation↳ Could also: A larger number of repetitions (e.g., 100–1000) with bootstrap confidence intervals around acc and MSE could also be used to quantify simulation uncertainty — More repetitions reduce Monte Carlo variance in performance metrics, allowing tighter uncertainty bounds and more reliable ranking of methods across scenarios, especially in high-variability settings such as small-N, small-n cases
-
Simulation evaluation used acc (ratio of estimated to true total map length) and MSE of individual map positions as quality measures↳ Could also: Pearson or Spearman correlation between estimated and true genetic-map positions, or the concordance correlation coefficient, could also quantify method agreement — Correlation-based metrics separate systematic bias (scale shift) from random error and can highlight whether deviations are uniformly distributed across the chromosome or concentrated in specific regions, complementing the information in acc and MSE
-
Breed differences in total map length and recombination patterns were summarized descriptively (ranges, PCA, clustering) without formal inferential tests↳ Could also: Bootstrap confidence intervals on breed-specific map lengths, or permutation tests for between-breed differences, could also quantify sampling uncertainty around observed differences — Uncertainty quantification distinguishes sampling variability from true breed differences, which is particularly relevant for smaller breeds (Angus: 4,181 animals; Limousin: 6,842) where estimates may be less stable than in larger breeds
-
Missing genotype data (up to ~48% in Holstein-CH) was accommodated within the LINKPHASE3 HMM framework after marker- and family-level QC filtering↳ Could also: Genome-wide imputation to a common high-density reference panel prior to map estimation could also be applied as a preprocessing step for all breeds, as was done for the Brown Swiss semi-real simulation using FImpute v3 — Pre-imputation can increase effective marker density and reduce the rate of systematically missing data, potentially improving map resolution and reducing estimation uncertainty, particularly for breeds genotyped on lower-density arrays
-
The GWAS for recombination activity traits was performed, with results described as showing 'strong evidence' and 'suggestive signals', implying threshold-based significance calling↳ Could also: A linear mixed model (e.g., GEMMA or BOLT-LMM) accounting for population structure and genomic relatedness could also be used; alternatively, a permutation-based genome-wide significance threshold could complement or replace a fixed threshold — Mixed-model GWAS controls for confounding due to breed stratification and family structure—both prominent features of this multi-breed half-sib dataset—which can reduce false-positive inflation; permutation-derived thresholds are empirically calibrated to the actual test-statistic distribution
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
scope.md — pmid-41942849
Paper: Wittenburg, Melzer, Abdollahi Sisi, Ding, Seefried (2026), A comparative study on recombination activity in cattle. Genet Sel Evol. DOI 10.1186/s12711-026-01041-0.
Data situation (important)
- Raw genotypes are RESTRICTED: "Genotype data are available from Qualitas AG (Zug) upon agreement." → the from-genotypes pipeline (PLINK QC → FImpute → LINKPHASE3 / hsphase / hsrecombi map estimation) CANNOT be re-run from scratch (no raw data).
- The brief's accession SRA PRJNA668863 is MIS-HARVESTED: it is an unrelated USDA-ARS long-read genome assembly of a single NZ Holstein bull, NOT this paper's SNP genotypes. → recorded as a data-link false positive; not usable.
- PROCESSED data IS public on Zenodo 10.5281/zenodo.17909700 ("Data derived from
sex-specific recombination rate analysis in cattle", 455 MB): per-breed
.Rdatawith the estimated genetic maps (per-marker cM positions for HMM / deterministic / likelihood methods, sex-specific), chromosome-levelgenetic_map_summary, adjacent-marker recombination rates, best map functions. This is exactly the data the CLARITY Shiny app (github.com/nmelzer/CLARITY, Zenodo 10.5281/zenodo.18668018) visualises. - GWAS scripts + data on Zenodo 10.5281/zenodo.17951753 (Table 4 hotspots).
Reproduction strategy (P16: run the shipped tool/data, verify deposited→reported)
The pipeline-derived summary results in Table 3 are deterministic functions of the deposited genetic maps. We recompute them from the public Zenodo processed data and compare to the printed paper values. This is a faithful 1:1 check of whether the deposited data reproduces the reported numbers.
IN SCOPE (recompute from Zenodo 17909700 processed maps)
- R1 — HMM-based total genetic map length (Morgan), male & female, per breed (Table 3). Sum chromosome-level HMM genetic length over autosomes BTA1–29. 6 breeds × 2 sexes = 12 values.
- R2 — cM:Mbp ratio range per chromosome (Results text): males ~0.9 (BTA1) to 1.7 (Brown Swiss BTA19); females ~0.8 (BTA1) to 1.5 (Angus BTA28).
- R3 — method comparison (Table 3): deterministic ≈9% shorter, likelihood ≈54% shorter than HMM (Limousin male: HMM 26.97 / det 24.52 / lik 12.51 M).
- R4 — number of shared SNPs across breeds = 34,755 (Results text).
- R5 — female shorter than male, up to 18% in Brown Swiss (derived from R1).
IN SCOPE if event-level data is deposited (per-breed zips)
- R6 — mean crossover count per gamete, sires & dams, per breed (Table 3, mean±SD).
Source:
recombination_traits.txt(2 MB) in Zenodo 10.5281/zenodo.17951753 — the per-gamete crossover/shuffling phenotypes. Mean/SD are deterministic (no ASReml needed).
OUT OF SCOPE — needs proprietary ASReml + non-public phenotypes
- R7 — GWAS hotspots (Table 4, BTA6/BTA10 RNF212/RNF212B p-values) and the GBLUP
heritabilities/genetic correlations: the deposited workflow (
Workflow_GBLUP_and_GWAS.R,random_gwas.as) drives ASReml 4.2 (proprietary, licensed); the deposit states the scripts "do not run without input files listed therein." The phenotype means (R6) are recomputable; the variance-component results are not, without ASReml. Recorded out-of-scope.
OUT OF SCOPE (cannot attempt)
- Raw-genotype QC/imputation/map estimation from scratch — raw data restricted (Qualitas AG).
- ASReml GBLUP heritabilities / genetic correlations — ASReml is proprietary/licensed AND needs per-animal phenotypes that are not in the public deposit; recorded out-of-scope.
- Simulation study (accuracy 1.01–1.05 etc.) — separate sim pipeline, attempt only if scripts shipped.
Pipelines named per result
- R1–R5: hsrecombi/LINKPHASE3 outputs (deposited) → simple summation in R (recompute & compare).
- R6: per-gamete recombination events (deposited per breed) → mean/SD.
- R7: GWAS (weighted LMM) from Zenodo 17951753.
Compute
All on «our HPC» («infra» workdir), R via conda pr
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Every attempted reported value reproduces exactly to 3 dp from the authors' two independent public Zenodo deposits — all 24 Table 3 map lengths, all 24 crossover mean/SD values, the non-obvious filtered 34,755 shared-SNP count, the 17.91% female shortening, the 67.5% likelihood shortening, and the cM:Mbp ranges with named chromosomes — with only sub-0.003 M rounding differences. The exact recovery of both mean AND SD of phenotypes from a file independent of the maps strongly argues the published values are faithfully derived (no fabrication indicators). The limitations are purely on data-availability/scope, not the authors' side: the raw genotypes are restricted (Qualitas AG) and the briefed SRA accession is mis-harvested, so the raw->map estimation step and the proprietary-ASReml GWAS/heritability results were out of scope. This is a clean deposited-data->reported-value confirmation; q1 is yellow only because the original raw input was unavailable.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.