Meta-analysis of six dairy cattle breeds reveals biologically relevant candidate genes for mastitis resistance.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
- Nothing in this column.
- 🔴Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH for the annotation step; core GWAS is data-restricted. The paper's central results (58 lead markers, 31 candidate genes, MR-MEGA/METAL/MTAG/MAGMA p-values) are NOT reproducible: all six breeds' genotype+phenotype are available only on reasonable request with breeding-company permission (Viking Genetics, Braunvieh Schweiz, DataGene, INRAE/Valogene, FBN, WUR); no public summary statistics exist -> data_restricted for the core. The smoove+duphold GC-CNV sub-analysis is public-tool but needs an un-accessioned 567-animal WGS cohort + tens-of-TB alignment -> out of 80/20. REPRODUCED 1:1 (P16, third-party tool on the paper's own reported variants): Ensembl VEP functional annotation of the lead-SNP (Table 1) and candidate-causal tables. On the public REST VEP (current Ensembl release) 47/69 variants reproduce EXACTLY; all 6 resolvable missense candidate-causal variants give the reported amino-acid change in the reported gene, and 3/6 SIFT scores match on the current release. The 18 mismatches are dominated by Ensembl annotation drift between the paper's pinned release-104 and the current release (newly-annotated ncRNA introns/regulatory regions, LOC618542->RBAK rename, v104-only gene ENSBTAG00000049290 retired). The version-faithful v104 offline-cache run is fully staged for «our HPC» (run.sbatch + resolve/compare scripts) but was NOT executed: it is gated on a phone-2FA VPN login (two fresh 2FA links expired before being tapped) and the operator chose to finalize. NOT ATTEMPTED: the GWAS meta-analysis itself, MAGMA/GARFIELD/MTAG numeric outputs, and the smoove CNV. No fabrication detected in the annotation columns checked (annotations are consistent with public gene models).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 69assessed: 2026-06-15 ⛓ 223a98e90a13
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator headless) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusA multi-breed meta-analysis of GWAS summary statistics for clinical mastitis (CM) and somatic cell score (SCS) across six dairy cattle breeds can increase power and precision to identify functional genetic variants and candidate genes affecting mastitis resistance in dairy cattle.
- ★ Meta-analysis of GWAS across multiple breeds for CM and SCS identified 58 lead markers associated with mastitis incidence, including 16 loci not overlapping previously identified QTL in AnimalQTLdb. finding
- ★ Post-GWAS analyses prioritized 31 candidate genes and 14 credible candidate causal variants affecting mastitis. finding
- ★ Combining single- and multi-trait meta-analysis methods that account for multi-breed structure increases power and precision to detect variants affecting mastitis-related traits. method
- ★ Putative causal genes were prioritized by integrating nearest-gene/gene-based analysis with GO, KEGG pathway, and mammalian phenotype database support. method
- The GC gene (Vitamin D-binding protein) on BTA6 (~88-89 Mb) is a plausible candidate for the recurrent CM/SCS QTL, with NPFFR2 as an additional potential causal gene. mechanism
- The candidate gene list helps elucidate the genetic architecture of mastitis resistance and supports breeding for improved resistance. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Genome-wide association study (GWAS) on imputed whole-genome sequence variants | Six dairy cattle breeds (Holstein, Jersey, Nordic Red/RDC, Montbéliarde, Normande, Brown Swiss/Original Braunvieh); Bos taurus | none | Association between sequence variants and clinical mastitis (CM) and somatic cell score (SCS) | GCTA-MLMA (mixed linear model) |
| Single-trait meta-analysis of GWAS summary statistics (fixed-effect) | Multi-breed dairy cattle GWAS datasets | none | Meta-analyzed association statistics for CM and SCS | METAL |
| Trans-ethnic meta-regression meta-analysis | Multi-breed dairy cattle GWAS datasets | none | Meta-analyzed association statistics accounting for allelic effect heterogeneity (MR-MEGA_CM, MR-MEGA_SCS) | MR-MEGA (4 PCs for CM, 12 PCs for SCS) |
| Multi-trait meta-analysis | Per-breed CM and SCS summary statistics | none | Combined multi-trait association statistics (MTAG_CM, MTAG_SCS) | MTAG |
| Gene-based analysis | Bovine GWAS summary statistics and gene location data | none | Gene-level association for candidate gene prioritization | MAGMA |
| Variant annotation | Significant sequence variants, ARS-UCD1.2 genome | none | Functional effect annotation of variants | Variant Effect Predictor (VEP) |
| Genomic feature enrichment analysis | Bovine GWAS variants | none | Enrichment of GWAS signals in genomic features / key variants | GARFIELD |
| Copy number variant (CNV) calling and quality control of summary statistics | Additional dataset; multi-breed summary statistics | none | CNV detection and QC metrics (allele frequency, lambda, MAF, imputation accuracy) | EasyQC |
- – 58 lead markers associated with mastitis incidence were identified by the meta-analyses 58 lead markers
- – 16 of the identified loci did not overlap with previously identified QTL in AnimalQTLdb 16 loci
- – 31 candidate genes were prioritized through post-GWAS analyses 31 genes
- – 14 credible candidate causal variants affecting mastitis were identified 14 variants
- – A recurrent QTL for CM and SCS was confirmed at ~88-89 Mb on BTA6 across many breeds/studies 88-89 Mb
- count 30,689 (Number of animals/records with phenotypes for clinical mastitis (CM))
- count 119,438 (Number of animals/records with phenotypes for somatic cell score (SCS))
- count 8 GWAS for CM and 14 GWAS for SCS (Number of GWAS combined in the meta-analyses)
- other −log10(p) > 8.5 (Significance threshold for single- and multi-trait meta-analysis)
- count 1869 QTL for SCS and 569 QTL for CM (QTL reported in AnimalQTLdb for the two traits)
- other 0.02 to 0.12 (Heritability range of clinical mastitis (CM))
- other 0.10 to 0.15 (Heritability range of somatic cell score (SCS))
- correlation 0.24 to 0.55 (Genetic correlation of CM with milk yield in Nordic dairy cattle)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper conducted sequence-level GWAS within each of six dairy cattle breeds using mixed linear models (GCTA-MLMA with a genomic relationship matrix), then meta-analysed summary statistics across eight CM and fourteen SCS GWAS using two complementary approaches: MR-MEGA (trans-ethnic meta-regression with principal components to accommodate allelic-effect heterogeneity across breeds) and METAL (fixed-effect inverse-variance-weighted). A multi-trait analysis (MTAG, per breed) followed by MR-MEGA was additionally performed to leverage the genetic correlation between CM and SCS. Post-GWAS analyses included MAGMA gene-based testing, VEP variant annotation, and GARFIELD genomic feature enrichment, with results integrated against CattleGTEx expression data to nominate candidate causal genes and variants; a uniform genome-wide threshold of −log10(p) > 8.5 was applied across all outputs.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Mixed linear model (GCTA-MLMA): y = 1μ + bx + g + e, with genomic relationship matrix modelling polygenic background | Within-breed single-variant GWAS for CM and SCS at all seven contributing institutions | 30,689 total (CM); 119,438 total (SCS) across all breeds; per-breed n not stated in text | not stated |
| Trans-ethnic meta-regression with principal components (MR-MEGA) | Primary meta-analysis combining 8 CM GWAS and 14 SCS GWAS across six breeds; also applied after MTAG for multi-trait outputs | 30,689 (CM); 119,438 (SCS) | not stated |
| Fixed-effect inverse-variance-weighted meta-analysis (METAL, STDERR method) | Secondary meta-analysis of CM and SCS summary statistics | 30,689 (CM); 119,438 (SCS) | not stated |
| Multi-trait meta-analysis (MTAG) followed by MR-MEGA combination across breeds | Joint analysis of CM and SCS per breed, then combined across breeds; outputs MTAG_CM and MTAG_SCS | 30,689 (CM); 119,438 (SCS) | not stated |
| Gene-based association test (MAGMA) | Post-GWAS prioritisation of candidate genes from meta-analysis summary statistics | — | not stated |
| Genomic feature enrichment analysis (GARFIELD) | Post-GWAS enrichment of significant variants in functional genomic annotations | — | not stated |
-
Fixed-effect inverse-variance-weighted meta-analysis (METAL) was used alongside MR-MEGA to combine within-breed GWAS results↳ Could also: A random-effects meta-analysis (e.g., DerSimonian–Laird or REML-based) could also have been applied — When allelic effects are expected to differ across genetically diverse breeds, a random-effects model explicitly quantifies between-study heterogeneity in effect size and yields confidence intervals that reflect that uncertainty; this would complement the heterogeneity statistics already available within MR-MEGA
-
The genome-wide significance threshold was set uniformly at −log10(p) > 8.5 across all analyses↳ Could also: A Bonferroni correction anchored to the effective number of independent sequence-level variants, or a permutation-based empirical threshold, could also have been used — Deriving the threshold empirically from the actual LD structure and total variant count of the dataset would formally calibrate the family-wise error rate and make the chosen threshold reproducible and transparent to readers
-
Post-GWAS fine-mapping relied on VEP functional annotation and GARFIELD enrichment to prioritise credible causal variants↳ Could also: Bayesian statistical fine-mapping methods such as SuSiE or FINEMAP applied to summary statistics with an LD reference could also have been used — These methods produce posterior inclusion probabilities and credible sets that quantify per-variant uncertainty about causality in a statistically principled way, complementing the annotation-based prioritisation already performed
-
Gene-based testing was performed using MAGMA on the GWAS summary statistics↳ Could also: Colocalization analysis (e.g., coloc) or a transcriptome-wide association study (TWAS) using the CattleGTEx eQTL data already accessed in the study could also have been applied — Coloc and TWAS explicitly test whether a GWAS signal and a cis-eQTL in a relevant tissue share a causal variant, providing a mechanistic link between the association signal and gene expression that p-value aggregation in MAGMA alone does not establish
-
Summary statistics from phenotypes defined in different ways across institutions (DRP, DYD, EBV with varying reliability weights) were combined without explicit modelling of phenotype-definition heterogeneity↳ Could also: A meta-regression model that includes phenotype type (DRP vs. DYD vs. EBV) as a moderator covariate could also have been applied — Different deregression methods and reliability weights may introduce systematic differences in effect-size scale and standard-error calibration across studies; including this as a covariate would allow assessment of how much it contributes to the heterogeneity observed in MR-MEGA
-
Within-breed GWAS used a single-variant mixed linear model (GCTA-MLMA) treating each SNP independently↳ Could also: Bayesian whole-genome regression methods (e.g., BayesR or BayesC) could also have been applied within each breed before meta-analysis — Bayesian approaches simultaneously model all variants and can have greater power for highly polygenic traits with many small-effect loci; the resulting per-variant posterior effect estimates could serve as an alternative or complementary input to the meta-analysis stage
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
A recurrent QTL for clinical mastitis and somatic cell score is confirmed at 88–89 Mb on BTA6 across multiple breeds and studies.other bos-taurus multi-breed dairy cattle 2024×1papers★ This paper is the founder (earliest)
-
58 lead sequence variants associated with clinical mastitis and somatic cell score were identified by multi-breed meta-GWAS.other bos-taurus multi-breed dairy cattle 2024×1papers★ This paper is the founder (earliest)
-
31 candidate genes for mastitis resistance were prioritized through post-GWAS gene-based and annotation analyses.other bos-taurus multi-breed dairy cattle 2024×1papers★ This paper is the founder (earliest)
-
14 credible candidate causal sequence variants affecting mastitis resistance were identified via fine-mapping.other bos-taurus multi-breed dairy cattle 2024×1papers★ This paper is the founder (earliest)
-
16 mastitis-associated loci identified by meta-GWAS do not overlap with previously reported QTL in AnimalQTLdb.other bos-taurus multi-breed dairy cattle 2024×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-39009986
Paper: Cai et al. 2024, Genet Sel Evol 56:54. "Meta-analysis of six dairy cattle breeds reveals biologically relevant candidate genes for mastitis resistance." DOI 10.1186/s12711-024-00920-8 · PMCID PMC11247842.
What the paper does (pipeline map)
A multi-breed (Holstein, Jersey, Nordic Red, Brown Swiss, Montbéliarde, Normande) sequence-based GWAS meta-analysis for clinical mastitis (CM, 30,689 animals) and somatic cell score (SCS, 119,438 animals). Each partner ran a within-breed GWAS (GCTA-MLMA) on WGS-imputed genotypes; summary statistics were QC'd (EasyQC) and combined by MR-MEGA (meta-regression), METAL (fixed-effect) and MTAG (multi-trait). Downstream: MAGMA gene-based test, GARFIELD functional enrichment, VEP v104 variant annotation, PLINK LD, and a smoove + duphold structural-variant sub-analysis on chr6 WGS.
In scope vs out of scope
| Reported result | Pipeline | In scope? | Why |
|---|---|---|---|
| 58 lead markers / QTL coordinates & −log10(p) (Table 1/2) | GCTA-MLMA → MR-MEGA/METAL/MTAG | OUT | Inputs are per-breed genotype+phenotype/summary-stats; all six restricted (on-request + breeding-company permission). No public sumstats deposit. |
| MAGMA gene-based, GARFIELD enrichment, MTAG novel signals | MAGMA/GARFIELD/MTAG | OUT | Same restricted GWAS inputs. |
| 12 kb CNV at BTA6:86,949,652–86,961,433 near GC | Trimmomatic→bwa→GATK→smoove→duphold (chr6) | OUT (hard >20%) | Tool public; data public in principle (1000 BGP, PRJNA431934 +19 BioProjects) BUT the exact 567-animal cohort is not individually accessioned, requires tens-of-TB WGS download + whole-genome alignment + joint calling. Infeasible in scope; and the result is itself a confirmation of a prior study's CNV, not a fresh pinnable number. |
| Functional annotation of lead SNPs + candidate causal variants (Tables 1, 2 & "candidate causal mutations"): consequence type, nearest gene, missense AA change + SIFT score | Ensembl VEP v104 on ARS-UCD1.2 | IN ✅ | Tool (VEP v104) and reference (Ensembl release-104 bos_taurus cache) are public; inputs are the rsIDs + ARS-UCD1.2 coordinates printed in the paper's own tables; compute is minutes. Re-annotating these variants is a clean, deterministic 1:1 check (and a fabrication probe on the reported gene/consequence/SIFT columns). |
Reproduction target (the "few clear data points")
Re-run Ensembl VEP v104 (offline cache, --sift b --symbol --nearest symbol)
on the lead SNPs (Table 1, by rsID/coordinate) and the candidate causal mutations
(Table "candidate causal mutations", the missense rows carry SIFT scores), and
compare 1:1 against the paper's reported Annotation / nearest-gene / SIFT columns.
P16 note: VEP is a third-party tool; applying it to the paper's reported variants is an equally valid reproduction of the paper's annotation step. We do NOT and cannot reproduce the GWAS that produced the variants (restricted data).
Honest non-attempt list
- The GWAS meta-analysis itself (restricted genotype/phenotype across 6 partners).
- The smoove/duphold CNV (cohort not accessioned; tens-of-TB WGS, out of 80/20).
- MAGMA/GARFIELD/MTAG numeric outputs (restricted inputs).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.