Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Improving genomic prediction accuracy for methane emission and feed efficiency in sheep: integrating rumen microbial PCA with host genomic variation using neura

Genet Sel Evol · 2025
L1 85/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for a clean 1:1 on the headline numbers. The GitHub repo is an R Shiny RESULTS-VIEWER, not the genomic-prediction pipeline; it ships precomputed result tables, and the phenotype+genotype data are RESTRICTED (only raw microbiome reads are public on SRA PRJNA859547). I reproduced the in-scope, clearly-specified results EXACTLY: the Train-Test genomic-prediction accuracies and dispersion biases (Table S1) were recomputed from the shipped per-animal predictions as accuracy=Pearson cor(Actual,Predicted) and bias=slope lm(Actual~Predicted). All 15 (trait x model) cells matched to 3 decimals (max accuracy diff 0.000, max bias diff 0.01), including the abstract's headline methane improvement G=0.114 -> NN-GBLUP(PC88)=0.301 and RFI best 0.396. Compute ran on «our HPC» («job»: clone+recompute on a compute node, cwd on «infra»; nothing on «host»). NOT attempted (the hard 20%, honestly out of scope): (1) microbiability REML estimates (~0.50 methane / ~0.71 RFI) - need restricted genotypes; shipped estimates only consistency-checked vs paper (0.496/0.707); (2) five-fold CV per-animal accuracies (Table S2) - no per-animal five-fold file shipped, so not independently recomputable; (3) microbiome PCA and the SRA-reads->tag pipeline - no pipeline code in the repo. Overall: partial, with the reproducible core graded exact. No fabrication detected for in-scope claims (all derivable from shipped data). Provisional - a human reviewer decides via AUDIT.md.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 85
    assessed: 2026-06-14 ⛓ 5109711f01bd
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether reducing the dimensionality of rumen microbiome composition (RMC) data via PCA and incorporating it as an intermediate trait in a Neural Network GBLUP (NN-GBLUP) model can improve genomic prediction accuracy for methane emissions and feed efficiency traits in sheep, compared to standard GBLUP.

Core claims
  • PCA effectively reduces the dimensionality of high-dimensional RMC data while retaining essential biological information. method
  • NN-GBLUP incorporating PCA-reduced RMC data as an intermediate trait improves genomic prediction accuracy for methane emissions and residual feed intake compared to standard GBLUP. finding
  • Microbiability estimates derived from PCA components explaining 100% of RMC variation closely match estimates from the full RMC dataset, validating the dimension reduction. finding
  • The optimal number of PCA components is trait-dependent, with components explaining 25% and 50% of variation giving the best prediction accuracy, while 75% and 95% reduce accuracy for methane traits. finding
  • Prediction accuracy did not improve for carbon dioxide emissions, live weight, and mid-trial intake, indicating trait-dependent microbiome influence. finding
  • A reference-free RMC profiling pipeline yields substantially higher genomic prediction accuracy than a reference-based approach. resource
Experimental setups
Assay System Perturbation Readout Platform
Genomic prediction (NN-GBLUP vs GBLUP) Sheep, Methane Grass group (grass lambs) none Prediction accuracy for scaled CH4, CH4Ratio, CO2, live weight NN-GBLUP / GBLUP models
Genomic prediction (NN-GBLUP vs GBLUP) Sheep, RFI Lucerne group (lucerne-fed lambs) none Prediction accuracy for residual feed intake and mid-trial intake NN-GBLUP / GBLUP models
Rumen microbiome sequencing (restriction enzyme-reduced representation sequencing) Rumen contents from sheep (both groups) none Microbial tag counts (rumen microbiome composition profile) Reference-free pipeline (Hess et al.)
SNP genotyping Sheep (both groups) none SNP genotypes imputed to 600K, subset to 15K panel OvineSNP50 / ISGC_SheepLD15K chips
Methane measurement Sheep, Methane Grass group none CH4, CH4Ratio, CO2 emissions, live weight Portable accumulation chambers (PAC)
Feed intake measurement Sheep, RFI Lucerne group diet transition to lucerne pellets Daily feed intake, derived RFI and mid-trial intake Feed intake facility (AgResearch Invermay)
Principal component analysis (dimensionality reduction) RMC data, both groups none Cumulative variance explained by PCs (25/50/75/95/100% thresholds) R
Microbiability estimation (variance component model) Sheep, both groups none Proportion of phenotypic variance attributable to RMC
Key results
  • Methane emission prediction accuracy increased using PCA components explaining 25% of RMC variation in train-test validation 0.09 to 0.30
  • Methane emission prediction accuracy increased in five-fold cross-validation 0.15 to 0.27
  • Residual feed intake prediction accuracy increased in train-test validation 0.25 to 0.37
  • Residual feed intake prediction accuracy increased in cross-validation 0.25 to 0.34
  • Microbiability estimates from PCs explaining 100% of variation closely matched full-dataset estimates
  • Higher PCA components (75% and 95%) reduced prediction accuracy for methane traits relative to 25%/50%
  • No improvement in prediction accuracy for CO2 emissions, live weight, and mid-trial intake
  • Reference-free RMC profiling improved prediction accuracy compared to reference-based approach twofold
Key statistics
  • correlation 0.09 to 0.30 (Methane emission prediction accuracy, train-test validation, PCA components explaining 25% of RMC variation (train n=690, test n=361))
  • correlation 0.15 to 0.27 (Methane emission prediction accuracy, five-fold cross-validation, Methane Grass group)
  • correlation 0.25 to 0.37 (Residual feed intake prediction accuracy, train-test validation (train n=587, test n=397))
  • correlation 0.25 to 0.34 (Residual feed intake prediction accuracy, five-fold cross-validation, RFI Lucerne group)
  • count 240,743 (Number of microbial tags identified in the Methane Grass group)
  • count 207,393 (Number of microbial tags identified in the RFI Lucerne group)
  • count 11,967 (Number of quality-controlled SNPs in the 15K genotype panel used for main analyses)
  • other 0.16 to 0.25 (methane), 0.11 to 0.3 (RFI) (Literature-reported heritability ranges for methane traits and residual feed intake, cited as background)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

Two sheep cohorts (Methane Grass, n=1051; RFI Lucerne, n=984) with SNP genotypes and high-dimensional rumen microbiome composition (RMC) data (240,743 and 207,393 microbial tags, respectively) were analysed. PCA was applied to RMC data at variance thresholds of 25%, 50%, 75%, 95%, and 100%, and microbiability estimates from PCA-reduced subsets were compared to full-dataset estimates to validate dimensionality reduction. Genomic prediction accuracy — reported as Pearson correlations between predicted and observed fixed-effect-adjusted phenotypes — was evaluated for three models (GBLUP; GBLUP with RMC as an additional independent random effect; NN-GBLUP with PCA-reduced RMC as intermediate traits) across both a year-based train-test split and a cohort-stratified five-fold cross-validation.

Replicationbiological Sample sizeTotal n=1051 (Methane Grass) and n=984 (RFI Lucerne), born 2014–2016 across three New Zealand flocks, 863 animals common to both groups; train-test split: 690 train / 361 test and 587 train / 397 test; five-fold CV: approximately 203–218 per validation fold (Methane Grass) and 195–199 (RFI Lucerne); no formal a priori power analysis mentioned GroupsGBLUP vs. microbiome-augmented GBLUP vs. NN-GBLUP, across five PCA variance thresholds, for six phenotypic traits in two cohorts Pairingna Randomization/blindingstated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Genomic Best Linear Unbiased Prediction (GBLUP) — genomic relationship matrix-based mixed model Baseline genomic prediction for all traits (CH4, CH4Ratio, CO2, LW, RFI, mid-intake) 1051 (Methane Grass), 984 (RFI Lucerne) not stated
GBLUP with additional independent rumen microbiome composition random effect (microbiome-augmented GBLUP) Second comparison model, fitting RMC as a non-genetic random term alongside the genomic effect 1051 (Methane Grass), 984 (RFI Lucerne) not stated
Neural Network GBLUP (NN-GBLUP) with PCA-reduced RMC as intermediate traits (Bayesian neural network mixed model) Primary novel model evaluated across all PCA variance thresholds and all traits 1051 (Methane Grass), 984 (RFI Lucerne) not stated
Mixed model variance component estimation for microbiability Objective 1: validation of PCA dimensionality reduction by comparing microbiability from PCA subsets vs. full RMC data 1051 (Methane Grass), 984 (RFI Lucerne) not stated
Principal Component Analysis (PCA) on standardised microbial tag counts Preprocessing step applied separately to Methane Grass and RFI Lucerne RMC data across the combined train+test population 1051 (Methane Grass), 984 (RFI Lucerne) not stated
Pearson correlation between predicted and observed adjusted phenotypes (prediction accuracy metric) Performance metric reported for all models, traits, and validation schemes 361 test (Methane Grass) / 397 test (RFI Lucerne) for train-test; approximately 203–218 per fold (Methane Grass) and 195–199 per fold (RFI Lucerne) for cross-validation not stated
Approaches that could also have been used
  • Model comparisons were based on point estimates of prediction accuracy (Pearson correlation) without confidence intervals or formal tests of the difference between correlations
    Could also: Bootstrap confidence intervals around each correlation, or a Steiger test / permutation test of the difference between two dependent correlations evaluated on the same animals, could also be used — Confidence intervals would quantify estimation uncertainty around each accuracy value, and a formal test of the difference would allow readers to assess whether observed gains (e.g., 0.09 → 0.30 for CH4) exceed what could arise from sampling variation, especially relevant given test-set sizes of roughly 360–400 animals
  • PCA was fitted on the combined train and test RMC data before the train-test split was applied
    Could also: PCA could also be fitted exclusively on the training partition and applied to the test partition as a fixed linear projection — Fitting PCA on the full dataset means test-set observations contribute to defining the principal components; a strictly train-only PCA projection is a more conservative approach that eliminates any indirect use of test-set information in the feature construction step, and is common practice in predictive pipelines
  • Linear PCA was the sole dimensionality-reduction method evaluated for the RMC data
    Could also: Partial least squares regression (PLSR) against the trait phenotype, sparse PCA, or kernel PCA could also reduce the RMC feature space — Linear PCA maximises retained total RMC variance irrespective of phenotypic relevance; supervised alternatives such as PLSR orient the reduction toward trait-associated variance and may select fewer components while retaining or improving predictive signal
  • Microbiability from the 100% PCA solution was compared to the full-data microbiability as a single scalar validation of dimensionality reduction
    Could also: Reconstruction error (e.g., Frobenius norm between original and PCA-reconstructed microbiome matrices) reported across all five variance thresholds could also be used, together with confidence intervals on the microbiability difference — A single point comparison at the 100% threshold confirms equivalence there but does not quantify how much information is lost at intermediate thresholds (25%, 50%, 75%); a reconstruction-error curve across thresholds would give a more complete picture of the information-retention trade-off
  • The 600K imputed SNP set was reduced to 15K for main analyses based on a preliminary comparison showing similar accuracy across densities, but those preliminary results were not presented
    Could also: Reporting prediction accuracy for all three SNP densities (15K, 50K, 600K) within the main results table could also be done — Showing the density comparison in the main results would allow readers to directly evaluate the equivalence claim and assess whether the choice of 15K affects conclusions about NN-GBLUP versus GBLUP gains
  • Cross-validation performance was reported as a single accuracy value (presumably averaged over five folds), without reporting fold-level variability
    Could also: Mean ± SD of prediction accuracy across the five folds, or the range, could also be reported alongside the mean — Fold-level variance quantifies the stability of predictions across cohort subsets and is informative when fold sizes are modest (~200 animals); high fold-to-fold variation would add context to the overall accuracy estimates
  • Traits without accuracy improvement (CO2, live weight, mid-intake) were noted descriptively but not compared statistically to traits that did improve
    Could also: A formal test of the difference between correlations (e.g., Meng-Rosenthal-Rubin test for comparing dependent correlations) could also be applied across trait-model combinations — Formal testing of trait-specific model improvements would distinguish signal from sampling variability and help characterise which traits reliably benefit from microbiome integration versus those where the null of no improvement cannot be rejected
Software: R

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
4
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40676527

Paper: Alemu et al. 2025, Genet Sel Evol 17:... "Improving genomic prediction accuracy for methane emission and feed efficiency in sheep: integrating rumen microbial PCA with host genomic variation using neural networks." DOI 10.1186/s12711-025-00987-x · PMCID PMC12273308.

Code: https://github.com/AgResearch/genomics-prediction-app (R Shiny app, commit 5701c17a7b540d32f7d30a974b0d89d7d7b197e3, 2025-06-29).

What the artifact actually is

The repo is not the genomic-prediction pipeline. It is an R Shiny results viewer (app.R + R/{data_loading,plotting_functions,ui_components}.R) that reads precomputed result tables shipped under data/ and renders the paper's figures/tables. The upstream pipeline that produced those tables (microbiome tag → taxonomy → PCA; REML microbiability; GBLUP / GM / NN-GBLUP genomic prediction; cross-validation) is not in the repo.

Data availability

  • Raw rumen-microbiome reads: NCBI SRA PRJNA859547 (public, but only the sequence reads).
  • Phenotypes + genotypes (15K SNP, 11,967 QC SNPs): RESTRICTED — "available from Sheep Improvement Limited or AgResearch ... restrictions apply ... not publicly available" (Data availability statement).

In scope (reproducible) vs out of scope

Reported result Pipeline Status
Genomic-prediction accuracy + dispersion bias, Train–Test validation (Table S1; Figs) — G, NN-GBLUP at PC levels, per trait accuracy = Pearson cor(Actual,Predicted); bias = slope lm(Actual~Predicted) IN SCOPE — recomputable from shipped per-animal predictions (data/supplementarymaterial/Scatterplot/*.csv)
Five-fold CV accuracies (Table S2) same metric partial — summary table shipped & matches paper, but no per-animal five-fold file shipped → cannot independently recompute
Microbiability (REML, data/Microbiability/*.csv) ASReml/REML on G+M kernels out of scope — needs restricted genotypes; only the precomputed estimates are shipped (consistency-checked vs paper only)
Microbiome PCA variance explained (data/PCA/*.csv) PCA of tag table out of scope — needs restricted/large microbiome tag matrix; only precomputed variance tables shipped
Microbiome tag generation from SRA reads bespoke (not in repo) out of scope — no pipeline code; genotype/pheno restricted anyway

Reproduction target (the 80%)

Recompute, for every (trait, PC-level) group in the shipped Train–Test per-animal prediction files, the genomic-prediction accuracy (Pearson r of Actual vs Predicted) and dispersion bias (regression slope), and compare 1:1 against the reported values in table_s1.csv / final_combined/TrainTest_Combined.csv and the paper. This tests whether the headline reported numbers are actually derivable from the shipped data (a fabrication/auditability check) — equally valid as a third-party-tool reproduction (P16).

The hard 20% (not attempted, and why)

Full pipeline from SRA reads → microbiome PCA → REML microbiability → trained GBLUP/GM/NN-GBLUP predictions is not attempted: (a) no pipeline code in the repo, (b) phenotype/genotype data are restricted (data_restricted). We do not fabricate those steps.

Figures / tables: Table
C1
Reported
0.114
Reproduced
0.114
exact
C2
Reported
0.301
Reproduced
0.301
exact
C3
Reported
0.216
Reproduced
0.216
exact
C4
Reported
0.298
Reproduced
0.298
exact
C5
Reported
0.396
Reproduced
0.396
exact
C6
Reported
1.70
Reproduced
1.70
exact
C7
Reported
Table S1 (15 cells)
Reproduced
max accuracy diff 0.000, max bias diff 0.01
exact
C8
Reported
0.50
Reproduced
0.496 (shipped estimate, not recomputed)
partial
C9
Reported
0.71
Reproduced
0.707 (shipped estimate, not recomputed)
partial
C10
Reported
0.267
Reproduced
0.267 (shipped table only, no per-animal file)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

96.3 k
tokens (I/O) · 5.9 M incl. cache
11 min
runtime · 0 CPU-h
0 GB
peak RAM
1
HPC jobs
hummel
machine