Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

On the holobiont 'predictome' of immunocompetence in pigs.

Genet Sel Evol · 2023
L1 40/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
40/100
Reproducibility score
1.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 4% of all assessed papers rank 1126 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

RAN TO COMPLETION (re-run of a previously requeued room). The one in-scope, fully-specified, public-data, pipeline-derived target -- the batch-2 (MiSeq, 400 samples) 16S DADA2 ASV count -- was computed end-to-end on «our HPC» («job», node n034, 1h51m) using the EXACT Additional-file-1 pipeline (QIIME2 2021.8 dada2 denoise-paired, trim-left-f 17, trim-left-r 21, trunc-len-f/r 250). RESULT = 11,867 ASV across 400/400 samples (16,782,904 reads) vs the paper's reported 2,566 -> MISMATCH (4.6x). The shipped script applies NO post-DADA2 feature filter to batch 2 (the samples2keep.txt filter is batch-1-only), so 2,566 should equal the raw DADA2 output; a diagnostic over the reproduced feature table (reproduction/outputs/asv_distribution_diagnostic.json) shows no documented or standard filter recovers 2,566 (singleton removal->11,345; total-freq>=10->9,453; prevalence>=20 samples->2,304, the closest but undocumented). Not a stochastic/version effect (QIIME2 2021.8 pinned per the paper; DADA2 is deterministic; thread count does not change ASVs). FLAGGED for human audit as a value not derivable from the shipped data+code as described (brief rule 5); benign explanations to weigh: SRA-deposited reads differing from the authors' exact input, an undocumented prevalence filter, or a genuine discrepancy -- a human signs off. NOT attempted (correctly out of reach): batch-1/merged ASV + merged reads/sample (need an unshipped run->pig map + samples2keep.txt), and ALL genotype/prediction headline results (host SNP genotypes + immune phenotypes were never deposited in PRJNA608629 -- which holds only 16S amplicon, verified via ENA API -- and the code link github.com/gdlc/BGLR-R is the generic BGLR R package, not the authors' analysis scripts). Engineering notes for the record: three «our HPC» jobs -- 2220354 died at conda-activate under set -u (fixed via conda run -p), 2220610 hit OUT_OF_MEMORY at the 125 GiB cgroup cap (fixed by taking a full 768 GiB node + bounding DADA2 to 32 threads), 2220639 completed. The shared «user» «infra»/home quota was exhausted by concurrent rooms, so the job runs all bulk work (reads + QIIME2 env + DADA2) on node-local /tmp (750 GB) and persists only the small ASV result to «infra».

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-15 ⛓ 18181753ae71
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether integrating host genotype and gut microbiota data (a 'holobiont' model) improves predictive accuracy for immunocompetence traits in pigs compared to genotype-only or microbiome-only models, and how different modelling strategies (statistical model, priors, microbial clustering) affect this prediction.

Core claims
  • Holobiont (combined genotype + microbiome) models performed better than partial models (genotype-only or microbiome-only) overall finding
  • Host genotype was especially relevant for predicting adaptive immunity traits (IgM and IgG concentrations) finding
  • Microbial composition was important for predicting innate immunity traits (haptoglobin, C-reactive protein, lymphocyte phagocytic capacity) finding
  • No single model was uniformly best across all six immunocompetence traits finding
  • Greater variability in predictive accuracies across models was observed when microbiability (variance explained by the microbiome) was high finding
  • Clustering microbial abundances (by phylogeny or by abundance) did not necessarily increase predictive accuracy finding
  • Host genome and gut microbiome jointly ('hologenome') act on complex traits including immunocompetence mechanism
  • A large catalogue of model choices ('predictome') combining RKHS, Bayes C, an ensemble method, varying priors, and clustering strategies was built to evaluate holobiont prediction method
Experimental setups
Assay System Perturbation Readout Platform
SNP genotyping (Porcine 70k GGP HD Array) 400 Duroc pigs none SNP genotype (coded 0/1/2 for AA/AB/BB) Illumina Infinium HD Assay Ultra
16S rRNA (V3-V4) amplicon metagenomic sequencing fecal samples from Duroc pigs none ASV relative abundance (CLR-transformed) Illumina NovaSeq (2x250) and MiSeq (2x300); QIIME2 v2021.8/dada2
Immunoglobulin concentration assay plasma of Duroc pigs none plasma IgM and IgG concentration
Acute-phase protein assay serum of Duroc pigs none serum haptoglobin (HP) and C-reactive protein (CRP) concentration
Phagocytosis assay lymphocytes from Duroc pigs none phagocytic capacity (LYM_PHAGO_FITC)
Immune cell subset quantification blood of Duroc pigs none percentage of gamma-delta T cells
Genomic/microbial prediction modelling (RKHS, Bayes C, ensemble) 400 Duroc pigs (genotype and/or microbiome data) model type, priors, feature set predictive accuracy (correlation between predicted and observed phenotype) BGLR R package
Phylogenetic and abundance-based clustering of ASV 16S ASV sequences from Duroc pig faecal samples clustering method (phylogeny vs abundance) number and composition of microbial clusters (k=232) MUSCLE, MEGA v11, QIIME2, R stats/hclust package
Key results
  • Holobiont models had better predictive accuracy than partial (genotype-only or microbiome-only) models
  • Host genotype was especially relevant for predicting IgM and IgG (adaptive immunity)
  • Microbial composition was important for predicting HP, CRP, and LYM_PHAGO_FITC (innate immunity)
  • No model was uniformly best across all six traits
  • Greater variability in predictive accuracies across models occurred when microbiability was high
  • Clustering microbial abundances did not necessarily increase predictive accuracy
  • SNP quality control retained 41,131 of 68,516 genotyped SNPs
  • ASV quality control retained 2,945 of 57,195 raw ASV for the microbial relationship matrix
Key statistics
  • count 400 (Duroc pigs used in the study (199 females, 201 males))
  • count 68,516 SNPs genotyped; 41,131 retained after QC (SNP genotyping and quality control with Plink v1.9)
  • count 0.19% (missing genotype rate, imputed with average allele frequency per SNP)
  • count 57,195 raw ASV filtered to n_B = 2,945 retained (ASV quality control filtering (present in <3 samples or <0.001% of total counts discarded))
  • count k = 232 (number of phylogeny-based and abundance-based microbial clusters generated)
  • count 22 boars and 132 sows (parents of the 400 Duroc piglets sampled)
  • other MAF < 5%; missing genotype data > 10% (SNP exclusion thresholds applied in Plink quality control)
  • other π0 = 0.5, 0.01, 0.001, 0.0001 (SNPs); 0.5, 0.1, 0.01, 0.001 (ASV); 0.5, 0.1, 0.01 (clusters) (Bayes C prior probabilities tested for feature inclusion)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This prediction study compared a catalogue of Bayesian regression models—Bayes C and RKHS (collectively called the 'predictome')—plus an ensemble approach, for predicting six porcine immunocompetence traits in 400 Duroc pigs genotyped for ~41k SNPs and gut-microbiome-profiled via 16S sequencing. Three data configurations were contrasted: genotype-only, microbiome-only, and combined holobiont models, each evaluated under multiple prior settings and two microbial feature representations (individual ASV and two clustering strategies). Predictive accuracy was defined as the Pearson correlation between predicted and observed pre-adjusted phenotypes, assessed via cross-validation.

Replicationbiological Sample size400 pigs (199 females, 201 males), a subset of 432 from a prior study; offspring of 22 boars and 132 sows; no formal power analysis or sample-size justification stated Groupsholobiont (genotype + microbiome) vs genotype-only vs microbiome-only prediction models, across six immunocompetence traits Pairingna Randomization/blindingnot stated Dispersionunclear Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Bayes C (spike-and-slab Bayesian regression, BGLR implementation) All six immunocompetence traits; genotype-only, microbiome-only, and holobiont model variants with varying exclusion prior π0 400 not stated
Reproducing Kernel Hilbert Space (RKHS) Bayesian kernel regression (BGLR implementation) All six immunocompetence traits; genotype-only, microbiome-only, and holobiont model variants with informative and REML-like (uninformative) variance-component priors 400 not stated
Ensemble method (details not provided in available text excerpt) All six immunocompetence traits (noted in abstract only) 400 not stated
Approaches that could also have been used
  • Phenotypes were pre-adjusted for fixed environmental effects (batch, sex) before cross-validation, and these residuals were the target of prediction
    Could also: Environmental covariates could also be included directly within the prediction model inside each cross-validation fold — Pre-adjusting outside the cross-validation loop may allow information from validation-fold individuals to influence the correction step; including covariates inside the loop is one way to avoid that information pathway and obtain accuracy estimates with a cleaner separation between training and validation data
  • Centered log-ratio (CLR) transformation was applied to raw ASV abundances to address compositionality
    Could also: Isometric log-ratio (ILR) transformation or Aitchison-distance-based kernels could also handle the compositional constraint — ILR preserves the full Aitchison geometry without requiring a reference category; Aitchison-distance kernels can be plugged directly into RKHS models and offer an alternative way to encode compositional structure without an explicit feature-space transformation
  • Predictive accuracy was quantified as the Pearson correlation between predicted and observed phenotypes
    Could also: Root mean squared error (RMSE), mean absolute error (MAE), or Lin's concordance correlation coefficient (CCC) could also characterize prediction performance — Correlation is scale-free and can remain high even with systematic bias; RMSE and CCC additionally quantify the magnitude and location of prediction error, which can be informative for assessing practical utility in a breeding context
  • Microbial features were summarized either as individual ASV or as hard clusters derived from hierarchical algorithms (phylogeny-based UPGMA and abundance-based Ward's), both at k=232
    Could also: Ordination-based dimensionality reduction (e.g., PCoA on Aitchison or UniFrac distances) or probabilistic latent-factor models could also produce low-dimensional microbial summaries — PCoA and latent-factor representations capture continuous community variation rather than discrete cluster assignments, which may better reflect ecological gradients and can be directly incorporated as features or kernels in the same BGLR framework
  • The exclusion prior π0 in Bayes C was explored over a predefined grid of values for each feature type (SNPs and ASV/clusters) rather than estimated from data
    Could also: π0 could also be treated as a hyperparameter with its own prior and estimated within the MCMC sampler, as in the Bayes Cπ extension — Data-adaptive estimation of π0 removes the need to pre-specify a sparsity grid and can improve model fit when the true proportion of influential features is unknown; Bayes Cπ is implemented in BGLR and would fit naturally within the same workflow
  • Model comparisons across the predictome were made descriptively by examining cross-validated accuracy values across many model-trait combinations, without a formal procedure for comparing predictors
    Could also: Diebold-Mariano tests or bootstrap confidence intervals on pairwise accuracy differences could also accompany the descriptive comparison — With many model-trait combinations evaluated simultaneously, interval estimation or formal paired tests on accuracy differences would quantify uncertainty around whether observed differences are consistent across resampling folds, complementing the descriptive pattern
Software: R/BGLR · QIIME2 2021.8 · R/dada2 · MEGA 11 · PLINK 1.9 · R stats/hclust

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
10
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37127575

Paper: Calle-García et al. (2023) "On the holobiont 'predictome' of immunocompetence in pigs." Genet Sel Evol 55:28. DOI 10.1186/s12711-023-00803-4. PMCID PMC10150480.

What the paper does (pipeline-derived results)

400 Duroc pigs; six immune traits (IgM, IgG, HP, CRP, LYM_PHAGO_FITC, %γδ T cell); host 70k SNP array; 16S V3–V4 gut microbiome. Two computational stages:

  1. 16S bioinformatics QC (QIIME2 v2021.8 + DADA2). Two sequencing batches (batch 1 = NovaSeq, from Ramayo-Caldas 2020; batch 2 = new MiSeq), denoised separately, merged. Reported counts: batch1 = 2971 ASV, batch2 = 2566 ASV, merged = 2945 ASV (53 genera), avg 136,616 reads/sample (merged), raw ASV = 57,195. Pipeline fully specified in Additional file 1 (QIIME2 .sh).
  2. Genomic/holobiont prediction (BGLR R package: RKHS, Bayes C, ensemble; 133 models/trait; 3-partition CV). Outputs = predictive accuracies (Figs 1–3), heritability h² + microbiability b² (Fig 4), per-ASV b² contributions (Fig 5).

In scope (attempted) vs out of scope (not attempted) — and why

Result Pipeline Public inputs available? Decision
Batch-2 (MiSeq) ASV count = 2566 QIIME2 2021.8 / DADA2 (Add. file 1) YES — 400 MiSeq runs in PRJNA608629, params fully specified ATTEMPTED
Batch-1 ASV (2971), merged (2945), reads/sample, raw ASV same needs run→pig mapping + samples2keep.txt (NOT provided); batch1 = 400 of 1539 NovaSeq runs, unidentifiable not attempted (20%)
Genotype QC: 68,516→41,131 SNPs Plink v1.9 NO — 70k SNP-array genotypes NOT deposited (SRA holds only 16S amplicon) not attempted (data)
Predictive accuracies (Figs 1–3) BGLR NO — phenotypes (6 traits) NOT deposited; no analysis scripts (only generic BGLR pkg) not attempted (data+code)
Heritability/microbiability (Fig 4), per-ASV b² (Fig 5) BGLR NO — same not attempted (data+code)

Key feasibility findings (verified)

  • Code link github.com/gdlc/BGLR-R is the generic BGLR R package (a third-party tool), not the authors' analysis scripts. Per brief P16 a third-party tool on the paper's data is valid — but BGLR needs the processed inputs (X, B, y), which are not public.
  • Data accession PRJNA608629 contains ONLY 16S amplicon reads (1939 runs: 1539 NovaSeq + 400 MiSeq; verified via ENA portal API). No SNP genotypes, no phenotypes. The "Availability of data" statement claims "raw sequencing data of host genotype" but no genotype/array data is present in the BioProject → the prediction results (the paper's headline) are not reproducible from public artifacts (genotype + phenotype data effectively data_restricted/on-request; analysis code no_code).
  • The one cleanly reproducible target is the batch-2 ASV count (2566): 400 MiSeq runs uniquely identifiable (sample_title RESEQ_*_fecal_bacteria), DADA2 params fully specified (trim-left-f 17, trim-left-r 21, trunc-len 250/250), version pinned (QIIME2 2021.8). This is the 80/20 result. Expected within DADA2 version/stochastic tolerance, not bit-exact.

Outcome class

Partial reproduction: one well-specified microbiome-QC data point reproduced on public data with the documented third-party pipeline; the prediction results (majority of the paper) are non-reproducible because the required phenotype and genotype data are not publicly deposited and no analysis code is shipped.

Figures / tables: Table
asv_batch2
Reported
2566 ASV (batch-2 MiSeq, after DADA2 denoising)
Reproduced
11867 ASV (400/400 samples, 16,782,904 reads) via QIIME2 2021.8 dada2 denoise-paired with the EXACT Additional-file-1 params (trim-left-f 17, trim-left-r 21, trunc-len-f/r 250) on the 400 public MiSeq runs of PRJNA608629
did not match
asv_batch1_merged_reads
Reported
2971 (batch1) / 2945 (merged) / 136616 reads-per-sample
Reproduced
not-attempted (needs run->pig map + samples2keep.txt, not shipped, to isolate 400 of 1539 NovaSeq runs and apply the batch-1 sample filter)
partial
snp_qc_41131
Reported
41,131 of 68,516 SNPs retained after Plink QC
Reproduced
not-reproducible (host SNP-array genotypes not deposited; PRJNA608629 holds only 16S amplicon)
partial
prediction_h2_b2
Reported
predictive accuracies + heritability/microbiability + per-ASV b2 (Figs 1-5)
Reproduced
not-reproducible (phenotypes+genotypes not deposited; no analysis code shipped, code link is the generic BGLR package)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 40/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

This is a data-availability + no-code reproduction shortfall, not a discrepancy or fabrication case. The deposited accession PRJNA608629 contains only 16S amplicon reads, while the host SNP-array genotypes and immune phenotypes that the headline genome+microbiome prediction results (Figs 1–5, SNP QC 41,131/68,516) depend on were never deposited — contradicting the 'Availability of data' statement — and the linked code is the generic BGLR package rather than the authors' scripts. The one fully-specified, public-data target (batch-2 MiSeq ASV count vs the reported 2566) was launched on «our HPC» but had not produced a number at finalization, so no value could actually be compared. The blocker sits on the authors'/data side (q4 red), leaving the central claim untestable (q7/q8 yellow) with no fabrication signal on the QC numbers.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

410.7 k
tokens (I/O) · 31.6 M incl. cache
233 min
runtime · 51.19 CPU-h
125 GB
peak RAM
3 (2 failed)
HPC jobs
hummel
machine