Mining the equine gut metagenome: poorly-characterized taxa associated with cardiovascular fitness in endurance athletes.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- 🟡A deviation arose in the data or preprocessing
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (partial, honest 1:1 within DB-version tolerance). Reran the named third-party tool Kaiju v1.8.0 (greedy, -e 3) on the deposited equine gut gene catalog (doi:10.15454/NGBSPC, 25,250,066 genes) vs the nr_euk index 2021-02-24 (closest prebuilt to the paper's 2020-05-25; exact snapshot no longer hosted) on a 192-core «our HPC» node («job», 14.6 min, 127 GiB RAM). 61.8% of genes classified (15.6M). CORE taxonomic composition reproduces well: kingdom split R1 Bacteria 95.91/Euk 2.67/Arch 1.19% vs paper 95.58/2.55/1.20 (within-tol); dominant bacterial phyla R4 Firmicutes 49.2% & Bacteroidetes 22.7% vs 47.1/21.8 (within-tol); catalog scale R5a/b EXACT (25,250,066 genes, 618 bp). Genera counts R2/R3a reproduce in the right direction but +7-11% (newer nr_euk has more reference genera); raw distinct phyla R3b 222 vs paper's curated 95 (counting-convention difference, MISMATCH). Independent cross-check against the authors' OWN deposited Kaiju Krona report (Root magnitude=25,250,066 exact; genera within ~3%) corroborates derivability -> NO fabrication signal. ENA confirms 11 WGS runs, 1.363B raw read pairs vs reported ~1.41B (within-tol). IMPORTANT scoping correction: the brief's named accession GSE163767 is a HOST EXPRESSION MICROARRAY (40 samples), NOT a Kaiju input. NOT attempted (out of scope): host array DE, metabolomics, MAG assembly/CheckM, PerMANOVA fitness stats, 16S DADA2 ASVs.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 68assessed: 2026-06-22 ⛓ fe7eba2d99b0
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe gut microbiome contributes to endurance exercise performance, and its taxonomic/functional composition and gene content are associated with cardiovascular fitness in elite endurance horses used as a model for exercise responsiveness.
- ★ Built an integrated horse gut microbiome gene catalog (~25 million unique genes) and 372 metagenome-assembled genomes (MAGs) spanning 4179 genera and 95 phyla resource
- ★ Gut microbiomes enriched in Lachnospiraceae taxa are negatively associated with cardiovascular capacity finding
- ★ More complex and functionally diverse microbiomes are associated with higher plasma glucose and reduced long-chain acylcarnitines and non-esterified fatty acids, suggesting increased mitochondrial ß-oxidation capacity finding
- ★ More fit athletes show upregulation of mitochondrial-related genes involved in energy metabolism, biogenesis, and Ca2+ cytosolic transport finding
- ★ Dominant gut phylotypes segregate horses into two distinct taxonomic and functional microbial community clusters that differ in cardiovascular fitness finding
- More than 90% of identified genera have not been previously described in horses finding
- Only 22.5% of genes overlapped with a previously published horse gut microbiome gene catalog, indicating vast under-sampled functional potential finding
- 57 clusters of antimicrobial resistance genes were identified across major antibiotic classes in the gene catalog resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| shotgun metagenomic sequencing | fecal samples from 11 elite endurance horses | none | non-redundant gene catalog, gene counts, assembly size | Illumina |
| taxonomic annotation of metagenomic reads/genes | horse gut metagenome | none | taxonomic classification (phylum to genus level) | Kaiju against NCBI nr_euk database |
| functional gene annotation (KEGG orthologs and CAZymes) | horse gut metagenome | none | KO counts, CAZyme family counts | EggNOG database; CAZy database |
| antimicrobial resistance gene screening | horse gut metagenome | none | AMR gene cluster counts by antibiotic class | ResFinder database |
| metagenome-assembled genome (MAG) binning and taxonomic classification | horse gut metagenome | none | number, completeness, contamination, and taxonomy of MAGs | GTDB-Tk |
| 16S rRNA gene sequencing | fecal samples from same horse cohort | none | genus-level taxonomic composition for cross-validation | — |
| community ordination and diversity analysis (NMDS, Bray-Curtis, PerMANOVA, Shannon/inverse Simpson) | dominant gut phylotypes from horse metagenomes | none | clustering of samples, alpha-diversity indices | — |
| plasma metabolite profiling (implied metabolomics) | plasma from endurance horses | none | glucose, long-chain acylcarnitines, non-esterified fatty acid concentrations | — |
- – Non-redundant gene catalog comprised 25,250,066 genes with average length 618 bp from total assembly of 21.68 Gb
- – Core genes present in all 11 individuals made up <7.2% of the gene pool but represented 29.7-38.4% of overall microbiome abundance <7.2% (n=1,805,693)
- – 61% of gene sequences were taxonomically annotated (39% unclassified), revealing 95 phyla, 1110 families, 4179 genera 61%
- – Bacteria dominated the assemblage by abundance and diversity, followed by eukaryotes and archaea 95.58% bacteria, 2.55% eukaryotes, 1.20% archaea
- – Firmicutes and Bacteroidetes were the most abundant bacterial phyla Firmicutes 47.1%, Bacteroidetes 21.8%
- – Only 22.5% of genes overlapped with a previously published horse gut gene catalog 22.5% (n=922,362)
- – Two distinct microbial community clusters were identified by ordination of dominant phylotypes, differing in diversity and taxonomic composition PerMANOVA p=0.008, R2=0.3716
- ▲ Cluster 1 (poorly-characterized, diverse taxa) showed significantly higher alpha-diversity than Cluster 2 (Lachnospiraceae-dominated) p=0.0134 (Shannon and inverse Simpson)
- count 25,250,066 genes (size of non-redundant gut gene catalog)
- count 372 MAGs (121 near-complete) (metagenome-assembled genomes recovered)
- other average length 618 bp (gene catalog gene length)
- fold_change core genes <7.2% of gene pool, 29.7-38.4% of abundance (conserved core gene set across individuals)
- other 22.5% overlap (n=922,362) (overlap with prior published horse gut gene catalog)
- pvalue PerMANOVA p=0.008, R2=0.3716 (clustering of dominant phylotype communities)
- pvalue p=0.0134 (two-sided Wilcoxon rank-sum) (Shannon and inverse Simpson diversity difference between clusters)
- count 95 phyla, 1110 families, 4179 genera (taxonomic diversity annotated via Kaiju)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study is an observational, cross-sectional metagenomic/holo-omics analysis of 11 endurance horses. Dominant gut phylotype composition was ordinated with NMDS (Bray-Curtis distances) and horses were split into two clusters (n=3 and n=8); cluster differences in diversity indices and a composite cardiovascular fitness score were tested with two-sided Wilcoxon rank-sum tests, cluster separation was tested with pairwise PerMANOVA, differential taxon abundance between clusters was assessed with DESeq2, and covariates explaining ordination variance were tested with the envfit() function, reporting r2 effect sizes.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| two-sided Wilcoxon rank-sum test (adjusted p-values) | Shannon diversity index and inverse Simpson index compared between cluster 1 and cluster 2 (Fig. 3c, d) | n=11 horses (cluster 1 n=3, cluster 2 n=8) | not stated |
| two-sided Wilcoxon rank-sum test (adjusted p-values) | composite cardiovascular fitness score compared between cluster 1 and cluster 2 (Fig. 3i) | n=11 horses (cluster 1 n=3, cluster 2 n=8) | not stated |
| pairwise PerMANOVA (permutational multivariate analysis of variance) on NMDS/Bray-Curtis ordination | testing separation of the two dominant-phylotype clusters (Fig. 3b); reported p=0.008, R2=0.3716 | n=11 horses | not stated |
| DESeq2 differential abundance testing (adjusted p<0.05) | taxa/genera (e.g., methanogens, Lachnospiraceae members) enriched in cluster 1 vs cluster 2 (Supplementary Data 12) | n=11 horses (cluster 1 n=3, cluster 2 n=8) | not stated |
| envfit() covariate-fitting test on NMDS ordination | identifying athletic-performance and mitochondrial-gene covariates significantly associated with ordination axes, with r2 effect sizes (Fig. 3g, h) | n=11 samples | not stated |
-
Cluster differences in diversity indices and fitness score (n=3 vs n=8) were assessed with the Wilcoxon rank-sum test.↳ Could also: An exact permutation test or bootstrap resampling approach — With very small group sizes such as n=3, exact or permutation-based approaches can give p-values that do not rely on large-sample approximations, which is often preferred when one group is very small.
-
Cluster structure in gut microbiome composition was defined using NMDS ordination on Bray-Curtis distances, tested with PerMANOVA.↳ Could also: Ordination/testing using a phylogenetically-informed distance metric (e.g., weighted or unweighted UniFrac) or alternative multivariate methods such as PCoA with ANOSIM — Phylogenetic distance metrics incorporate evolutionary relatedness among taxa, which can complement compositional (Bray-Curtis) distances and is a widely used alternative in microbiome community analysis.
-
Differential taxon abundance between clusters was tested with DESeq2, a tool originally developed for RNA-seq count data.↳ Could also: Microbiome-specific compositional methods such as ANCOM-BC, ALDEx2, or MaAsLin2 — These methods are designed specifically to address compositionality and sparsity in microbiome count data, offering an alternative framework for differential abundance testing.
-
Adjusted p-values are reported for multiple comparisons (Wilcoxon tests, DESeq2) without stating the specific correction method used.↳ Could also: Explicitly reporting the correction method (e.g., Benjamini-Hochberg FDR or Bonferroni) alongside the adjusted values — Naming the specific multiplicity-control method allows readers to judge how conservatively the family-wise error rate or false discovery rate was controlled across the taxa or comparisons tested.
-
Cluster comparisons rely on a small and unbalanced number of biological replicates (n=3 vs n=8).↳ Could also: Reporting effect sizes with confidence intervals, or performing a sensitivity/power analysis — For small, unbalanced groups, confidence intervals around effect estimates can convey the precision of the comparison in a way that complements p-values alone.
-
Spread of the data is summarized via boxplots/violin plots showing median, IQR, and min-max range.↳ Could also: Reporting means with SD, SEM, or 95% confidence intervals alongside the median/IQR summaries — Parametric summary statistics such as mean ± SD or a 95% CI can provide an additional, complementary view of central tendency and variability, particularly useful when comparing with studies that report data this way.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36192523
Paper: Mach N, Midoux C, Leclercq S, et al. "Mining the equine gut metagenome: poorly-characterized taxa associated with cardiovascular fitness in endurance athletes." Communications Biology 5:1015 (2022). DOI 10.1038/s42003-022-03977-7. PMCID PMC9529974.
Named code (brief): https://github.com/bioinformatics-centre/kaiju (Kaiju — a third-party protein-level taxonomic classifier). Per BRIEF rule P16, running an existing third-party tool on the paper's own data is a valid reproduction.
Named data (brief): geo:GSE163767.
⚠️ Data-pairing caveat (important)
The brief pairs the named tool Kaiju with accession GSE163767, but these do not match: GSE163767 is a host gene-expression microarray series (40 samples = 20 horses × pre/post race, platform GPL20908, Equus caballus; title "Crosstalk between gut microbiota and mitochondria during endurance…"). Microarray expression is not an input to Kaiju and is out of scope for a metagenomic taxonomic reproduction. It is profiled (as required) but not used for the Kaiju run.
The data Kaiju actually consumes in this paper are the shotgun metagenomic reads and the deposited gene catalog, found under BioProject PRJNA438436 and the Migale/INRAE dataverse doi:10.15454/NGBSPC — see datasets below.
Datasets the paper relies on
| # | Repo / accession | Data type | Role | In scope for Kaiju? |
|---|---|---|---|---|
| D1 | GEO GSE163767 | host microarray expression, 40 samples | host transcriptomics | No (wrong modality) |
| D2 | SRA/ENA PRJNA438436, WGS runs SRR17543904–17543914 (11) | shotgun metagenomics, 1.41 B raw reads | Kaiju / assembly input | Yes (raw) |
| D3 | INRAE/Migale doi:10.15454/NGBSPC — Genecatalog_with-note.fna.gz (≈25.25 M genes) + Genecatalog_KO.tab + Genecatalog_CAZy.tab |
non-redundant gene catalog (assembly output) | direct Kaiju input for the reported taxonomic composition | Yes (primary) |
| D4 | SRA PRJNA438436, 16S AMPLICON runs (SRR13664916–931 = 11 experimental; SRR20568504–555 = validation/extra) | 16S rRNA amplicon | QIIME2/DADA2 ASVs | secondary (different tool) |
| D5 | MAG assemblies PRJNA438436 (JAKSHS…) | metagenome-assembled genomes | ATLAS/binning output | not Kaiju |
| D6 | Metabolomics Workbench UrqK1489 | NMR/MS plasma metabolites | wet-lab | out of scope |
In-scope (pipeline-derived) results — reproduction targets
The paper states the taxonomic composition of the gene catalog was obtained with
Kaiju v1.8.0, greedy mode, -a greedy -e 3 (≤3 mismatches), against the NCBI
nr_euk reference database released 2020-05-25. Reproduction = run that tool
with those parameters on the deposited gene catalog (D3) and recompute:
- R1 Kingdom split of classified genes: Bacteria 95.58 % / Eukaryota 2.55 % / Archaea 1.20 %.
- R2 Genus counts per kingdom: Bacteria 2927, Eukaryota 1081, Archaea 170.
- R3 Overall taxonomic diversity: 4179 genera spanning 95 phyla (biorxiv preprint states 4696 genera — discrepancy to note).
- R4 Dominant phyla among bacteria: Firmicutes 47.1 %, Bacteroidetes 21.8 %.
- R5 Gene-catalog scale (sanity): 25,250,066 genes, mean length 618 bp, assembly size 21.68 Gb — directly checkable from the deposited FASTA (counts, not Kaiju).
Secondary (different tool, "equally valid" data points if time allows):
- R6 16S: 364,026 high-quality reads for 11 horses (mean 33,093, range 12,052–62,670); 5412 ASVs (QIIME2 + DADA2, Greengenes 13.8). Reproduce read counts from D4 raw FASTQ (count, tractable) and optionally DADA2 ASV count.
- R7 Shotgun QC: 1124 M high-quality paired reads after host decontamination (raw 1.41 B from ENA → checkable QC-survival ratio).
Out of scope (not attempted)
- Host microarray DE analysis (GSE163767) — wet-lab/array, different modality.
- Metabolomic
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a solid, honest partial reproduction of the deposited equine gut gene catalog's taxonomic composition. The input data (catalog, doi:10.15454/NGBSPC, 25,250,066 genes) is identical 1:1 — R5a/R5b reproduce exactly and core composition (R1 kingdom split, R4 dominant phyla, R7a read totals) is within-tol. Deviations sit on the input/reference side: genera counts +7–11% from nr_euk DB-version drift and phyla 222 vs 95 a counting convention, with R5c a different-metric (assembly vs catalog) artifact — none on the authors' computation. An independent cross-check against the authors' own deposited Kaiju/Krona output confirms the headline numbers are derivable from shared data (no fabrication signal); the only real limitation is that the paper's central fitness-association claim was out of scope, hence q7 yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.