Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Mining the equine gut metagenome: poorly-characterized taxa associated with cardiovascular fitness in endurance athletes.

Commun Biol · 2022
L1 68/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
68/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 32% of all assessed papers rank 765 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (partial, honest 1:1 within DB-version tolerance). Reran the named third-party tool Kaiju v1.8.0 (greedy, -e 3) on the deposited equine gut gene catalog (doi:10.15454/NGBSPC, 25,250,066 genes) vs the nr_euk index 2021-02-24 (closest prebuilt to the paper's 2020-05-25; exact snapshot no longer hosted) on a 192-core «our HPC» node («job», 14.6 min, 127 GiB RAM). 61.8% of genes classified (15.6M). CORE taxonomic composition reproduces well: kingdom split R1 Bacteria 95.91/Euk 2.67/Arch 1.19% vs paper 95.58/2.55/1.20 (within-tol); dominant bacterial phyla R4 Firmicutes 49.2% & Bacteroidetes 22.7% vs 47.1/21.8 (within-tol); catalog scale R5a/b EXACT (25,250,066 genes, 618 bp). Genera counts R2/R3a reproduce in the right direction but +7-11% (newer nr_euk has more reference genera); raw distinct phyla R3b 222 vs paper's curated 95 (counting-convention difference, MISMATCH). Independent cross-check against the authors' OWN deposited Kaiju Krona report (Root magnitude=25,250,066 exact; genera within ~3%) corroborates derivability -> NO fabrication signal. ENA confirms 11 WGS runs, 1.363B raw read pairs vs reported ~1.41B (within-tol). IMPORTANT scoping correction: the brief's named accession GSE163767 is a HOST EXPRESSION MICROARRAY (40 samples), NOT a Kaiju input. NOT attempted (out of scope): host array DE, metabolomics, MAG assembly/CheckM, PerMANOVA fitness stats, 16S DADA2 ASVs.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 68
    assessed: 2026-06-22 ⛓ fe7eba2d99b0
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The gut microbiome contributes to endurance exercise performance, and its taxonomic/functional composition and gene content are associated with cardiovascular fitness in elite endurance horses used as a model for exercise responsiveness.

Core claims
  • Built an integrated horse gut microbiome gene catalog (~25 million unique genes) and 372 metagenome-assembled genomes (MAGs) spanning 4179 genera and 95 phyla resource
  • Gut microbiomes enriched in Lachnospiraceae taxa are negatively associated with cardiovascular capacity finding
  • More complex and functionally diverse microbiomes are associated with higher plasma glucose and reduced long-chain acylcarnitines and non-esterified fatty acids, suggesting increased mitochondrial ß-oxidation capacity finding
  • More fit athletes show upregulation of mitochondrial-related genes involved in energy metabolism, biogenesis, and Ca2+ cytosolic transport finding
  • Dominant gut phylotypes segregate horses into two distinct taxonomic and functional microbial community clusters that differ in cardiovascular fitness finding
  • More than 90% of identified genera have not been previously described in horses finding
  • Only 22.5% of genes overlapped with a previously published horse gut microbiome gene catalog, indicating vast under-sampled functional potential finding
  • 57 clusters of antimicrobial resistance genes were identified across major antibiotic classes in the gene catalog resource
Experimental setups
Assay System Perturbation Readout Platform
shotgun metagenomic sequencing fecal samples from 11 elite endurance horses none non-redundant gene catalog, gene counts, assembly size Illumina
taxonomic annotation of metagenomic reads/genes horse gut metagenome none taxonomic classification (phylum to genus level) Kaiju against NCBI nr_euk database
functional gene annotation (KEGG orthologs and CAZymes) horse gut metagenome none KO counts, CAZyme family counts EggNOG database; CAZy database
antimicrobial resistance gene screening horse gut metagenome none AMR gene cluster counts by antibiotic class ResFinder database
metagenome-assembled genome (MAG) binning and taxonomic classification horse gut metagenome none number, completeness, contamination, and taxonomy of MAGs GTDB-Tk
16S rRNA gene sequencing fecal samples from same horse cohort none genus-level taxonomic composition for cross-validation
community ordination and diversity analysis (NMDS, Bray-Curtis, PerMANOVA, Shannon/inverse Simpson) dominant gut phylotypes from horse metagenomes none clustering of samples, alpha-diversity indices
plasma metabolite profiling (implied metabolomics) plasma from endurance horses none glucose, long-chain acylcarnitines, non-esterified fatty acid concentrations
Key results
  • Non-redundant gene catalog comprised 25,250,066 genes with average length 618 bp from total assembly of 21.68 Gb
  • Core genes present in all 11 individuals made up <7.2% of the gene pool but represented 29.7-38.4% of overall microbiome abundance <7.2% (n=1,805,693)
  • 61% of gene sequences were taxonomically annotated (39% unclassified), revealing 95 phyla, 1110 families, 4179 genera 61%
  • Bacteria dominated the assemblage by abundance and diversity, followed by eukaryotes and archaea 95.58% bacteria, 2.55% eukaryotes, 1.20% archaea
  • Firmicutes and Bacteroidetes were the most abundant bacterial phyla Firmicutes 47.1%, Bacteroidetes 21.8%
  • Only 22.5% of genes overlapped with a previously published horse gut gene catalog 22.5% (n=922,362)
  • Two distinct microbial community clusters were identified by ordination of dominant phylotypes, differing in diversity and taxonomic composition PerMANOVA p=0.008, R2=0.3716
  • Cluster 1 (poorly-characterized, diverse taxa) showed significantly higher alpha-diversity than Cluster 2 (Lachnospiraceae-dominated) p=0.0134 (Shannon and inverse Simpson)
Key statistics
  • count 25,250,066 genes (size of non-redundant gut gene catalog)
  • count 372 MAGs (121 near-complete) (metagenome-assembled genomes recovered)
  • other average length 618 bp (gene catalog gene length)
  • fold_change core genes <7.2% of gene pool, 29.7-38.4% of abundance (conserved core gene set across individuals)
  • other 22.5% overlap (n=922,362) (overlap with prior published horse gut gene catalog)
  • pvalue PerMANOVA p=0.008, R2=0.3716 (clustering of dominant phylotype communities)
  • pvalue p=0.0134 (two-sided Wilcoxon rank-sum) (Shannon and inverse Simpson diversity difference between clusters)
  • count 95 phyla, 1110 families, 4179 genera (taxonomic diversity annotated via Kaiju)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study is an observational, cross-sectional metagenomic/holo-omics analysis of 11 endurance horses. Dominant gut phylotype composition was ordinated with NMDS (Bray-Curtis distances) and horses were split into two clusters (n=3 and n=8); cluster differences in diversity indices and a composite cardiovascular fitness score were tested with two-sided Wilcoxon rank-sum tests, cluster separation was tested with pairwise PerMANOVA, differential taxon abundance between clusters was assessed with DESeq2, and covariates explaining ordination variance were tested with the envfit() function, reporting r2 effect sizes.

Replicationbiological Sample sizeCohort of n=11 endurance horses, subsequently split into cluster 1 (n=3) and cluster 2 (n=8) based on microbiome ordination; no explicit power calculation or a priori sample-size justification is described GroupsTwo microbiome-composition-defined clusters of horses (n=3 vs n=8) Pairingunpaired Randomization/blindingnot stated DispersionIQR Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionyes
Statistical tests used
Test Applied to n Assumptions
two-sided Wilcoxon rank-sum test (adjusted p-values) Shannon diversity index and inverse Simpson index compared between cluster 1 and cluster 2 (Fig. 3c, d) n=11 horses (cluster 1 n=3, cluster 2 n=8) not stated
two-sided Wilcoxon rank-sum test (adjusted p-values) composite cardiovascular fitness score compared between cluster 1 and cluster 2 (Fig. 3i) n=11 horses (cluster 1 n=3, cluster 2 n=8) not stated
pairwise PerMANOVA (permutational multivariate analysis of variance) on NMDS/Bray-Curtis ordination testing separation of the two dominant-phylotype clusters (Fig. 3b); reported p=0.008, R2=0.3716 n=11 horses not stated
DESeq2 differential abundance testing (adjusted p<0.05) taxa/genera (e.g., methanogens, Lachnospiraceae members) enriched in cluster 1 vs cluster 2 (Supplementary Data 12) n=11 horses (cluster 1 n=3, cluster 2 n=8) not stated
envfit() covariate-fitting test on NMDS ordination identifying athletic-performance and mitochondrial-gene covariates significantly associated with ordination axes, with r2 effect sizes (Fig. 3g, h) n=11 samples not stated
Approaches that could also have been used
  • Cluster differences in diversity indices and fitness score (n=3 vs n=8) were assessed with the Wilcoxon rank-sum test.
    Could also: An exact permutation test or bootstrap resampling approach — With very small group sizes such as n=3, exact or permutation-based approaches can give p-values that do not rely on large-sample approximations, which is often preferred when one group is very small.
  • Cluster structure in gut microbiome composition was defined using NMDS ordination on Bray-Curtis distances, tested with PerMANOVA.
    Could also: Ordination/testing using a phylogenetically-informed distance metric (e.g., weighted or unweighted UniFrac) or alternative multivariate methods such as PCoA with ANOSIM — Phylogenetic distance metrics incorporate evolutionary relatedness among taxa, which can complement compositional (Bray-Curtis) distances and is a widely used alternative in microbiome community analysis.
  • Differential taxon abundance between clusters was tested with DESeq2, a tool originally developed for RNA-seq count data.
    Could also: Microbiome-specific compositional methods such as ANCOM-BC, ALDEx2, or MaAsLin2 — These methods are designed specifically to address compositionality and sparsity in microbiome count data, offering an alternative framework for differential abundance testing.
  • Adjusted p-values are reported for multiple comparisons (Wilcoxon tests, DESeq2) without stating the specific correction method used.
    Could also: Explicitly reporting the correction method (e.g., Benjamini-Hochberg FDR or Bonferroni) alongside the adjusted values — Naming the specific multiplicity-control method allows readers to judge how conservatively the family-wise error rate or false discovery rate was controlled across the taxa or comparisons tested.
  • Cluster comparisons rely on a small and unbalanced number of biological replicates (n=3 vs n=8).
    Could also: Reporting effect sizes with confidence intervals, or performing a sensitivity/power analysis — For small, unbalanced groups, confidence intervals around effect estimates can convey the precision of the comparison in a way that complements p-values alone.
  • Spread of the data is summarized via boxplots/violin plots showing median, IQR, and min-max range.
    Could also: Reporting means with SD, SEM, or 95% confidence intervals alongside the median/IQR summaries — Parametric summary statistics such as mean ± SD or a 95% CI can provide an additional, complementary view of central tendency and variability, particularly useful when comparing with studies that report data this way.
Software: DESeq2 · R (envfit() function, package not explicitly named in text) · Kaiju (taxonomic annotation) · GTDB-Tk (MAG taxonomic assignment)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36192523

Paper: Mach N, Midoux C, Leclercq S, et al. "Mining the equine gut metagenome: poorly-characterized taxa associated with cardiovascular fitness in endurance athletes." Communications Biology 5:1015 (2022). DOI 10.1038/s42003-022-03977-7. PMCID PMC9529974.

Named code (brief): https://github.com/bioinformatics-centre/kaiju (Kaiju — a third-party protein-level taxonomic classifier). Per BRIEF rule P16, running an existing third-party tool on the paper's own data is a valid reproduction.

Named data (brief): geo:GSE163767.


⚠️ Data-pairing caveat (important)

The brief pairs the named tool Kaiju with accession GSE163767, but these do not match: GSE163767 is a host gene-expression microarray series (40 samples = 20 horses × pre/post race, platform GPL20908, Equus caballus; title "Crosstalk between gut microbiota and mitochondria during endurance…"). Microarray expression is not an input to Kaiju and is out of scope for a metagenomic taxonomic reproduction. It is profiled (as required) but not used for the Kaiju run.

The data Kaiju actually consumes in this paper are the shotgun metagenomic reads and the deposited gene catalog, found under BioProject PRJNA438436 and the Migale/INRAE dataverse doi:10.15454/NGBSPC — see datasets below.


Datasets the paper relies on

# Repo / accession Data type Role In scope for Kaiju?
D1 GEO GSE163767 host microarray expression, 40 samples host transcriptomics No (wrong modality)
D2 SRA/ENA PRJNA438436, WGS runs SRR17543904–17543914 (11) shotgun metagenomics, 1.41 B raw reads Kaiju / assembly input Yes (raw)
D3 INRAE/Migale doi:10.15454/NGBSPCGenecatalog_with-note.fna.gz (≈25.25 M genes) + Genecatalog_KO.tab + Genecatalog_CAZy.tab non-redundant gene catalog (assembly output) direct Kaiju input for the reported taxonomic composition Yes (primary)
D4 SRA PRJNA438436, 16S AMPLICON runs (SRR13664916–931 = 11 experimental; SRR20568504–555 = validation/extra) 16S rRNA amplicon QIIME2/DADA2 ASVs secondary (different tool)
D5 MAG assemblies PRJNA438436 (JAKSHS…) metagenome-assembled genomes ATLAS/binning output not Kaiju
D6 Metabolomics Workbench UrqK1489 NMR/MS plasma metabolites wet-lab out of scope

In-scope (pipeline-derived) results — reproduction targets

The paper states the taxonomic composition of the gene catalog was obtained with Kaiju v1.8.0, greedy mode, -a greedy -e 3 (≤3 mismatches), against the NCBI nr_euk reference database released 2020-05-25. Reproduction = run that tool with those parameters on the deposited gene catalog (D3) and recompute:

  • R1 Kingdom split of classified genes: Bacteria 95.58 % / Eukaryota 2.55 % / Archaea 1.20 %.
  • R2 Genus counts per kingdom: Bacteria 2927, Eukaryota 1081, Archaea 170.
  • R3 Overall taxonomic diversity: 4179 genera spanning 95 phyla (biorxiv preprint states 4696 genera — discrepancy to note).
  • R4 Dominant phyla among bacteria: Firmicutes 47.1 %, Bacteroidetes 21.8 %.
  • R5 Gene-catalog scale (sanity): 25,250,066 genes, mean length 618 bp, assembly size 21.68 Gb — directly checkable from the deposited FASTA (counts, not Kaiju).

Secondary (different tool, "equally valid" data points if time allows):

  • R6 16S: 364,026 high-quality reads for 11 horses (mean 33,093, range 12,052–62,670); 5412 ASVs (QIIME2 + DADA2, Greengenes 13.8). Reproduce read counts from D4 raw FASTQ (count, tractable) and optionally DADA2 ASV count.
  • R7 Shotgun QC: 1124 M high-quality paired reads after host decontamination (raw 1.41 B from ENA → checkable QC-survival ratio).

Out of scope (not attempted)

  • Host microarray DE analysis (GSE163767) — wet-lab/array, different modality.
  • Metabolomic
Figures / tables: figure
R1a
Reported
Bacteria 95.58% of classified catalog genes (Kaiju nr_euk)
Reproduced
95.91%
within tolerance
R1b
Reported
Eukaryota 2.55%
Reproduced
2.67%
within tolerance
R1c
Reported
Archaea 1.20%
Reproduced
1.19%
within tolerance
R2a
Reported
2927 bacterial genera
Reproduced
3259
partial
R2b
Reported
1081 eukaryotic genera
Reproduced
1155
partial
R2c
Reported
170 archaeal genera
Reproduced
182
within tolerance
R3a
Reported
4179 genera total
Reproduced
4596 (Bact+Euk+Arch)
partial
R3b
Reported
95 phyla
Reproduced
222 distinct (154 bacterial)
did not match
R4a
Reported
Firmicutes 47.1% of bacteria
Reproduced
49.19%
within tolerance
R4b
Reported
Bacteroidetes 21.8% of bacteria
Reproduced
22.69%
within tolerance
R5a
Reported
25,250,066 non-redundant genes
Reproduced
25,250,066
exact
R5b
Reported
mean gene length 618 bp
Reproduced
618.0 bp
exact
R5c
Reported
assembly size 21.68 Gb
Reproduced
gene catalog sum = 15.61 Gb
did not match
R7a
Reported
shotgun raw ~1.41B (vs 1124M post-QC)
Reproduced
1,362,661,344 raw read pairs (11 WGS runs, ENA)
within tolerance
R7b
Reported
depth 93-107M paired/sample
Reproduced
110.7-145.7M reads/run (raw)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 68/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

This is a solid, honest partial reproduction of the deposited equine gut gene catalog's taxonomic composition. The input data (catalog, doi:10.15454/NGBSPC, 25,250,066 genes) is identical 1:1 — R5a/R5b reproduce exactly and core composition (R1 kingdom split, R4 dominant phyla, R7a read totals) is within-tol. Deviations sit on the input/reference side: genera counts +7–11% from nr_euk DB-version drift and phyla 222 vs 95 a counting convention, with R5c a different-metric (assembly vs catalog) artifact — none on the authors' computation. An independent cross-check against the authors' own deposited Kaiju/Krona output confirms the headline numbers are derivable from shared data (no fabrication signal); the only real limitation is that the paper's central fitness-association claim was out of scope, hence q7 yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

429.7 k
tokens (I/O) · 36.7 M incl. cache
93 min
runtime · 26.32 CPU-h
126.9 GB
peak RAM
1
HPC jobs
hummel
machine