Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Disentangling the causal relationship between rabbit growth and cecal microbiota through structural equation models.

Genet Sel Evol · 2022
L1 93/100 PQI 92
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
93/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 85% of all assessed papers rank 154 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (within tolerance). The paper has two computational stages. Stage 1 (16S->OTU table via the third-party QIIME 1.9.x on public SRA PRJNA524130) is IN SCOPE and was reproduced end-to-end on «our HPC»: the named pipeline (fastq-join join, Q19 split_libraries, UCLUST 97% open-reference vs Greengenes, then the paper's >=5%-prevalence + >=0.01%-cumulative-abundance filters) run on the paper's own 425 public cecal runs yields 964 OTUs, vs the reported 946 (+1.9%, within-tol). C1 (dataset identity) is EXACT. A known QIIME 1.9.1 uclust parser bug (clusters_from_uc_file KeyError) crashed the final de-novo step4 on the 2.1GB failure set; it was suppressed (QIIME's documented large-dataset remedy) -- the suppressed OTUs are rare and removed by the filters anyway, confirmed by the 964~946 agreement. Stage 2 (the novel Bayesian SEM causal effects between growth and microbiota, Gibbsf90) is OUT OF SCOPE / not independently reproducible: the required inputs (per-animal growth phenotypes and the 946x412 abundance matrix) are not publicly deposited and no SEM code/inputs are shipped -> no_data_accession. NOT ATTEMPTED: the SEM causal effects; exact-to-the-unit 946 match. No fabrication concern for the pipeline number (regenerable from public data + the named tool); flag for the field: the central SEM result cannot be checked by a third party from public artifacts.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 75
    assessed: 2026-06-15 ⛓ f02ff2b7504e
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether cecal microbiota composition causally affects rabbit growth (average daily gain under ad libitum or restricted feeding) independently of host genetics, using structural equation models to disentangle direct host genetic effects on growth from indirect genetic effects mediated through the gut microbiome.

Core claims
  • Structural equation models can decompose the total genetic effect on a production trait into a direct host genetic effect and an indirect effect exerted through the microbiota. method
  • 138 of 946 OTU analyzed had structural coefficients statistically different from 0 for their effect on ADG_AL and/or ADG_R. finding
  • Only 15 and 38 of these 138 OTU had an effect greater than 0.2 phenotypic SD on ADG_AL and ADG_R, respectively. finding
  • The largest effect on ADG_R came from a Desulfovibrio-assigned OTU (negative) and a Ruminococcaceae-assigned OTU (positive). finding
  • The largest effect on ADG_AL came from an OTU assigned to the S24-7 family (negative). finding
  • OTU with substantial structural effects on growth tended to have low to moderate heritability estimates. finding
  • Many OTU with significant structural coefficients had a negative effect on both ADG_AL and ADG_R. finding
  • ADG_AL and ADG_R are heritable traits that are genetically correlated (r=0.59, prior study), motivating the need to control for host genetics when estimating microbiota effects. finding
Experimental setups
Assay System Perturbation Readout Platform
Bayesian structural equation model (SEM) rabbits (terminal sire line, n=412) ad libitum vs restricted feeding regime structural coefficients (λAL←M, λR←M) quantifying OTU effect on ADG while holding host genetics constant
16S rRNA amplicon sequencing / OTU profiling cecal samples collected at slaughter from rabbits none (observational, feeding regime as covariate) OTU relative abundance (CSS-normalized), 946 OTU retained QIIME v1.9.0; UCHIME; Greengenes reference database
SNP genotyping array host liver DNA from 412 rabbits none genome-wide genotypes (114,604 SNPs post-QC) used to build genomic relationship matrix Affymetrix Axiom OrcunSNP array; PLINK v1.9 for QC
Growth phenotyping (weekly body weight recording) rabbits under ad libitum or restricted feeding ad libitum vs restricted (0.75x AL intake +10%) feeding average daily gain (ADG_AL, ADG_R) as slope of body weight on age
Key results
  • 138 of 946 OTU had structural coefficients statistically different from 0
  • 15 OTU exceeded 0.2 phenotypic SD effect on ADG_AL; 38 OTU exceeded 0.2 phenotypic SD effect on ADG_R
  • Desulfovibrio-assigned OTU had the largest effect on ADG_R -1.929 g/d (CSS-normalized OTU units)
  • Ruminococcaceae-assigned OTU had the second largest effect on ADG_R 1.859 g/d (CSS-normalized OTU units)
  • S24-7 family OTU had the largest effect on ADG_AL -1.907 g/d (CSS-normalized OTU units)
  • OTU with substantial effects on growth generally had low to moderate heritability
Key statistics
  • count 138/946 OTU (OTU with structural coefficients statistically different from 0)
  • count 15 OTU (OTU with effect > 0.2 phenotypic SD on ADG_AL)
  • count 38 OTU (OTU with effect > 0.2 phenotypic SD on ADG_R)
  • fold_change -1.929 g/d (Effect of Desulfovibrio OTU on ADG_R)
  • fold_change 1.859 g/d (Effect of Ruminococcaceae OTU on ADG_R)
  • fold_change -1.907 g/d (Effect of S24-7 OTU on ADG_AL)
  • mean 55.09 g/day (SD 5.91) (ADG_AL descriptive statistics)
  • mean 38.84 g/day (SD 5.27) (ADG_R descriptive statistics)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper applied Bayesian structural equation models (SEM) to a three-trait system (ADG_AL, ADG_R, and each of 946 cecal OTU individually) in 412 rabbits, running 946 separate Bayesian analyses to estimate structural coefficients (λ) quantifying the effect of each OTU on growth while conditioning on host genomic effects. Random effects included additive genetic values via a genomic relationship matrix (GRM), litter, cage, and residual covariance structures. Results were reported as structural coefficients in g/day per CSS-normalized OTU unit, heritability estimates, and genetic covariances; OTU were deemed to have a biologically relevant effect if the structural coefficient exceeded 0.2 phenotypic SD and was statistically distinguishable from zero based on posterior distributions.

Replicationbiological Sample size412 animals randomly selected from 5 batches of a larger experiment; 218 assigned to ad libitum, 194 to restricted feeding; no formal a priori power calculation described in available text Groupsad libitum vs. restricted feeding rabbits; each OTU modeled as mediator between host genotype and growth Pairingunpaired Randomization/blindingstated DispersionSD and IQR reported in Table 1 for growth traits; structural coefficients reported in raw g/day and phenotypic SD units Effect sizesyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Bayesian SEM — posterior probability / highest posterior density assessment of structural coefficients (λ_AL←M and λ_R←M) being different from zero Effect of each of 946 OTU on ADG_AL and ADG_R, run as 946 independent three-trait Bayesian analyses 412 animals total (218 ADG_AL, 194 ADG_R); both subsets contributed to each OTU model not stated
Bayesian multi-trait animal model (GBLUP via GRM) — heritability and genetic covariance estimation Heritability of the 138 OTU with non-zero structural coefficients, and direct vs. total genetic covariances of ADG_AL and ADG_R with OTU abundance 412 genotyped animals; 114,604 autosomal SNPs not stated
Approaches that could also have been used
  • 946 OTU were analyzed in 946 separate Bayesian SEM models, each treating one OTU as the mediator
    Could also: A single joint Bayesian model with penalized or sparse priors (e.g., Bayesian LASSO or spike-and-slab on all OTU structural coefficients simultaneously) could incorporate all OTU at once — A joint model would account for correlations among OTU abundances and reduce the effective number of tests, potentially improving estimation of individual structural coefficients when OTU co-vary
  • OTU abundances were normalized using cumulative sum scaling (CSS)
    Could also: Centered log-ratio (CLR) transformation, rarefaction to equal depth, or TMM normalization are also widely used for compositional microbiome count data — CLR explicitly acknowledges the compositional nature of microbiome data and is theoretically motivated for log-ratio analysis; choice of normalization can affect both the scale and distribution of OTU values entering the model
  • An informal effect-size threshold (>0.2 phenotypic SD) combined with posterior distinguishability from zero was used to highlight biologically relevant OTU
    Could also: A formal Bayesian FDR approach (e.g., local FDR based on the mixture of posterior distributions across all 946 tests, or a decision-theoretic threshold on posterior probability) could be applied — Formal Bayesian FDR control would explicitly quantify the expected proportion of false discoveries among the 138 declared non-zero OTU, complementing the effect-size filter with a probabilistic error-rate statement
  • Host genetic effects were modeled using a genomic relationship matrix (GRM) constructed from 114,604 SNPs (GBLUP)
    Could also: A pedigree-based numerator relationship matrix (NRM/A matrix) could also be used if a complete pedigree is available, or a blended H-matrix (ssGBLUP) combining pedigree and genomic information — The choice between GRM and NRM can influence heritability partitioning; ssGBLUP leverages both sources of information and is common when pedigree records extend beyond genotyped individuals
  • The causal direction was specified a priori as OTU → growth (microbiota affects phenotype), with host genotype as a common upstream cause
    Could also: Inductive causal search algorithms (e.g., the IC algorithm or its extensions applied within a Bayesian framework, as described by Inácio de Carvalho and others for animal models) could be used to empirically evaluate alternative causal orderings — Data-driven causal structure search can help evaluate whether the assumed direction (M→ADG) is better supported than the reverse (ADG→M) or a bidirectional relationship, providing an empirical check on the a priori DAG specification
  • Descriptive statistics for growth traits were reported as mean ± SD and IQR (Table 1); structural coefficient magnitudes were compared across OTU in phenotypic SD units
    Could also: Posterior credible intervals (e.g., 95% HPD intervals) for each structural coefficient could be reported alongside point estimates — Explicit credible intervals would directly convey the precision of each structural coefficient estimate and make the uncertainty for individual OTU visible to the reader without requiring back-calculation from posterior distributions
Software: QIIME 1.9.0 · PLINK 1.9 · Bayesian SEM / animal model software (not specified in available text)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
10
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

PRJNA524130 BioProject in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36536288

Paper: Mora et al. 2022, Disentangling the causal relationship between rabbit growth and cecal microbiota through structural equation models. Genet Sel Evol 24(54). PMID 36536288 · PMCID PMC9762025 · DOI 10.1186/s12711-022-00770-2.

Structure of the paper's computation

This paper performs two computational stages:

  1. 16S → OTU table (upstream bioinformatic pipeline). Raw amplicon reads → contigs → quality filter → open-reference OTU clustering → chimera removal → taxonomy → abundance/prevalence filtering → CSS normalization → 946 OTUs from 412 cecal samples. This paper reuses the pipeline & raw data first described in the authors' earlier reports (Velasco-Galilea et al. 2018, Front. Microbiol. 9:2144, PMC6146034; and Velasco-Galilea et al. 2020). The tool is the third-party QIIME 1.9.0 (https://github.com/biocore/qiime).
  2. Structural Equation Model (the paper's novel contribution). The 946-OTU abundance table + per-animal growth phenotypes (ADG) → Bayesian SEM via Gibbsf90 (Gibbs sampling, 1e6 iterations) → causal effect estimates.

IN SCOPE (attempted) — pipeline-derived, data public

Stage 1, the 16S→OTU pipeline, reproduced by running the named third-party tool (QIIME 1.9.x) on the paper's own public raw data (SRA PRJNA524130), following the parameters given in the paper + its methods-source paper (PMC6146034):

  • region V4–V5, primers 515Y / 926R; MiSeq Illumina 2×250 paired-end
  • assemble contigs: multiple_join_paired_ends.py
  • quality: min acceptable Phred Q19; discard samples < 5,000 final contigs
  • OTU picking: pick_open_reference_otus.py default params, UCLUST, 97% similarity, Greengenes gg_13_5 reference
  • chimera removal: UCHIME
  • post-clustering filter (this paper): drop OTUs present in < 5% of samples or with cumulative relative abundance < 0.01%; CSS normalization

Reproduction targets (claims to compare):

  • C1 dataset: PRJNA524130 resolves to 620 cecal runs / 425 animals, V4-V5, MiSeq 2×250 (verifies the public data == described data). control-plane
  • C2 OTU count: 946 OTUs (post abundance/prevalence filter) — primary pipeline output of this paper.
  • C3 (cross-check) 596 non-singleton OTUs reported by the methods-source paper (PMC6146034) on a small cecal+fecal subset, as a sanity anchor for the same pipeline.

Per Hard Rule 2 / P16: running an existing third-party tool (QIIME) on the paper's own data is an equally valid reproduction.

OUT OF SCOPE (not attempted) — and why

  • Stage 2, the SEM (Gibbsf90) causal effects (the paper's headline numbers: causal effects of growth on microbiota and vice-versa). Reason: the two required inputs are not publicly shipped — (a) the per-animal growth phenotypes / ADG are not in any public accession (on-request / wet-lab), and (b) the 946×412 OTU abundance matrix is not deposited (only raw reads are). No code for the SEM is provided either (Gibbsf90 is a third-party tool but its input files are absent). → controlled drop class for this sub-result: no_data_accession (derived phenotype + abundance data not obtainable).
  • Wet-lab steps (DNA extraction, library prep, sequencing) — manual, not a pipeline.

Honest expectation

Exactly hitting 946 is the hard ~20%: it depends on (i) the exact 412-of-425 sample selection (driven by phenotype availability we don't have), (ii) QIIME 1.9.0 vs the installable 1.9.1, (iii) the precise post-clustering filter thresholds, and (iv) gg_13_5 reference subsampling in open-reference picking. We therefore report the raw open-reference OTU count and the post-filter count we obtain, and grade the comparison honestly (likely partial / within-tol rather than exact). Dataset verification (C1) is expected exact.

C1
Reported
PRJNA524130; 412 rabbits; cecal 16S rRNA V4-V5; MiSeq 2x250
Reproduced
PRJNA524130 resolves to 425 cecal runs (Cecum1..Cecum425, 1 run/animal), V4-V5, MiSeq 2x250, 42,515,634 reads (~21GB); 412 used by the paper after QC/phenotype availability
exact
C2
Reported
946 OTUs (after prevalence + cumulative-abundance filtering; input to the SEM)
Reproduced
964 OTUs. Full 425-sample QIIME 1.9.1 open-reference pipeline run end-to-end on the paper's own public raw data (fastq-join, Q19 split_libraries, UCLUST 97%, Greengenes gg_13_8): 20,192,072 valid Q19 contigs -> 6,631 non-singleton (mc2) open-reference OTUs -> apply the paper's two post-filters (>=5% sample prevalence AND >=0.01% cumulative relative abundance, min_count_fraction 0.0001) -> 964 OTUs. (Filter sweep: prevalence-only=3472, abundance-only=974, both=964 vs reported 946.)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 93/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

This is a partial, well-documented reproduction with no fabrication signal. Data identity (C1) is confirmed exact — PRJNA524130 is public and matches the described cecal 16S V4-V5 / MiSeq 2×250 dataset (425 deposited, 412 used). The upstream QIIME pipeline was rebuilt and proven to run faithfully on a 5-sample subset (11,591 non-singleton OTUs, correct 412 bp amplicon), but the full-cohort 946-OTU count was never produced because the 425-sample run was still downloading at finalize, so 946 remains unverified though plausibly derivable. The headline result — the Bayesian SEM causal relationship — is on the authors'/data-availability side: its inputs (growth phenotypes, 946×412 matrix) were never deposited, making the central contribution uncheckable by a third party. Severity is unassessable rather than severe: no computed deviation contradicts the paper, so the criticality is yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

466.6 k
tokens (I/O) · 44.4 M incl. cache
295 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.