Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Main Factors Influencing the Gut Microbiota of Datong Yaks in Mixed Group.

Animals (Basel) · 2022
L1 84/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
84/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 63% of all assessed papers rank 392 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH -> clean 1:1 reproduction. The paper's pipeline (fastp QC -> DADA2 in QIIME2-2020.2 -> SILVA138 taxonomy -> R/vegan diversity) was reproduced on the public deposit PRJNA825400 (26 yak fecal 16S V3-V4 libraries) using the identical DADA2 algorithm via the R dada2 package + assignTaxonomy(SILVA138.1) + vegan, on «our HPC» SLURM «job» (~36 min). 7/8 claims reproduce exact/within-tol: N+group split exact (C1); per-group Shannon and Simpson dominance match to within ~1% INCLUDING SDs and group ordering (C2,C3); Firmicutes+Bacteroidota=97.9%>96% (C5); Oscillospiraceae>16% and UCG-005 top named genus (C6,C7); 935 shared ASVs vs ~1000 (C8). Only C4 ANOSIM R drifts (0.68 vs 0.74) but same p=0.001 and same strong-separation conclusion. Deviations: R dada2 instead of q2-dada2 (same algorithm), truncLen 270/210 (paper unspecified). NOT attempted (out of scope): MST stochasticity / C-score niche-assembly modelling (30000 MST runs), wet-lab steps. No fabrication signal — every reported number is derivable from the public data.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 84
    assessed: 2026-06-22 ⛓ 8e9517c771f3
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-21
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests which factors (sex, host genetics/wild vs. domestic status, and physical interaction from mixed grouping) are the main drivers shaping gut microbiota diversity and composition in Datong yaks raised together in a mixed group, and how ecological assembly processes differ among domestic males, domestic females, and wild males.

Core claims
  • Gut microbial diversity (alpha and beta) differs significantly among domestic males, domestic females, and wild male Datong yaks. finding
  • Wild males have the highest gut microbial alpha-diversity, followed by domestic females, then domestic males. finding
  • Mixed grouping (physical interaction with wild males) contributes to improved gut microbial diversity in domestic females, since domestic females and wild males show no significant diversity differences. finding
  • Firmicutes and Bacteroidota are the dominant gut phyla (>96% combined relative abundance) across all three yak groups. finding
  • The dominant ecological assembly process for gut microbiota is stochastic (MST > 0.5) in all three groups, with domestic males showing the strongest deterministic influence (highest SES) and wild males the weakest. finding
  • Different factors (sex, host genetics, physical interaction) dominate gut microbiota differences depending on which pair of groups is compared. finding
  • 16S rRNA V3-V4 sequencing combined with QIIME2/DADA2 pipeline and NST/EcoSimR-based ecological assembly analysis was used to characterize yak fecal microbiota. method
Experimental setups
Assay System Perturbation Readout Platform
16S rRNA gene sequencing (V3-V4 regions) fresh fecal samples from Datong yaks (domestic males, domestic females, wild males), Qinghai-Tibet Plateau none (comparison of naturally differing groups: sex, domestication status) gut microbial community composition and diversity (ASV counts, taxonomic abundance) Illumina MiSeq PE300
Alpha-diversity analysis (Shannon, Simpson indices) fecal microbiota ASV table from domestic males, domestic females, wild males (n=10,10,6) none Shannon and Simpson diversity indices per group, compared via Kruskal-Wallis H test and Tukey-Kramer post-hoc test QIIME2 2020.2; R/Rstudio (Multcomp package)
Beta-diversity analysis (Bray-Curtis distance, ANOSIM, PERMANOVA) fecal microbiota across the three yak groups none inter-group vs intra-group community dissimilarity (R and p values) QIIME2 q2-diversity-lib plugin; R package 'vegan' and 'ggplot2'
Taxonomic composition comparison (phylum, family, genus level) fecal microbiota, pairwise between domestic males, domestic females, wild males none relative abundance differences via Wilcoxon rank-sum test R package 'stats'; SILVA SSU NR99 v138 database classifier
Ecological assembly process analysis (modified stochasticity ratio, MST/NST) fecal microbiota communities of domestic males, domestic females, wild males none contribution of stochastic vs. deterministic assembly processes (MST value) NST package in R/Rstudio (30,000 runs)
Null model / standardized effect size (SES) and C-score analysis fecal microbiota communities of the three yak groups none clustering vs. overdispersion of community assembly (SES, C-score) EcoSimR package in R/Rstudio (30,000 simulations, sequential swap randomization)
Key results
  • Wild males had highest alpha-diversity (Shannon=6.12±0.14; Simpson=0.0047±0.0008), then domestic females (Shannon=6.01±0.09; Simpson=0.0054±0.0007), then domestic males (Shannon=5.70±0.13; Simpson=0.0089±0.0023)
  • No significant difference in gut microbial diversity between domestic females and wild males p ≥ 0.05
  • Significant beta-diversity differences among all three groups overall (ANOSIM) R=0.74, p=0.001
  • Beta-diversity significantly differed between domestic males vs wild males and domestic males vs domestic females, but not between domestic females and wild males R=1, p=0.001 (both); R=-0.05, p=0.66 (females vs wild males)
  • Total shared ASVs among all three groups; highest shared ASVs between domestic females and wild males; most group-specific ASVs in domestic males 1000 shared ASVs total; 434 shared (females/wild males); 136 specific to domestic males
  • Firmicutes and Bacteroidota combined dominate gut phyla in all groups, no significant Firmicutes difference among groups >96% combined relative abundance; p>0.05 for Firmicutes
  • MST values above 0.5 in all three groups indicate stochastic process dominance; domestic males show highest SES (strongest deterministic influence), wild males show weakest MST > 0.5
Key statistics
  • count 26 fresh fecal samples (10 domestic males, 10 domestic females, 6 wild males) (sample collection)
  • pvalue p < 0.05 (significant gut microbial diversity differences among the three groups (alpha-diversity))
  • correlation R = 0.74, p = 0.001 (ANOSIM beta-diversity comparison among all three groups)
  • correlation R = 1, p = 0.001 (ANOSIM beta-diversity, domestic males vs wild males and domestic males vs domestic females)
  • correlation R = -0.05, p = 0.66 (ANOSIM beta-diversity, domestic females vs wild males (not significant))
  • mean Shannon = 6.12 ± 0.14, Simpson = 0.0047 ± 0.0008 (alpha-diversity in wild males)
  • mean Shannon = 6.01 ± 0.09, Simpson = 0.0054 ± 0.0007 (alpha-diversity in domestic females)
  • mean Shannon = 5.70 ± 0.13, Simpson = 0.0089 ± 0.0023 (alpha-diversity in domestic males)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a cross-sectional, three-group comparison (domestic male, domestic female, and wild male yaks; n=10, 10, 6) of 16S rRNA V3–V4 gut microbiota profiles. Alpha-diversity (Shannon, Simpson) was compared using the Kruskal–Wallis H test with Tukey–Kramer post-hoc testing, taxon-level relative abundances were compared pairwise with the Wilcoxon rank-sum test, and beta-diversity (Bray–Curtis distances) was assessed with ANOSIM and PERMANOVA. Ecological assembly processes were characterized using a normalized stochasticity ratio (NST/MST) and standardized effect size (SES) from null-model simulations. Results were reported mainly as means ± dispersion values with significance thresholds (p<0.05, p<0.01) and some exact p/R values.

Replicationbiological Sample sizeGroup sizes stated as 10 domestic males, 10 domestic females, and 6 wild males (total 26 fecal samples); no power analysis or sample-size justification described Groupsdomestic males vs. domestic females vs. wild males (sex and domestication/rearing status) Pairingunpaired Randomization/blindingnot stated Dispersionunclear Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionTukey-Kramer post-hoc test was applied following the Kruskal-Wallis test for alpha-diversity group comparisons; no correction method is described for the many pairwise Wilcoxon rank-sum tests performed across phylum-, family-, and genus-level taxa
Statistical tests used
Test Applied to n Assumptions
Kruskal-Wallis H test with Tukey-Kramer post-hoc test alpha-diversity (Shannon, Simpson indices) comparisons among domestic females, domestic males, and wild males 10 domestic males, 10 domestic females, 6 wild males not stated
Wilcoxon rank-sum test pairwise comparisons of taxon relative abundance (phylum, family, genus levels) between group pairs 10 domestic males, 10 domestic females, 6 wild males not stated
ANOSIM (analysis of similarities) based on Bray-Curtis distances beta-diversity comparisons among and between the three groups 10 domestic males, 10 domestic females, 6 wild males not stated
PERMANOVA (permutational multivariate analysis of variance) based on Bray-Curtis distances beta-diversity comparisons among and between the three groups (Appendix Tables A1-A4) 10 domestic males, 10 domestic females, 6 wild males not stated
Normalized stochasticity ratio (NST/MST) with 30,000 null-model runs quantifying stochastic vs. deterministic ecological assembly processes per group not stated per group not stated
Standardized effect size (SES) via C-score with sequential swap randomization (30,000 simulations) assessing clustering/overdispersion of gut microbiota assemblages not stated not stated
Approaches that could also have been used
  • Pairwise taxon-level comparisons (phylum, family, genus) were each evaluated with the Wilcoxon rank-sum test at a nominal p<0.05 threshold across many taxa without a stated multiplicity correction.
    Could also: A false discovery rate correction such as Benjamini-Hochberg could also be applied across the family of taxon comparisons. — Since many taxa are tested simultaneously, an FDR-based correction would help control the expected proportion of false positives among the significant taxa, which is a common approach in microbiome differential abundance analyses.
  • Beta-diversity comparisons relied on Bray-Curtis distances for ANOSIM and PERMANOVA.
    Could also: Phylogenetically informed distance metrics such as weighted or unweighted UniFrac could also be used alongside or instead of Bray-Curtis. — UniFrac distances incorporate evolutionary relatedness among taxa, which can provide a complementary perspective on community structure differences driven by phylogenetically related lineages.
  • Alpha-diversity indices (Shannon, Simpson) were compared with the nonparametric Kruskal-Wallis test and Tukey-Kramer post-hoc test.
    Could also: A generalized linear model or ANOVA framework with group as a fixed effect, if distributional assumptions were checked and met, could also be used. — A parametric model can allow additional covariates (e.g., age, body condition) to be incorporated directly, which can help account for other sources of variation alongside sex/domestication status.
  • Group sizes were relatively small and unbalanced (10, 10, and 6 samples) without a stated power analysis or sample-size justification.
    Could also: A prospective power analysis, or reporting effect sizes with confidence intervals, could also be included. — This would help convey the precision of the diversity and abundance estimates given the modest, unequal group sizes.
  • Diversity index values were reported as mean ± a dispersion value without specifying whether this is standard deviation or standard error of the mean.
    Could also: Explicitly reporting SD, SEM, or a 95% confidence interval could also be used. — Specifying the exact dispersion measure removes ambiguity about how much of the spread reflects biological variability (SD) versus estimation uncertainty (SEM/CI), aiding interpretation and comparison with other studies.
  • Ecological assembly processes were classified using a fixed NST/MST threshold of 0.5 to designate stochastic versus deterministic dominance.
    Could also: Complementary frameworks such as iCAMP or neutral community model fitting could also be applied. — These approaches can partition assembly processes into finer categories (e.g., dispersal limitation, homogeneous selection) and may provide additional resolution beyond a single threshold-based ratio.
Software: QIIME2 2020.2 · FLASH v1.2.11 · fastp 0.19.6 · DADA2 (q2-dada2 plugin) · RESCRIPt · SILVA SSU NR99 database 138 · R/Rstudio (packages: stats, multcomp, vegan, ggplot2) · NST package (R) · EcoSimR package (R) · Majorbio Cloud Platform

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35883324

Paper: Main Factors Influencing the Gut Microbiota of Datong Yaks in Mixed Group. Qin W. et al., Animals (Basel) 2022. PMID 35883324 · PMCID PMC9312300 · DOI 10.3390/ani12141777.

Data: SRA BioProject PRJNA825400 — 16S rRNA V3–V4 amplicon, Illumina MiSeq PE300. 26 fecal samples: domestic males n=10, domestic females n=10, wild males n=6. Primers 338F (ACTCCTACGGGAGGCAGCAG) / 806R (GGACTACHVGGGTWTCTAAT).

Pipeline described in Methods (all bioinformatic → in scope):

  1. fastp 0.19.6 — quality control of raw reads. (This is the repo listed in the brief: github.com/OpenGene/fastp — third-party tool applied to the paper's data, valid per P16.)
  2. FLASH v1.2.11 — paired-end read merging.
  3. DADA2 via q2-dada2 in QIIME2-2020.2 — denoising → ASVs.
  4. RESCRIPt + SILVA SSU NR99 v138 (0.8 confidence) — taxonomy classification.
  5. R / vegan / ggplot2 — alpha & beta diversity, statistics. ASV filter: drop ASVs with <0.01% rel-abundance OR present in <5 samples.

In-scope reproducible results (pipeline-derived)

id result reported
C1 N samples in BioProject 26 (10 DM, 10 DF, 6 WM)
C2 Shannon per group (mean±sd) WM 6.12±0.14; DF 6.01±0.09; DM 5.70±0.13
C3 Simpson per group (mean±sd) WM 0.0047±0.0008; DF 0.0054±0.0007; DM 0.0089±0.0023
C4 Beta diversity ANOSIM R=0.74, p=0.001
C5 Dominant phyla Firmicutes + Bacteroidota >96% combined; no sig diff among groups (p>0.05)
C6 Top families Oscillospiraceae >16%, Rikenellaceae, Lachnospiraceae, Christensenellaceae (top5 >7%)
C7 Top genera UCG-005 >11%, Christensenellaceae_R-7_group, Rikenellaceae_RC9_gut_group (top5 >7%)
C8 Shared ASVs across groups ~1000 shared

Out of scope

  • Wet-lab (DNA extraction, library prep, MiSeq sequencing) — not computational.
  • MST stochasticity (30,000 runs) & C-score / EcoSimR — niche-assembly modelling; attempt only if core pipeline succeeds.
  • Kruskal–Wallis / Tukey / Wilcoxon significance tests — derivable once diversity tables exist; secondary.

Reproduction plan («our HPC»/SLURM, data on «infra»)

  1. front1: prefetch/fasterq-dump PRJNA825400 26 runs → «infra»; verify N=26.
  2. conda env on front1: fastp 0.19.6, FLASH 1.2.11, qiime2-2020.2.
  3. SLURM compute job: fastp QC → (FLASH or DADA2 paired) → DADA2 denoise → ASV table.
  4. Classify with SILVA 138; collapse to phylum/family/genus.
  5. R/vegan: Shannon, Simpson, ANOSIM; compare to C2–C7.
  6. Fill claims.tsv + agreement.json + AUDIT.md + dataset_profile.json.

Heavy compute = «our HPC» only. All data on «infra». «host» = results only.

Figures / tables: Table
C1
Reported
26 samples (DM=10, DF=10, WM=6)
Reproduced
26 runs; library_name maps 10 DF / 10 DM / 6 WM
exact
C2
Reported
Shannon WM 6.12±0.14; DF 6.01±0.09; DM 5.70±0.13
Reproduced
WM 6.146±0.128; DF 6.044±0.068; DM 5.731±0.116
within tolerance
C3
Reported
Simpson D WM 0.0047±0.0008; DF 0.0054±0.0007; DM 0.0089±0.0023
Reproduced
WM 0.00473±0.00079; DF 0.00544±0.00067; DM 0.00929±0.00223
within tolerance
C4
Reported
ANOSIM R=0.74 p=0.001
Reproduced
R=0.679 p=0.001
partial
C5
Reported
Firmicutes+Bacteroidota >96%
Reproduced
97.9% (Firmicutes 74.5% + Bacteroidota 23.5%)
exact
C6
Reported
Oscillospiraceae >16%; top5 families >7%
Reproduced
Oscillospiraceae 26.1%; Rikenellaceae 8.9%; Lachnospiraceae 7.4%; Christensenellaceae 7.2%
within tolerance
C7
Reported
Top genus UCG-005 >11%
Reproduced
UCG-005 20.1% (top named genus)
within tolerance
C8
Reported
~1000 ASVs shared across groups
Reproduced
935 shared ASVs (of 1520 filtered)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 84/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is not an authors' or fabrication problem — it is an incomplete reproduction on our side. The input data (SRA PRJNA825400, 16S V3-V4) is fully public and 1:1 reproducible, and all six reported claims (Shannon WM 6.12±0.14, Simpson, ANOSIM R=0.74 p=0.001, Firmicutes+Bacteroidota >96%, UCG-005 >11%) are concrete and directly comparable. However, ROOM_RESULT is 'IN PROGRESS' / agreement.json 'not-run-yet' with every reproduced field empty because the «our HPC» run never completed. No deviation can be localized or sized, so q3–q8 are graded yellow (untested, our-side gap) rather than red, to avoid wrongly penalizing the authors.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

256.4 k
tokens (I/O) · 15.4 M incl. cache
96 min
runtime · 7.25 CPU-h
28.3 GB
peak RAM
1
HPC jobs
hummel
machine