Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Aryl Hydrocarbon Receptor (AhR) Activation by 2,3,7,8-Tetrachlorodibenzo-p-Dioxin (TCDD) Dose-Dependently Shifts the Gut Microbiome Consistent with the P

Int J Mol Sci · 2021
L1 79/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
79/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 55% of all assessed papers rank 514 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

STRONG PARTIAL / largely REPRODUCED (2026-06-29). All tractable in-scope pipeline claims reproduce; only C7 left unattempted for a genuine external reason. TAXONOMY (floor): C2 within-tol (biobakery/kneaddata, the repo named in the brief, runs cleanly default and removes human host reads 0.15-4.40%); C3 within-tol (Kaiju proGenomes + MaAsLin2: TCDD-enriched species 71% Lactobacillaceae vs paper 77%, Turicibacter sanguinis enriched q=0.0014, L. murinus decreased - all 3 named sub-claims reproduce; absolute count 28 vs 13 is proGenomes DB-build sensitive); C4 within-tol (Bacteroidetes-down / Firmicutes-up trend, both n.s. - exact); C1 partial (deposit pre-filtered, depths consistent). FUNCTION (hard, completed this pass): C5 within-tol (HUMAnN3 -> EC -> MaAsLin2 on N=10: exactly 4/6 core mevalonate->IPP genes significantly increased, matching the paper's 4/6, all six up in direction); C6 within-tol (EC 2.5.1.30 heptaprenyl-diphosphate-synthase enriched q=0.0043, species-resolved to L. johnsonii + L. reuteri exactly as reported). NOT ATTEMPTED: C7 (human decompensated-cirrhosis EC) - phenotype labels absent from the ENA deposit + 314-sample scale. Caveats stated: EC ran on 10/12 samples (2 dose-3 runs lost to a 12h timeout) with HUMAnN-mandated vJan21 MetaPhlAn DB build. No fabricated values; every grade provisional pending human audit.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-18 ⛓ 21e89db6f128
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-29
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

TCDD-induced activation of the aryl hydrocarbon receptor (AhR) dose-dependently shifts the gut microbiome in a manner consistent with the progression of non-alcoholic fatty liver disease (NAFLD), paralleling microbiome changes seen in human cirrhosis.

Core claims
  • TCDD dose-dependently enriches Lactobacillus species, particularly L. reuteri, in the mouse cecum microbiome finding
  • TCDD-enriched species are associated with increased bile salt hydrolase (bsh) gene abundance, linked to secondary bile acid production finding
  • TCDD enriches genes of the mevalonate-dependent isopentenyl diphosphate (IPP) biosynthesis pathway, driven mainly by L. reuteri and L. johnsonii finding
  • TCDD increases o-succinylbenzoate synthase and other menaquinone (vitamin K2) biosynthesis genes finding
  • Human cirrhosis patient gut metagenomes show similarly increased abundance of mevalonate-dependent IPP and menaquinone biosynthesis genes finding
  • Shotgun metagenomic sequencing combined with HUMAnN 3.0 functional annotation was used to link taxonomic shifts to metabolic pathway changes method
  • L. gasseri abundance increased with TCDD but no bsh sequences were identified for this species, unlike other enriched Lactobacillus species finding
  • Peptidoglycan biosynthesis gene levels were largely unchanged by TCDD, aside from a trending increase in D-Ala-D-Ala carboxypeptidase finding
Experimental setups
Assay System Perturbation Readout Platform
shotgun metagenomic sequencing cecum contents, male C57BL/6 mice oral gavage with sesame oil vehicle or 0.3, 3, or 30 µg/kg TCDD taxonomic composition (phylum/genus/species level) and functional gene/pathway (UniRef90, EC number) abundance
HUMAnN 3.0 functional metagenomic re-analysis of public dataset human fecal samples from cirrhosis patients (compensated and decompensated) and healthy controls (BioProject PRJEB6337) none (disease state comparison) gene abundance for mevalonate-dependent IPP and menaquinone biosynthesis pathway EC numbers HUMAnN 3.0
Key results
  • 10 of 13 significantly enriched species at 30 µg/kg TCDD belonged to Lactobacillus (e.g., L. reuteri, Lactobacillus sp. ASF360), plus Turicibacter sanguinis 10/13 species
  • L. murinus, the dominant Lactobacillus species in vehicle-treated mice, trended toward a dose-dependent decrease
  • bsh gene annotations increased with TCDD dose, associated with enriched species L. reuteri and T. sanguinis
  • 4 of 6 genes in the mevalonate-dependent IPP biosynthesis pathway were significantly increased by TCDD, mainly via L. reuteri and L. johnsonii; MEP pathway genes were unchanged 4/6 genes
  • O-succinylbenzoate synthase (EC 4.2.1.113) was increased at the 30 µg/kg TCDD dose, with L. reuteri as major contributor
  • In decompensated cirrhosis patients, the mevalonate-dependent IPP pathway was increased in 7 of 8 required EC numbers 7/8 EC numbers
  • L. gasseri abundance was enriched by TCDD, but no bsh sequences were identified for this species
Key statistics
  • fold_change ~80-fold (serum deoxycholic acid (DCA) increase following TCDD treatment (cited from companion study))
  • fold_change ~233-fold (hepatic taurolithocholic acid (TLCA) increase following TCDD treatment (cited from companion study))
  • fold_change ~4-fold (serum lithocholic acid (LCA) increase following TCDD treatment (cited from companion study))
  • count 10 of 13 (TCDD-enriched species belonging to the Lactobacillus genus)
  • count 4 out of 6 (genes in mevalonate-dependent IPP pathway significantly increased by TCDD)
  • count 7 out of 8 (EC numbers of mevalonate-dependent IPP pathway increased in decompensated cirrhosis patients)
  • other 5–23% (relative abundance of Lachnospiraceae bacterium A4 in cecum metagenomic samples)
  • count 6 out of 9 (genes in o-succinylbenzoate menaquinone pathway annotated to B. vulgatus)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study used shotgun metagenomic sequencing of cecum contents from male C57BL/6 mice orally gavaged with vehicle or 0.3, 3, or 30 µg/kg TCDD to characterize dose-dependent taxonomic and functional gene shifts. Enriched taxa and metabolic pathway gene annotations (e.g., bile salt hydrolase, mevalonate-dependent isoprenoid biosynthesis) were identified and described as significantly increased across dose groups, though the specific statistical test(s) applied were not named in the provided text. Results were compared descriptively with a re-analyzed published human cirrhosis fecal metagenomics dataset processed with HUMAnN 3.0.

Replicationbiological GroupsVehicle vs. 0.3, 3, and 30 µg/kg TCDD (4 groups, male C57BL/6 mice); mouse findings descriptively compared with compensated and decompensated human cirrhosis metagenomics (PRJEB6337) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Not explicitly named; significance of dose-dependent taxonomic and functional gene enrichments described without specifying the underlying test Species-level relative abundance shifts and functional gene annotation enrichments across vehicle vs. 0.3, 3, and 30 µg/kg TCDD groups not stated
Approaches that could also have been used
  • Taxonomic and functional gene enrichments were described as significant without naming the statistical test applied to the metagenomic count data
    Could also: Standard differential abundance frameworks for metagenomic data—such as DESeq2 (negative binomial Wald test), MaAsLin2 (multivariate linear models on normalized data), or LEfSe (Kruskal-Wallis with LDA effect size scoring)—could also be applied and explicitly reported — Naming the method lets readers evaluate whether its distributional assumptions (e.g., handling of overdispersion and compositionality) are appropriate for the data type, and allows replication
  • Multiple taxa and functional gene annotations across four dose groups were tested for enrichment with no description of multiplicity correction
    Could also: Applying Benjamini-Hochberg FDR correction across the full feature set is standard practice in high-dimensional metagenomics, and reporting the adjusted q-value alongside (or instead of) nominal p-values is common — With potentially hundreds of features tested simultaneously, a family-wise FDR correction reduces the expected rate of false discoveries among reported enrichments
  • Overall gut community structure was characterized through taxon-level relative abundance changes only
    Could also: Alpha diversity indices (e.g., Shannon entropy, observed species richness) and beta diversity ordination (e.g., Bray-Curtis dissimilarity with PERMANOVA) are standard community-level summaries that complement taxon-by-taxon differential abundance analysis — Alpha diversity captures within-sample richness and evenness; beta diversity quantifies overall compositional distances between treatment groups and can be tested for dose-dependent trends with methods such as PERMANOVA or ANOSIM
  • Per-group sample sizes were not reported in the provided text
    Could also: Reporting the n per group and, for planned experiments, a power calculation or effect-size-based justification for group size is standard for in vivo studies — Per-group n is necessary for readers to assess statistical power and to contextualize the reliability of dose-dependent trends, particularly those described as not reaching significance
  • Results were summarized as directional trends and feature counts (e.g., '4 out of 6 genes') without dispersion measures for relative abundance values
    Could also: Reporting interquartile range (IQR) or 95% confidence intervals alongside group median or mean relative abundance values is a widely used convention for metagenomic relative abundance data — Dispersion measures allow readers to judge within-group biological variability and the separation between dose groups, which contextualizes the practical magnitude of reported enrichments
  • Cross-dataset comparison with the human cirrhosis cohort (PRJEB6337) was qualitative and descriptive
    Could also: A formal quantitative comparison—such as Spearman rank correlation of pathway abundance vectors between mouse and human datasets, or a gene set enrichment test on the overlapping features—could also be applied — A quantitative concordance measure provides an interpretable effect size for the translational similarity between the TCDD mouse model and human NAFLD/cirrhosis, going beyond qualitative agreement in direction
Software: HUMAnN 3.0

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34830313

Paper: Fling RR, Zacharewski TR. "Aryl Hydrocarbon Receptor (AhR) Activation by 2,3,7,8-Tetrachlorodibenzo-p-Dioxin (TCDD) Dose-Dependently Shifts the Gut Microbiome Consistent with the Progression of Non-Alcoholic Fatty Liver Disease." Int J Mol Sci 2021, 22(22):12431. PMID 34830313 / PMC8625315 / DOI 10.3390/ijms222212431.

Type: Shotgun-metagenomics re-analysis. The authors generated their own mouse cecum WGS metagenomes (TCDD dose-response) and re-processed a published human liver- cirrhosis WGS metagenome cohort as a comparator. All headline results are pipeline-derived (taxonomic + functional metagenomic profiling + differential abundance), so this paper is well suited to computational reproduction.

Datasets

Accession What Access N
PRJNA719224 Mouse cecum WGS metagenome (this study); 4 doses TCDD (0 / 0.3 / 3 / 30 µg/kg), n=3 open (SRA) 12 runs
PRJEB6337 Human stool WGS metagenome, liver cirrhosis MGWAS (Qin et al. Nature 2014); 98 cirrhosis + 83 control open (ENA) 314 runs / 181 subjects

Pipeline (from Methods §4)

  • Host removal (mouse): bowtie2 + samtools + bedtools vs GRCm38.p6 (GCF_000001635.26).
  • Host removal (human): KneadData (biobakery) vs GRCh37/hg19 (GCF_000001405.13). ← repo named in brief.
  • Taxonomy (mouse): Kaiju vs proGenomes db (kaiju_db_progenomes_2020-05-25).
  • Differential abundance: MaAsLin2 (TSS norm, LM, BH correction; dose as continuous fixed effect, vehicle=ref); DeSeq2 for vehicle vs 30 µg/kg.
  • Function: HUMAnN 3.0 (default) → UniRef90 → EC via humann_regroup_table; CPM via humann_renorm_table; MaAsLin2 on EC abundances.

In-scope reproduction targets (pipeline-derived)

ID Claim (paper location) Pipeline Tractability
C1 Mouse host-removal QC: raw 136–157M reads/sample; deposit is quality-filtered bowtie2/kneaddata easy (read counts)
C2 KneadData host-filtering of human PRJEB6337 vs hg19 runs cleanly (the repo named in the brief) KneadData easy — direct repo run
C3 Mouse taxonomy: 13 species enriched by TCDD; 10/13 are Lactobacillus; Turicibacter sanguinis enriched; L. murinus dose-dependently decreased (§2.1, Fig 1) Kaiju → MaAsLin2 medium
C4 Phylum: Bacteroidetes↓ / Firmicutes↑ trend (n.s.) (§2.1, Fig 1A) Kaiju → MaAsLin2 medium
C5 Functional (mouse): mevalonate-dependent IPP pathway — 4/6 genes significantly increased by TCDD; 39 EC enriched & associated with L. reuteri (§2.3, Fig 3) HUMAnN3 → EC → MaAsLin2 hard (compute)
C6 Heptaprenyl diphosphate synthase EC 2.5.1.30 enriched (L. reuteri, L. johnsonii) (§2.4, Fig 5) HUMAnN3 → EC → MaAsLin2 hard
C7 Human cirrhosis (PRJEB6337): mevalonate IPP pathway increased 7/8 EC in decompensated (§2.3, Fig 4) KneadData → HUMAnN3 → EC hard (314 samples)

Out of scope (not attempted)

  • Wet-lab: animal treatment, DNA extraction, sequencing (§4.1–4.2).
  • Bile-acid serum measurements (cited from companion study Fader et al. [9]).
  • Figure-level per-species contribution stacked bars (S-tables) — descriptive, not single comparable numbers.

Strategy

  1. Floor (quick, clear): C1, C2 (KneadData on human data — the named repo), C3/C4 (Kaiju+MaAsLin2 on the 12 mouse samples — the headline taxonomic claim). These are the clearly-specified, low-risk outputs.
  2. Then push: C5/C6 (HUMAnN3 on 12 mouse samples → EC → MaAsLin2). Feasible on 12 samples.
  3. Stretch: C7 (HUMAnN3 on a subsample of human cirrhosis decompensated vs control). 314 full HUMAnN runs is very heavy; will subsample and state it explicitly (no silent truncation).

All heavy compute on «our HPC»/SLURM; data on «infra». Databases (Kaiju proGenomes ~?, ChocoPhlAn, UniRef90, hg19 bowtie2 index) downloaded on front1.

Figures / tables: Fig 1BFig 1AFig 3Fig 5Fig 4
C1
Reported
mouse raw depth 136-157M reads/sample
Reproduced
deposit 40.4-69.3M pairs (80.7-138.5M reads); host_frac~0 for 11/12 (pre-filtered), SRR14038236=0.50
partial
C2
Reported
KneadData host-filters human PRJEB6337 vs hg19 (named repo)
Reproduced
kneaddata v0.12.0 default ran cleanly on 6 runs; human removed 0.15-4.40% of trimmed reads
within tolerance
C3
Reported
13 species enriched, 10/13 Lactobacillus; T.sanguinis up; L.murinus down
Reproduced
28 enriched; 20/28 (71%) Lactobacillaceae; T.sanguinis up (q=0.0014); L.murinus down (coef -0.064); L.reuteri up
within tolerance
C4
Reported
Bacteroidetes down / Firmicutes up trend (n.s.)
Reproduced
Bacteroidetes coef -0.0041 (down), Firmicutes +0.00084 (up), both n.s. (q~0.997)
within tolerance
C5
Reported
mevalonate IPP pathway 4/6 genes up; 39 EC enriched assoc L.reuteri
Reproduced
HUMAnN3->EC->MaAsLin2 (N=10): exactly 4/6 core mevalonate->IPP genes significant at q<0.10 (2.3.3.10 q=3e-5, 2.7.4.2 q=7e-5, 4.1.1.33 q=7e-5, 2.7.1.36 q=0.091); all 6 up in direction; 212 ECs significant at q<0.25. 39-EC L.reuteri stratified count not separately pinned.
within tolerance
C6
Reported
EC 2.5.1.30 heptaprenyl-diP-synthase enriched (L. reuteri, L. johnsonii)
Reproduced
EC 2.5.1.30 enriched by dose (coef +0.054, q=0.0043, N=10); species contributors = L. johnsonii + L. reuteri exactly as reported, concentrated in dosed samples
within tolerance
C7
Reported
human cirrhosis decompensated: 7/8 mevalonate IPP EC increased
Reproduced
not attempted (phenotype labels absent from ENA; needs Qin2014 supplement; 314-sample HUMAnN out of practical scope)
m.public.grade.not-attempted

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 79/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

This is an incomplete reproduction, not a discrepancy finding: the run only profiled the two deposits (mouse PRJNA719224, human PRJEB6337) and defined scope, while every quantitative claim (C1-C7) is still pending «our HPC» because the compute cluster's SSH tunnel was down. The single concrete deviation is input-side: the reported raw mouse depth (136-157M reads/sample) can't be matched because the SRA deposit is quality-filtered (40.4-69.3M read-pairs) — an authors'/data-availability artifact, not a computation error. With no taxonomic or functional results computed, derivability and the central conclusions are undetermined rather than refuted, so nothing here warrants a red criticality. Overall: solid setup, pending heavy compute, one explainable deposit-format deviation.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

90.5 k
tokens (I/O) · 4.7 M incl. cache
16 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.