Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Verrucomicrobia are prevalent in north-temperate freshwater lakes and display class-level preferences between lake habitats.

PLoS One · 2018
L1 99/100 PQI 98
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
99/100
Reproducibility score
1.4 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 55 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> clean 1:1. Reproduced the headline computational results of the paper's downstream R analysis by re-running the authors' verruco-analysis.Rmd code (copied verbatim) on the shipped VerrucoData.RData phyloseq object (repo DenefLab/Verruco @ 8455714) on «our HPC» using an existing R 4.3.3 + phyloseq 1.46 + vegan 2.6.8 conda env. 15/15 graded values match: 14 exact, 1 within-tol (PERMANOVA R2 0.112 vs authors' 0.10811 because vegan's adonis() was replaced by adonis2(); p-value exact, both round to the paper's 0.11). Every authors' inline-documented p-value reproduced to all printed digits -> no fabrication signal. NOT attempted (the ~20%): the upstream mothur raw-reads->OTU pipeline (OTU table ships; trusted as-is), plus DESeq2, phylogenetic-signal (geiger/picante), bioenv (Table 3), multiple-linear-regression model selection (Table 1), Procrustes, and the IToL tree (web GUI) -- skipped per 80/20 due to heavy/finicky deps or manual steps; rationale in scope.md.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 99
    assessed: 2026-06-16 ⛓ ae18b6245b9b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

What physical and geochemical drivers determine the relative abundance and within-phylum community composition of the bacterial phylum Verrucomicrobia in north-temperate freshwater lakes, and how is habitat differentiation structured within the phylum? Given the phylum's diverse metabolism, the authors expected distinct relative abundances and community compositions across different lake environments.

Core claims
  • Verrucomicrobia is highly prevalent in north-temperate freshwater lakes, on average the 4th most abundant phylum (range 1.7–41.7%). finding
  • Within-phylum habitat preference between lake habitats is phylogenetically conserved at the class level. finding
  • Fraction, season, station, and depth explain up to ~70% of the variance in Verrucomicrobia community composition. finding
  • A majority of the phylum exhibits preference for the particle-associated fraction, and two classes (Opitutae and Verrucomicrobiae) are more abundant in spring. finding
  • Verrucomicrobia and non-Verrucomicrobia bacterial community composition correlate to similar quantitative environmental parameters, with lake-system-dependent differences and >55% of variance unexplained. finding
  • 16S rRNA gene V4 hypervariable region sequencing was used to characterize Verrucomicrobia distribution and diversity across lake systems. method
  • Verrucomicrobia relative abundance is significantly lower in Laurentian (Lake Michigan) samples than in estuary and inland samples. finding
  • mothur output files, metadata, and code to replicate analyses are publicly available as a resource. resource
Experimental setups
Assay System Perturbation Readout Platform
16S rRNA gene amplicon sequencing (V4 hypervariable region) 12 southeastern Michigan inland lakes (water, free-living and particle-associated fractions) none relative abundance and community composition of Verrucomicrobia and total bacteria
16S rRNA gene amplicon sequencing (V4 hypervariable region) Lake Michigan (Laurentian Great Lake) near-to-offshore transect, water none Verrucomicrobia relative abundance and community composition Joint Genome Institute sequencing center
16S rRNA gene amplicon sequencing (V4 hypervariable region) Muskegon Lake freshwater estuary, water and sediment none Verrucomicrobia relative abundance, diversity, and community composition University of Michigan sequencing center
Within-phylum alpha diversity (inverse Simpson index) All lake-type samples (water and sediment) none Verrucomicrobia within-phylum diversity
Phylogenetic diversity (standardized effect size mean pairwise distance, SES MPD) All lake-type samples (water and sediment) none phylogenetic clustering/evenness of Verrucomicrobia communities
Multiple linear regression modeling Laurentian, estuary, and inland lake systems with environmental data none Verrucomicrobia relative abundance vs geochemical parameters
PERMANOVA / nested PERMANOVA and Procrustes analysis Laurentian, estuary, inland samples none variance in community composition explained by categorical factors; correlation between Verrucomicrobia and non-Verrucomicrobia ordinations
bioenv analysis on Bray-Curtis dissimilarity / PCoA Lake Michigan, estuary, and inland lake bacterial communities none physicochemical parameters correlating with community composition shifts
Key results
  • Verrucomicrobia was the 4th most abundant phylum with median relative abundance 9.3% (range 1.7–41.7%) median 9.3%, range 1.7–41.7%
  • Verrucomicrobia relative abundance significantly lower in Laurentian (4.7±3.0%) vs estuary (11.0±9.4%) and inland (10.3±8.5%) samples 4.7% vs 11.0% vs 10.3%
  • Fraction, season, station, and depth explained up to ~70% of Verrucomicrobia community composition variance (e.g. ~60% in Laurentian) up to 70%
  • Estuary Verrucomicrobia and non-Verrucomicrobia ordinations most strongly correlated (Procrustes correlation 0.96), then Laurentian (0.80) and inland (0.77) r=0.96, 0.80, 0.77
  • Estuary multiple linear regression model for relative abundance had R2=0.60; temperature negatively related (coefficient -0.97 individually) R2=0.60; coeff=-0.97
  • Inland surface samples had significantly higher relative abundance than bottom samples; fall higher than spring/summer
  • Verrucomicrobia relative abundance higher in PA than FL in Laurentian, but opposite in inland samples; sediment lower than water in estuary
  • Sediment samples harbored more diverse Verrucomicrobia communities than water; estuary summer less diverse than spring and fall
Key statistics
  • count 228 sequencing data sets generated; mean 26,850 reads/sample (range 1–439,926) (V4 16S rRNA sequencing output)
  • other median relative abundance 9.3% (range 1.7–41.7%) (Verrucomicrobia relative abundance across all samples)
  • pvalue p < 0.001 (Kruskal-Wallis: Verrucomicrobia relative abundance differs among Laurentian/estuary/inland)
  • pvalue p = 0.033 (KW: inland surface vs bottom relative abundance)
  • pvalue p = 0.001 (KW: inland fall higher than spring/summer relative abundance)
  • correlation Procrustes correlation 0.96 (estuary), 0.80 (Laurentian), 0.77 (inland), all p=0.001 (Verrucomicrobia vs non-Verrucomicrobia ordination correlation)
  • correlation R2 = 0.60 (estuary MLR); adjusted R2 = 0.27 for temperature alone, p=0.002 (multiple linear regression of relative abundance vs environment)
  • other PERMANOVA lake type R2 = 0.11 (Verrucomicrobia), 0.10 (non-Verrucomicrobia), p=0.001; season R2=0.433 Laurentian Verruco (variance in community composition by lake type and factors)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is an observational 16S rRNA amplicon survey of Verrucomicrobia across 14 freshwater systems, with samples stratified by lake type, season, depth, and free-living vs particle-associated fraction. Relative abundance and diversity (inverse Simpson, SES MPD) differences across categorical groups were assessed with non-parametric Kruskal-Wallis tests, community composition was analyzed with nested PERMANOVA and PCoA on Bray-Curtis dissimilarity, Procrustes tests compared Verrucomicrobia vs non-Verrucomicrobia ordinations, and bioenv plus multiple linear regression linked environmental variables to abundance and composition. Results were reported with R-squared values, exact or threshold p-values, and medians with interquartile ranges; counts were normalized by scaling to the smallest library size (McMurdie and Holmes approach).

Replicationbiological Sample size228 sequencing data sets generated after combining biological replicates; 2 samples with <2,000 reads removed; per-system n reported in PERMANOVA table (35, 55, 126); no formal power analysis described Groupslake types (Inland, Laurentian, Estuary), depth, season, fraction (FL vs PA), source (water vs sediment) Pairingunpaired Randomization/blindingna DispersionIQR Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Kruskal-Wallis test Verrucomicrobia relative abundance and diversity (inverse Simpson, SES MPD) across lake type, depth, season, fraction, and source (e.g. Fig 2, S1, S4) not stated
Multiple linear regression Verrucomicrobia relative abundance modeled against quantitative environmental parameters per lake type (Table 1) stated
Nested PERMANOVA Verrucomicrobia and non-Verrucomicrobia community composition vs source, fraction, season, station, depth (Table 2) Laurentian n=35, Estuary n=55, Inland n=126 not stated
Procrustes test correlation between Verrucomicrobia and non-Verrucomicrobia community ordinations per lake type na
bioenv analysis identifying physicochemical parameters correlating with community composition per lake type (Table 3) na
PCoA on Bray-Curtis dissimilarity visualizing compositional differences (Fig 3) na
Approaches that could also have been used
  • Group differences in relative abundance and diversity were assessed with Kruskal-Wallis tests for each panel/category separately.
    Could also: A single omnibus model (e.g. one Kruskal-Wallis or ANOVA per factor) followed by a defined post-hoc multiple-comparison correction such as Benjamini-Hochberg FDR or Dunn's test could also be applied across the family of comparisons. — An explicit correction across the many comparisons would control the family-wise or false-discovery rate, which can be helpful when reporting numerous simultaneous tests.
  • Counts were normalized by scaling to the smallest library size (2,072 sequences) and rounding, following McMurdie and Holmes.
    Could also: Variance-stabilizing or model-based normalization approaches (e.g. DESeq2/edgeR transformations, CSS in metagenomeSeq, or rarefaction-curve-based comparisons) could also be used. — Model-based normalization retains more reads and can improve power and variance handling, offering an alternative way to account for differing library sizes.
  • Verrucomicrobia relative abundance (a proportion) was analyzed with linear regression and Kruskal-Wallis tests.
    Could also: Compositional or proportion-aware models (e.g. beta regression, or analyses on centered-log-ratio transformed data) could also be used. — Such approaches are designed for bounded proportion or compositional data and would model the constrained range and variance structure explicitly.
  • Dispersion of relative abundance was summarized as median ± interquartile range.
    Could also: A bootstrap 95% confidence interval for the median or location estimate could also accompany the IQR. — A confidence interval conveys uncertainty about the estimated central tendency in addition to describing the spread of the data.
  • Best regression parameters were selected by testing variables that exhibited linearity and were not co-correlated.
    Could also: Penalized or information-criterion-based model selection (e.g. LASSO, or AIC/BIC comparison with cross-validation) could also guide variable selection. — These approaches provide a formalized selection procedure and can help characterize out-of-sample predictive performance alongside the reported R-squared.
  • PERMANOVA was used to partition community-composition variance among categorical factors.
    Could also: A complementary test of multivariate dispersion (e.g. PERMDISP/betadisper) could also be reported alongside PERMANOVA. — Checking group dispersion helps distinguish differences in community centroid location from differences in within-group variability, adding interpretive context to PERMANOVA results.
Software: mothur (mothu output files referenced) · R (analysis code provided on GitHub; phyloseq-style normalization per McMurdie and Holmes implied)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
79
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

PRJNA412983 BioProject in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PRJNA414423 BioProject in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-29590198

Paper: Chiang et al. 2018, PLoS ONE 13(3):e0195112. "Verrucomicrobia are prevalent in north-temperate freshwater lakes and display class-level preferences between lake habitats." PMCID PMC5874073.

Code: https://github.com/DenefLab/Verruco @ 8455714 (only commit, 2018-05-12). Data: SRA PRJNA412983 (16S rRNA V4 amplicons, MiSeq).

What the pipeline is

Two computational stages:

  1. mothur 16S pipeline (raw fastq -> OTU table). Output ships in the repo (mothur_output/*.shared, *.cons.taxonomy) and is baked into analysis/VerrucoData.RData (a phyloseq object verruco).
  2. R analysis (analysis/verruco-analysis.Rmd, 5249 lines) consuming the shipped phyloseq object -> all manuscript figures, tables, and the reported statistics (relative abundance, Kruskal-Wallis, PERMANOVA/adonis, Procrustes, bioenv, DESeq2, phylogenetic-signal).

In scope (attempted) — the R analysis on shipped data

The Rmd is deterministic (set.seed(1)) and consumes the shipped VerrucoData.RData. Crucially, the Rmd embeds the authors' own expected outputs as inline comments (e.g. # p-value = 1.728e-10) directly after each test, so we can do a precise 1:1 comparison against both the paper text/tables AND the authors' documented runs.

Targeted, clearly-specified headline results (minimal deps: phyloseq, vegan, plyr, dplyr, scales):

  • C1 n samples retained after rarefying/scaling (paper: 226).
  • C2 n Verrucomicrobia OTUs in the scaled object (Rmd inline: "taxa = 329").
  • C3 Verruco relative abundance: overall median + by lake type (paper: median 9.3%, range 1.7-41.7%; Laurentian/Estuary/Inland medians).
  • C4 Kruskal-Wallis, Verruco rel-abund ~ Lake_Type (Rmd: 1.728e-10; paper p<0.001).
  • C5-C7 KW pairwise lake types: Laur-Est 1.152e-08, Laur-Inl 1.206e-10, Est-Inl 0.4838.
  • C8 KW ~ Fraction (8.544e-05), C9 KW ~ Source (0.0001279), C10 KW ~ Season (0.6316).
  • C11 PERMANOVA (adonis, Bray-Curtis) Verruco community ~ Lake_Type (Rmd: R2=0.10811, p=0.001; paper R2=0.11, p=0.001).

All eleven are reproduced by re-running the authors' exact code on the shipped VerrucoData.RData. This is the 80%: deterministic, shipped-input, base-R + two well-known ecology packages.

Out of scope (not attempted — the hard ~20%) and why

  • mothur raw-reads -> OTU table (stage 1). Needs mothur + SILVA v119 refs + exact alignment/cluster params; the OTU table already ships, so re-deriving it is the expensive, low-marginal-value 20%. We trust the shipped OTU table (the same one the authors analysed) and reproduce downstream of it.
  • DESeq2 differential-abundance, phylogenetic-signal (geiger/picante), bioenv (Table 3), multiple-linear-regression model selection (Table 1), Procrustes, IToL tree — each adds many heavy/finicky deps (phytools, geiger, picante, labdsv, leaps, DESeq2, ade4) and/or manual steps (IToL is a web GUI). Skipped to honour the 80/20 rule; the headline community-ecology claims above are the clearest and most central to the paper's thesis.
  • Wet-lab / sampling / chemistry measurements — not computational.

No completeness claim

We reproduce 11 specific computational values, not the whole paper.

C1
Reported
226
Reproduced
226
exact
C2
Reported
329
Reproduced
329
exact
C3a
Reported
9.3%
Reproduced
9.34%
exact
C3b
Reported
1.7-41.7%
Reproduced
1.72-41.65%
exact
C3c
Reported
4.7%
Reproduced
4.68%
exact
C3d
Reported
11.0%
Reproduced
11.01%
exact
C3e
Reported
10.3%
Reproduced
10.26%
exact
C4
Reported
1.728e-10 (paper p<0.001)
Reproduced
1.728e-10
exact
C5
Reported
1.152e-08
Reproduced
1.1515e-08
exact
C6
Reported
1.206e-10
Reproduced
1.2059e-10
exact
C7
Reported
0.4838
Reproduced
0.48381
exact
C8
Reported
8.544e-05
Reproduced
8.544e-05
exact
C9
Reported
0.0001279
Reproduced
0.0001279
exact
C10
Reported
0.6316
Reproduced
0.6316
exact
C11
Reported
R2=0.11, p=0.001 (Rmd adonis R2=0.10811)
Reproduced
R2=0.112, F=13.43, p=0.001
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 99/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

This is a clean 1:1 computational reproduction: the authors shipped both the phyloseq object (VerrucoData.RData) and the analysis Rmd, and re-running the verbatim code reproduced 14/15 values exactly, including every inline p-value to all printed digits — no fabrication signal. The lone deviation (PERMANOVA R2 0.112 vs 0.10811) is an expected library-version artifact (vegan's adonis() replaced by adonis2()), with the p-value exact and both R2 values rounding to the paper's 0.11. The central conclusion — Verrucomicrobia prevalence and habitat-level differences — holds fully under reproduction; all deviations are technical/rounding on our side, none on the authors'.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

124.8 k
tokens (I/O) · 8.8 M incl. cache
15 min
runtime · 0 CPU-h
1.5 GB
peak RAM
1
HPC jobs
hummel
machine