The effect of 16S rRNA region choice on bacterial community metabarcoding results.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🔴The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for the region-intrinsic amplicon-length checks; NOT for the headline pooled OTU/diversity numbers. Code is the authors' own helper repo (yuragal/mothur-scripts @88236b6 — demux/quality Perl + shared-OTU R); the core pipeline is mothur v1.34.4 MiSeq SOP described in the paper. Data is public on SRA (SRP145556 sediment + SRP102494 water). REPRODUCED on «our HPC» with mothur 1.48.5 make.contigs on the deposited explicitly region-labeled runs: (1) sediment read counts EXACT (24410/6305); (2) V3-V4 amplicon length ~472 bp -> ~443 after primer trim = matches paper's 443 (within-tol); (3) V2-V3 amplicon length ~410 bp across all 14 *_V23 runs at 100% merge (reads are 2x300 bp, ~190 bp overlap -> genuine 410 bp amplicon), which is EXACTLY 2x the paper's reported 205 bp -> graded mismatch and FLAGGED as a candidate discrepancy for human audit (paper's V2-V3 length is half of what its own deposited data produce; benign explanations like single-read-length-mislabel or sub-region trim are possible; NOT asserted as fabrication). NOT attempted (the hard ~20%): the pooled filtered-read totals (118232/22191), species-level OTU counts (2716/1615), multi-read/singleton fractions, top-95% pools, and all diversity indices (Shannon/Simpson/Chao1/ACE/p-distance) -- these require the full multi-sample pooled SILVA-123 clustering AND an undeposited run->14-sample sheet (13 large *-Lib1 runs carry no region label), so they are not 1:1 derivable from the public record. checkfastq.pl was invoked but its 2015 report-column format does not match mothur 1.48.5 output. Fabrication assessment: read counts and V3-V4 length show no concern; the V2-V3 205 vs 410 factor-2 gap is the one item warranting human review.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 59assessed: 2026-06-16 ⛓ dc77b4a1f012
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study compares the resolution of V2-V3 versus V3-V4 16S rRNA gene fragments for estimating bacterial community diversity via Illumina MiSeq metabarcoding, assuming the region that produces the highest diversity will be most useful, and aiming to determine the read coverage needed for statistically sound diversity estimates.
- ★ The V2-V3 16S rRNA fragment has higher resolution for lower-rank taxa (genera and species) than V3-V4, allowing more precise distance-based clustering of reads into species-level OTUs. finding
- ★ Convergent estimates of major-species diversity (OTUs covering 95% of reads) are achieved at sample sizes of 10000-15000 reads, with Shannon index relative error below 4%. finding
- ★ Shannon and Simpson species-diversity estimates do not differ significantly between V2-V3 and V3-V4 regions. finding
- ★ Hidden species richness (Chao1, ACE) is significantly higher for the V2-V3 region than V3-V4, indicating greater resolution at the species-level clustering threshold (0.03). finding
- ★ The fragment including V2 and V3 regions accumulates mutations faster than V3-V4 during early stages of bacterial speciation (higher average genetic distance). mechanism
- Differences in taxonomic composition between regions increase at lower taxonomic ranks (from class to family) and are minor at the phylum level. finding
- A quality-filtering and analysis pipeline using Mothur, SILVA 123, bootstrap convergence and R-based diversity statistics, with deposited sequence datasets as a resource. method
- 16S rRNA amplicon datasets for V2-V3 and V3-V4 fragments from Lake Baikal water and sediment deposited in NCBI SRA. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| 16S rRNA amplicon metabarcoding (V2-V3 fragment, B_V23 primers) | Lake Baikal water column and bottom sediment bacterial communities | none | species-level OTU abundance / community diversity | Illumina MiSeq Standard Kit v.3; Phusion Hot Start II polymerase; NEBNext Ultra II DNA Library Prep Kit |
| 16S rRNA amplicon metabarcoding (V3-V4 fragment, Pro_V34 primers) | Lake Baikal water column and bottom sediment bacterial communities | none | species-level OTU abundance / community diversity | Illumina MiSeq Standard Kit v.3; Phusion Hot Start II polymerase; NEBNext Ultra II DNA Library Prep Kit |
| DNA extraction (phenol-chloroform, lysozyme treatment) | Lake Baikal water filtrate (nitrocellulose filters) and homogenized bottom sediment | none | DNA concentration and quality | SmartSpec Plus spectrophotometer (BioRad) |
| PCR product concentration control (capillary electrophoresis) | PCR amplicons from extracted DNA | none | PCR product concentration | Shimadzu Multi-NA (DNA-12000 reagent kit) |
| Bioinformatic OTU clustering and taxonomic assignment | MiSeq paired-end reads | none | OTUs at 0.03 genetic distance, taxonomy | Mothur v.1.34.4, SILVA 123 database |
| Diversity/statistical analysis (Shannon, Simpson, Chao1, ACE, NMDS, p-distance) | species-rank OTU abundance tables | none | diversity indices, relative errors, genetic distances, community composition | R packages ape, pegas, gplots, vegan |
- – V2-V3 dataset: 118232 reads, average length 205 bp, clustered into 2716 species-level OTUs (1070 with >1 read; singletons 1.4% of reads) 118232 reads; 2716 OTUs
- – V3-V4 dataset: 22191 reads, average length 443 bp, clustered into 1615 species-level OTUs (509 with >1 read; singletons 5.2% of reads) 22191 reads; 1615 OTUs
- ▲ Chao1 and ACE hidden species richness higher for V2-V3 than V3-V4 (Chao1 mean 124 vs 91; ACE mean 123 vs 89), significant by Wilkinson-Mann-Whitney Chao1 124 vs 91; ACE 123 vs 89
- ▼ Average genetic distance (nucleotide diversity) lower for V2-V3 than V3-V4 region, significantly different 0.204 vs 0.228
- ▼ Shannon index relative error decreases with read count; <8% at 5000 reads, <4% at 10000 reads, ~3% at 10000-15000 reads <4% at 10000 reads
- – No significant correlation between Shannon index and read count, indicating sufficient coverage of 95% major OTUs r=0.098, P=0.61
- – Shannon indices did not differ significantly between regions (mean 2.79 V2-V3 vs 2.72 V3-V4); Simpson means 0.79 vs 0.82, also not significant Shannon 2.79 vs 2.72; Simpson 0.79 vs 0.82
- – Taxonomic composition differences small at phylum level and increase at lower ranks (Bray-Curtis R2=0.12 p=0.04; Gower R2=0.14 p=0.04 at phylum) R2=0.12-0.14 at phylum
- correlation r = 0.098, P-value = 0.61 (Shannon index vs read counts (no correlation))
- correlation r = −0.697, P-value = 0.00 (Shannon index relative error vs read count (significant negative))
- correlation r = 0.015, p = 0.93 (Simpson index vs read counts (no correlation))
- correlation r = −0.59, p ~ 10^-5 (Simpson index relative error vs read count)
- mean Shannon 2.79 (V2-V3) vs 2.72 (V3-V4) (average Shannon biodiversity index across 14 samples)
- mean genetic distance 0.204 (V2-V3) vs 0.228 (V3-V4) (average pairwise p-distance of representative sequences, P<0.05)
- other R2 = 0.12 (Bray-Curtis), p = 0.04; R2 = 0.14 (Gower), p = 0.04 (covariance of taxonomic composition by region at phylum level)
- count 251 OTUs (V2-V3) and 171 OTUs (V3-V4) cover 95% of reads (top OTUs retained after convergence filtering (underestimation <16%))
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This Data Descriptor compares bacterial community diversity estimates derived from two 16S rRNA gene fragments (V2-V3 vs V3-V4) sequenced on Illumina MiSeq, using paired samples from Lake Baikal. Reads were clustered into species-level OTUs in Mothur, and diversity was summarized with Shannon, Simpson, Chao1, and ACE indices whose confidence intervals/relative errors were obtained by bootstrap. Differences between regions were tested with a paired Wilcoxon-Mann-Whitney test, associations were assessed with Spearman correlation, and community composition was compared with NMDS plus permutation-based R2 (PERMANOVA-style) on Bray-Curtis, Gower, and Jaccard distances. Results were reported with index means, ranges, p-values, relative errors, and 95% confidence intervals.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Paired Wilcoxon-Mann-Whitney nonparametric test | differences between average Shannon, Simpson, Chao1 and ACE indices for V2-V3 vs V3-V4 fragments (Table 3); also average p-distances/genetic distances (Fig. 4) | 14 samples (Shannon/Simpson, paired across regions); n for genetic-distance comparison not stated | not stated |
| Spearman rank correlation (with Spearman significance test) | correlations between read counts and index values / relative errors (e.g., Shannon r=0.098 P=0.61; Shannon relative error r=-0.697 P~0.00; Simpson r=0.015 p=0.93; Simpson relative error r=-0.59) | — | not stated |
| Permutation test (1000 permutations) on R2 covariation coefficient (PERMANOVA-style) | degree of difference in taxonomic composition by V2-V3 vs V3-V4 factor at phylum/class/order/family levels for Bray-Curtis, Gower, Jaccard distances (Table 4) | 1000 permutations | na |
| Bootstrap index / bootstrap resampling | α-diversity convergence (proportion of undetected species) and confidence intervals/relative errors of Shannon and Simpson indices | — | na |
| NMDS ordination (Bray-Curtis, Gower, Jaccard distances) | qualitative/quantitative comparison of community composition and similarity of OTU/species sets (Fig. 3) | — | na |
-
Region differences across four diversity indices were each tested with a paired Wilcoxon-Mann-Whitney test, with no stated multiple-comparison correction.↳ Could also: A family-wise or false-discovery-rate adjustment (e.g., Benjamini-Hochberg or Holm) could also be applied across the set of index comparisons. — Such an adjustment would control the overall error rate across the family of related tests and is often reported alongside multiple parallel comparisons.
-
Composition differences by region were assessed with a permutation test on the R2 covariation coefficient (a PERMANOVA-style approach).↳ Could also: The same permutation framework could also be reported as a formal adonis/PERMANOVA with its pseudo-F statistic, and accompanied by a multivariate dispersion check (e.g., betadisper/PERMDISP). — Reporting the test statistic and a dispersion check would help distinguish location (centroid) differences from differences in group spread under the same distance-based design.
-
Confidence intervals and relative errors for Shannon and Simpson indices were obtained by bootstrap resampling.↳ Could also: Analytic variance estimators or rarefaction/coverage-based standardization could also be used to summarize diversity uncertainty. — Coverage-based standardization is a widely used complement for comparing diversity across samples with differing read depths and would offer an alternative view of sampling completeness.
-
Associations between read counts and indices/relative errors were quantified with Spearman rank correlations.↳ Could also: A regression model (e.g., on log read count) could also be fitted to describe these relationships. — A model-based approach would additionally provide a slope/effect magnitude and predicted values, complementing the rank-based association measure.
-
Data were normalized by the average number of reads per sample before composition analysis.↳ Could also: Rarefaction to a common depth or model-based normalization (e.g., DESeq2/edgeR-style or CSS in metagenomeSeq) could also be used. — These alternatives are common in amplicon studies for handling uneven library sizes and would offer a different way to account for variable sequencing depth.
-
Index estimates were reported with means, ranges, SE (Chao1/ACE) and 95% CIs/relative errors (Shannon/Simpson), a mix of dispersion summaries.↳ Could also: Uniformly reporting 95% confidence intervals (or SD) for all indices could also be done. — A consistent dispersion measure across indices makes spread directly comparable across the reported metrics, which is often preferred for small sample sizes.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Chao1 and ACE estimated species richness higher for the V2-V3 region than the V3-V4 regionother bacteria lake-baikal water sediment up 2019×1papers★ This paper is the founder (earliest)
-
Average genetic distance (nucleotide diversity) lower for the V2-V3 region than the V3-V4 regionother bacteria lake-baikal water sediment down 2019×1papers★ This paper is the founder (earliest)
-
Shannon and Simpson diversity indices did not differ significantly between the V2-V3 and V3-V4 regionsother bacteria lake-baikal water sediment none 2019×1papers★ This paper is the founder (earliest)
-
No significant correlation between Shannon index and read count, indicating sufficient sequencing coverageother bacteria lake-baikal water sediment none 2019×1papers★ This paper is the founder (earliest)
-
Relative error of the Shannon index decreases as read count increasesother bacteria lake-baikal water sediment down 2019×1papers★ This paper is the founder (earliest)
-
Taxonomic composition differences between regions small at phylum level and increase at lower taxonomic ranksother bacteria lake-baikal water sediment mixed 2019×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-30720800
Paper: Bukin et al. 2019, Sci Data — "The effect of 16S rRNA region choice on
bacterial community metabarcoding results." (data descriptor)
Code: https://github.com/yuragal/mothur-scripts (commit 88236b6, 2015 — authors'
helper Perl/R scripts: demu.{1,2}.pl demultiplex by primer, checkfastq.pl quality
filter of contigs in non-overlapping regions, count_shared_OUTs_seqs.R).
Data: SRA SRP145556 (bottom sediment, 2 runs) + SRP102494 (water column, 41 runs).
Pipeline (from Methods / Technical Validation)
mothur v1.34.4, MiSeq SOP: (1) merge R1/R2 into contigs (make.contigs);
(2) quality-filter — discard contigs with >5 sites of Phred ≤15 in the
non-overlapping regions (authors' checkfastq.pl); (3) align + cluster at 0.03
genetic distance → species-level OTUs; (4) taxonomy vs SILVA 123.
Two amplicons: V2-V3 (B_V23, primers 16S_BV2f/16S_BV3r) and V3-V4
(Pro_V34, primers MiCSQ_343FL/MiCSQ_806R).
Reported values (Technical Validation, pooled/averaged across the dataset)
| metric | V2-V3 | V3-V4 |
|---|---|---|
| reads after filtering | 118232 | 22191 |
| average contig length (bp) | 205 | 443 |
| species-level OTUs (0.03) | 2716 | 1615 |
| multi-read OTUs | 1070 | 509 |
| singleton % of reads | 1.4% | 5.2% |
| OTUs covering top 95% | 251 | 171 |
| Shannon (mean of 14 samples) | 2.79 | 2.72 |
| Simpson | 0.79 | 0.82 |
| Chao1 | 124 | 91 |
| ACE | 123 | 89 |
| avg p-distance | 0.204 | 0.228 |
IN SCOPE (clean, low-hanging, 1:1) — what we reproduce
- C1 — average contig length per region (205 / 443 bp). Region-intrinsic
(set by primer pair + amplicon), so robustly reproducible. Method: download the
explicitly region-labeled runs (
*_V23,*_V34, plus sediment15_B_V23/15_Pro_V34), merge R1/R2 with mothurmake.contigs, pool per region, measure mean contig length. This is the cleanest 1:1 target. - C2 — data integrity / read counts for the designated SRP145556 (sediment): 24410 pairs (15_B_V23), 6305 pairs (15_Pro_V34) per ENA — verify deposit present and counts as a data-availability check.
- C3 — quality-filter behavior using the authors' own
checkfastq.pl(≤5 sites Phred≤15 in non-overlap): demonstrate it runs on the deposited reads and report retained fraction (paper gives no per-sample filtered count → graded partial, mechanism-level).
OUT OF SCOPE (the hard ~20%, not attempted — why)
- Pooled filtered read totals (118232 / 22191) and OTU counts (2716 / 1615),
multi-read OTUs, singleton %, top-95% pool, all diversity indices. These are
computed over the full 14-sample dataset pooled / averaged, but the mapping
of the 43 SRA runs → the paper's 14 samples × 2 regions is not documented:
SRP102494 mixes explicitly region-labeled R6/R9 replicate runs (
*_V23/*_V34) with 13 large unlabeled*-Lib1runs (0.35–0.86 M reads each) whose region assignment and inclusion are unstated. Exact pooled counts therefore depend on an undeposited author sample-sheet, and clustering needs the SILVA 123 reference + the full multi-sample alignment. Not derivable 1:1 from the public record → recorded honestly, not as mismatch. (Per-sample sediment OTUs would not 1:1 match pooled 2716/1615 anyway.)
Grading note
Average contig length is the one fully-specified, sample-independent number the public data alone can confirm; it is our primary 1:1 claim.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This Sci Data descriptor reproduces partly: sediment read counts are exact (24410/6305) and the V3-V4 length matches 443 bp after primer trimming, but the reported V2-V3 length of 205 bp is not derivable from the deposited reads — every one of the 14 *_V23 runs yields ~410 bp, exactly 2x the reported value (flagged for human audit; benign explanations like single-read-length mislabel or sub-region trim are plausible, fabrication NOT asserted). The headline conclusion (region choice alters OTU/diversity results) rests on pooled OTU counts and diversity indices that could not be reproduced because the run->14-sample mapping is undeposited (13 *-Lib1 runs unlabeled). The most severe deviation (factor-2 length) sits on the authors' side as a value not derivable from their own shared data, but the overall reproduction is honest and solid where the public record allowed.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.