Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

The effect of 16S rRNA region choice on bacterial community metabarcoding results.

Sci Data · 2019
L1 59/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🔴The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
59/100
Reproducibility score
0.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 19% of all assessed papers rank 925 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for the region-intrinsic amplicon-length checks; NOT for the headline pooled OTU/diversity numbers. Code is the authors' own helper repo (yuragal/mothur-scripts @88236b6 — demux/quality Perl + shared-OTU R); the core pipeline is mothur v1.34.4 MiSeq SOP described in the paper. Data is public on SRA (SRP145556 sediment + SRP102494 water). REPRODUCED on «our HPC» with mothur 1.48.5 make.contigs on the deposited explicitly region-labeled runs: (1) sediment read counts EXACT (24410/6305); (2) V3-V4 amplicon length ~472 bp -> ~443 after primer trim = matches paper's 443 (within-tol); (3) V2-V3 amplicon length ~410 bp across all 14 *_V23 runs at 100% merge (reads are 2x300 bp, ~190 bp overlap -> genuine 410 bp amplicon), which is EXACTLY 2x the paper's reported 205 bp -> graded mismatch and FLAGGED as a candidate discrepancy for human audit (paper's V2-V3 length is half of what its own deposited data produce; benign explanations like single-read-length-mislabel or sub-region trim are possible; NOT asserted as fabrication). NOT attempted (the hard ~20%): the pooled filtered-read totals (118232/22191), species-level OTU counts (2716/1615), multi-read/singleton fractions, top-95% pools, and all diversity indices (Shannon/Simpson/Chao1/ACE/p-distance) -- these require the full multi-sample pooled SILVA-123 clustering AND an undeposited run->14-sample sheet (13 large *-Lib1 runs carry no region label), so they are not 1:1 derivable from the public record. checkfastq.pl was invoked but its 2015 report-column format does not match mothur 1.48.5 output. Fabrication assessment: read counts and V3-V4 length show no concern; the V2-V3 205 vs 410 factor-2 gap is the one item warranting human review.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 59
    assessed: 2026-06-16 ⛓ dc77b4a1f012
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study compares the resolution of V2-V3 versus V3-V4 16S rRNA gene fragments for estimating bacterial community diversity via Illumina MiSeq metabarcoding, assuming the region that produces the highest diversity will be most useful, and aiming to determine the read coverage needed for statistically sound diversity estimates.

Core claims
  • The V2-V3 16S rRNA fragment has higher resolution for lower-rank taxa (genera and species) than V3-V4, allowing more precise distance-based clustering of reads into species-level OTUs. finding
  • Convergent estimates of major-species diversity (OTUs covering 95% of reads) are achieved at sample sizes of 10000-15000 reads, with Shannon index relative error below 4%. finding
  • Shannon and Simpson species-diversity estimates do not differ significantly between V2-V3 and V3-V4 regions. finding
  • Hidden species richness (Chao1, ACE) is significantly higher for the V2-V3 region than V3-V4, indicating greater resolution at the species-level clustering threshold (0.03). finding
  • The fragment including V2 and V3 regions accumulates mutations faster than V3-V4 during early stages of bacterial speciation (higher average genetic distance). mechanism
  • Differences in taxonomic composition between regions increase at lower taxonomic ranks (from class to family) and are minor at the phylum level. finding
  • A quality-filtering and analysis pipeline using Mothur, SILVA 123, bootstrap convergence and R-based diversity statistics, with deposited sequence datasets as a resource. method
  • 16S rRNA amplicon datasets for V2-V3 and V3-V4 fragments from Lake Baikal water and sediment deposited in NCBI SRA. resource
Experimental setups
Assay System Perturbation Readout Platform
16S rRNA amplicon metabarcoding (V2-V3 fragment, B_V23 primers) Lake Baikal water column and bottom sediment bacterial communities none species-level OTU abundance / community diversity Illumina MiSeq Standard Kit v.3; Phusion Hot Start II polymerase; NEBNext Ultra II DNA Library Prep Kit
16S rRNA amplicon metabarcoding (V3-V4 fragment, Pro_V34 primers) Lake Baikal water column and bottom sediment bacterial communities none species-level OTU abundance / community diversity Illumina MiSeq Standard Kit v.3; Phusion Hot Start II polymerase; NEBNext Ultra II DNA Library Prep Kit
DNA extraction (phenol-chloroform, lysozyme treatment) Lake Baikal water filtrate (nitrocellulose filters) and homogenized bottom sediment none DNA concentration and quality SmartSpec Plus spectrophotometer (BioRad)
PCR product concentration control (capillary electrophoresis) PCR amplicons from extracted DNA none PCR product concentration Shimadzu Multi-NA (DNA-12000 reagent kit)
Bioinformatic OTU clustering and taxonomic assignment MiSeq paired-end reads none OTUs at 0.03 genetic distance, taxonomy Mothur v.1.34.4, SILVA 123 database
Diversity/statistical analysis (Shannon, Simpson, Chao1, ACE, NMDS, p-distance) species-rank OTU abundance tables none diversity indices, relative errors, genetic distances, community composition R packages ape, pegas, gplots, vegan
Key results
  • V2-V3 dataset: 118232 reads, average length 205 bp, clustered into 2716 species-level OTUs (1070 with >1 read; singletons 1.4% of reads) 118232 reads; 2716 OTUs
  • V3-V4 dataset: 22191 reads, average length 443 bp, clustered into 1615 species-level OTUs (509 with >1 read; singletons 5.2% of reads) 22191 reads; 1615 OTUs
  • Chao1 and ACE hidden species richness higher for V2-V3 than V3-V4 (Chao1 mean 124 vs 91; ACE mean 123 vs 89), significant by Wilkinson-Mann-Whitney Chao1 124 vs 91; ACE 123 vs 89
  • Average genetic distance (nucleotide diversity) lower for V2-V3 than V3-V4 region, significantly different 0.204 vs 0.228
  • Shannon index relative error decreases with read count; <8% at 5000 reads, <4% at 10000 reads, ~3% at 10000-15000 reads <4% at 10000 reads
  • No significant correlation between Shannon index and read count, indicating sufficient coverage of 95% major OTUs r=0.098, P=0.61
  • Shannon indices did not differ significantly between regions (mean 2.79 V2-V3 vs 2.72 V3-V4); Simpson means 0.79 vs 0.82, also not significant Shannon 2.79 vs 2.72; Simpson 0.79 vs 0.82
  • Taxonomic composition differences small at phylum level and increase at lower ranks (Bray-Curtis R2=0.12 p=0.04; Gower R2=0.14 p=0.04 at phylum) R2=0.12-0.14 at phylum
Key statistics
  • correlation r = 0.098, P-value = 0.61 (Shannon index vs read counts (no correlation))
  • correlation r = −0.697, P-value = 0.00 (Shannon index relative error vs read count (significant negative))
  • correlation r = 0.015, p = 0.93 (Simpson index vs read counts (no correlation))
  • correlation r = −0.59, p ~ 10^-5 (Simpson index relative error vs read count)
  • mean Shannon 2.79 (V2-V3) vs 2.72 (V3-V4) (average Shannon biodiversity index across 14 samples)
  • mean genetic distance 0.204 (V2-V3) vs 0.228 (V3-V4) (average pairwise p-distance of representative sequences, P<0.05)
  • other R2 = 0.12 (Bray-Curtis), p = 0.04; R2 = 0.14 (Gower), p = 0.04 (covariance of taxonomic composition by region at phylum level)
  • count 251 OTUs (V2-V3) and 171 OTUs (V3-V4) cover 95% of reads (top OTUs retained after convergence filtering (underestimation <16%))

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This Data Descriptor compares bacterial community diversity estimates derived from two 16S rRNA gene fragments (V2-V3 vs V3-V4) sequenced on Illumina MiSeq, using paired samples from Lake Baikal. Reads were clustered into species-level OTUs in Mothur, and diversity was summarized with Shannon, Simpson, Chao1, and ACE indices whose confidence intervals/relative errors were obtained by bootstrap. Differences between regions were tested with a paired Wilcoxon-Mann-Whitney test, associations were assessed with Spearman correlation, and community composition was compared with NMDS plus permutation-based R2 (PERMANOVA-style) on Bray-Curtis, Gower, and Jaccard distances. Results were reported with index means, ranges, p-values, relative errors, and 95% confidence intervals.

Replicationmixed Sample size14 samples total (8 from zone I, 5 from zone II water column, plus bottom sediment); four independent DNA extractions per sample; final analysis on top OTUs covering 95% of reads (251 for V2-V3, 171 for V3-V4); no formal power analysis described GroupsV2-V3 vs V3-V4 16S rRNA fragments (paired within sample) Pairingpaired Randomization/blindingna Dispersionmixed Exact p-valuesyes Effect sizesyes Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Paired Wilcoxon-Mann-Whitney nonparametric test differences between average Shannon, Simpson, Chao1 and ACE indices for V2-V3 vs V3-V4 fragments (Table 3); also average p-distances/genetic distances (Fig. 4) 14 samples (Shannon/Simpson, paired across regions); n for genetic-distance comparison not stated not stated
Spearman rank correlation (with Spearman significance test) correlations between read counts and index values / relative errors (e.g., Shannon r=0.098 P=0.61; Shannon relative error r=-0.697 P~0.00; Simpson r=0.015 p=0.93; Simpson relative error r=-0.59) not stated
Permutation test (1000 permutations) on R2 covariation coefficient (PERMANOVA-style) degree of difference in taxonomic composition by V2-V3 vs V3-V4 factor at phylum/class/order/family levels for Bray-Curtis, Gower, Jaccard distances (Table 4) 1000 permutations na
Bootstrap index / bootstrap resampling α-diversity convergence (proportion of undetected species) and confidence intervals/relative errors of Shannon and Simpson indices na
NMDS ordination (Bray-Curtis, Gower, Jaccard distances) qualitative/quantitative comparison of community composition and similarity of OTU/species sets (Fig. 3) na
Approaches that could also have been used
  • Region differences across four diversity indices were each tested with a paired Wilcoxon-Mann-Whitney test, with no stated multiple-comparison correction.
    Could also: A family-wise or false-discovery-rate adjustment (e.g., Benjamini-Hochberg or Holm) could also be applied across the set of index comparisons. — Such an adjustment would control the overall error rate across the family of related tests and is often reported alongside multiple parallel comparisons.
  • Composition differences by region were assessed with a permutation test on the R2 covariation coefficient (a PERMANOVA-style approach).
    Could also: The same permutation framework could also be reported as a formal adonis/PERMANOVA with its pseudo-F statistic, and accompanied by a multivariate dispersion check (e.g., betadisper/PERMDISP). — Reporting the test statistic and a dispersion check would help distinguish location (centroid) differences from differences in group spread under the same distance-based design.
  • Confidence intervals and relative errors for Shannon and Simpson indices were obtained by bootstrap resampling.
    Could also: Analytic variance estimators or rarefaction/coverage-based standardization could also be used to summarize diversity uncertainty. — Coverage-based standardization is a widely used complement for comparing diversity across samples with differing read depths and would offer an alternative view of sampling completeness.
  • Associations between read counts and indices/relative errors were quantified with Spearman rank correlations.
    Could also: A regression model (e.g., on log read count) could also be fitted to describe these relationships. — A model-based approach would additionally provide a slope/effect magnitude and predicted values, complementing the rank-based association measure.
  • Data were normalized by the average number of reads per sample before composition analysis.
    Could also: Rarefaction to a common depth or model-based normalization (e.g., DESeq2/edgeR-style or CSS in metagenomeSeq) could also be used. — These alternatives are common in amplicon studies for handling uneven library sizes and would offer a different way to account for variable sequencing depth.
  • Index estimates were reported with means, ranges, SE (Chao1/ACE) and 95% CIs/relative errors (Shannon/Simpson), a mix of dispersion summaries.
    Could also: Uniformly reporting 95% confidence intervals (or SD) for all indices could also be done. — A consistent dispersion measure across indices makes spread directly comparable across the reported metrics, which is often preferred for small sample sizes.
Software: Mothur 1.34.4 · R package ape · R package pegas · R package gplots · R package vegan · SILVA database 123

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
395
Impact: very high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-30720800

Paper: Bukin et al. 2019, Sci Data — "The effect of 16S rRNA region choice on bacterial community metabarcoding results." (data descriptor) Code: https://github.com/yuragal/mothur-scripts (commit 88236b6, 2015 — authors' helper Perl/R scripts: demu.{1,2}.pl demultiplex by primer, checkfastq.pl quality filter of contigs in non-overlapping regions, count_shared_OUTs_seqs.R). Data: SRA SRP145556 (bottom sediment, 2 runs) + SRP102494 (water column, 41 runs).

Pipeline (from Methods / Technical Validation)

mothur v1.34.4, MiSeq SOP: (1) merge R1/R2 into contigs (make.contigs); (2) quality-filter — discard contigs with >5 sites of Phred ≤15 in the non-overlapping regions (authors' checkfastq.pl); (3) align + cluster at 0.03 genetic distance → species-level OTUs; (4) taxonomy vs SILVA 123.

Two amplicons: V2-V3 (B_V23, primers 16S_BV2f/16S_BV3r) and V3-V4 (Pro_V34, primers MiCSQ_343FL/MiCSQ_806R).

Reported values (Technical Validation, pooled/averaged across the dataset)

metric V2-V3 V3-V4
reads after filtering 118232 22191
average contig length (bp) 205 443
species-level OTUs (0.03) 2716 1615
multi-read OTUs 1070 509
singleton % of reads 1.4% 5.2%
OTUs covering top 95% 251 171
Shannon (mean of 14 samples) 2.79 2.72
Simpson 0.79 0.82
Chao1 124 91
ACE 123 89
avg p-distance 0.204 0.228

IN SCOPE (clean, low-hanging, 1:1) — what we reproduce

  • C1 — average contig length per region (205 / 443 bp). Region-intrinsic (set by primer pair + amplicon), so robustly reproducible. Method: download the explicitly region-labeled runs (*_V23, *_V34, plus sediment 15_B_V23 / 15_Pro_V34), merge R1/R2 with mothur make.contigs, pool per region, measure mean contig length. This is the cleanest 1:1 target.
  • C2 — data integrity / read counts for the designated SRP145556 (sediment): 24410 pairs (15_B_V23), 6305 pairs (15_Pro_V34) per ENA — verify deposit present and counts as a data-availability check.
  • C3 — quality-filter behavior using the authors' own checkfastq.pl (≤5 sites Phred≤15 in non-overlap): demonstrate it runs on the deposited reads and report retained fraction (paper gives no per-sample filtered count → graded partial, mechanism-level).

OUT OF SCOPE (the hard ~20%, not attempted — why)

  • Pooled filtered read totals (118232 / 22191) and OTU counts (2716 / 1615), multi-read OTUs, singleton %, top-95% pool, all diversity indices. These are computed over the full 14-sample dataset pooled / averaged, but the mapping of the 43 SRA runs → the paper's 14 samples × 2 regions is not documented: SRP102494 mixes explicitly region-labeled R6/R9 replicate runs (*_V23/*_V34) with 13 large unlabeled *-Lib1 runs (0.35–0.86 M reads each) whose region assignment and inclusion are unstated. Exact pooled counts therefore depend on an undeposited author sample-sheet, and clustering needs the SILVA 123 reference + the full multi-sample alignment. Not derivable 1:1 from the public record → recorded honestly, not as mismatch. (Per-sample sediment OTUs would not 1:1 match pooled 2716/1615 anyway.)

Grading note

Average contig length is the one fully-specified, sample-independent number the public data alone can confirm; it is our primary 1:1 claim.

C2_reads_sediment
Reported
SRR7160311=24410, SRR7160312=6305 read pairs (SRP145556)
Reproduced
24410, 6305
exact
C1_len_v34
Reported
443 bp average length, V3-V4 (Pro_V34)
Reproduced
~472 bp (well-merged contigs); ~443 bp after primer (343F/806R) removal
within tolerance
C1_len_v23
Reported
205 bp average length, V2-V3 (B_V23)
Reproduced
~410 bp (all 14 *_V23 runs, 100% merged) = EXACTLY 2x reported
did not match
C3_qfilter_checkfastq
Reported
authors' checkfastq.pl: discard contigs with >5 sites Phred<=15 in non-overlap
Reproduced
tool obtained & invoked; mothur-1.48 report-format mismatch blocked numeric completion
partial
OUT_pooled_otus_diversity
Reported
118232/22191 reads; 2716/1615 OTUs; Shannon 2.79/2.72; Simpson 0.79/0.82; Chao1 124/91; ACE 123/89; p-dist 0.204/0.228
Reproduced
not attempted (out of scope)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 59/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🔴6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

This Sci Data descriptor reproduces partly: sediment read counts are exact (24410/6305) and the V3-V4 length matches 443 bp after primer trimming, but the reported V2-V3 length of 205 bp is not derivable from the deposited reads — every one of the 14 *_V23 runs yields ~410 bp, exactly 2x the reported value (flagged for human audit; benign explanations like single-read-length mislabel or sub-region trim are plausible, fabrication NOT asserted). The headline conclusion (region choice alters OTU/diversity results) rests on pooled OTU counts and diversity indices that could not be reproduced because the run->14-sample mapping is undeposited (13 *-Lib1 runs unlabeled). The most severe deviation (factor-2 length) sits on the authors' side as a value not derivable from their own shared data, but the overall reproduction is honest and solid where the public record allowed.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

163.5 k
tokens (I/O) · 9.9 M incl. cache
23 min
runtime · 0.24 CPU-h
1.9 GB
peak RAM
2
HPC jobs
hummel
machine