High-resolution metagenomic reconstruction of the freshwater spring bloom.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL reproduction (P16 third-party tool on the authors' deposited data) — paper is described well enough and data are openly deposited; pipeline steps reproduce. DESIGN (D1, from open ENA metadata): 57 metagenomes EXACT, 39 epi/18 hypo split EXACT, ~830 Gb volume within-tol. PRIMARY P16 (R1): barrnap 0.9 (the designated github.com/tseemann/barrnap tool, default params) ran on ALL 1898 deposited MAGs -> rRNA table (159 w/16S, 30 w/full complement); process reproduced (paper gives no headline rRNA number). SECONDARY (R2): CheckM v1.0.18 (exact paper version) lineage_wf on all 1898 deposited MAGs (run in 10 batches to dodge a large-N multiprocessing deadlock) -> completeness DISTRIBUTION reproduces Suppl Fig S3A within ~1.5 percentage points on every tier (HQ 21.8% vs 22.6%, MQ 28.3% vs 26.8%, partial 49.9% vs 50.0%) when using the paper's >=40% denominator. TWO HONEST FLAGS for human review: (a) reported 'ca. 1.96 billion reads' vs ~6.42 billion deposited (read_count ~3.3x; 830 Gb volume matches); (b) deposited MAG set = 1898 GCA (1217 >=40% completeness, 681 below 40%) != reported 2214 recovered / 855 dereplicated -> the public deposit is NOT the QC-filtered set. NOT attempted: full 830 Gb MEGAHIT reassembly (non-deterministic, would not reproduce the 2214 MAG identities -> not a faithful 1:1 test), STRETCH R3 GTDB-Tk phylum split + R4 dRep (heavy ref data, descoped), wet-lab CARD-FISH/microscopy, NCLDV/phage/euk phylogenomics, growth-rate estimates.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-30no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether high-frequency, high-resolution metagenomic sampling of a freshwater reservoir spring bloom (integrated with CARD-FISH and microscopy) can recover genomes of prokaryotes, eukaryotes, and viruses to reveal novel taxa, distributional dynamics, and drivers of bloom collapse that lower-resolution or single-method approaches miss.
- ★ High-frequency metagenomic sampling (57 samples, 3 filter sizes, 2 depths, 37 days) recovers thousands of prokaryotic, eukaryotic, and viral genomes tracking spring bloom succession in fine detail method
- ★ MAG relative-abundance dynamics combined with CARD-FISH yield concordant in situ doubling time estimates for dominant genome-streamlined microbial lineages finding
- ★ Discordance between sequence-based and microscopy-based cryptophyte quantitation reveals hidden, highly abundant aplastidic cryptophytes, confirmed by CARD-FISH finding
- ★ Aplastidic cryptophytes are prevalent throughout the water column and have not previously been included in plankton dynamics models finding
- ★ First metagenome-assembled genomes of freshwater protists (a diatom and a haptophyte) and thousands of giant viral contigs were recovered resource
- ★ Contrasting distributions of giant viruses (whole water column) versus parasitic perkinsids (deeper waters) suggest giant viruses, together with nutrient limitation, drive top-down control and collapse of the epilimnetic bloom mechanism
- 2214 bacterial MAGs (≥40% completeness, ≤5% contamination) were recovered, dereplicating to 855 genomes, with no archaeal MAGs detected resource
- Rhodopsin genes (mostly proteorhodopsin-type) occur in ~38% of dereplicated bacterial MAGs, indicating widespread light-dependent energy acquisition in the community finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| 16S rRNA gene sequencing | freshwater reservoir microbial community (epilimnion/hypolimnion, 0.22/5/0.8 μm filters) | none (natural seasonal bloom) | relative abundance of prokaryotic taxonomic groups over time | — |
| 18S rRNA gene sequencing | freshwater reservoir eukaryotic plankton community | none (natural seasonal bloom) | relative abundance of eukaryotic taxonomic groups over time | — |
| Shotgun metagenomic sequencing and MAG binning | freshwater reservoir, 3 size fractions (5 μm, 0.8 μm, 0.22 μm), 2 depths, 57 samples | none (natural seasonal bloom) | genome recovery (completeness/contamination), taxonomic composition, gene content (e.g., rhodopsins) | — |
| CARD-FISH | freshwater reservoir plankton, specific bacterial and cryptophyte lineages | none (natural seasonal bloom) | in situ cell abundance and doubling time estimates | — |
| Microscopy | phytoplankton, ciliates, zooplankton (rotifers, crustaceans) in reservoir water column | none (natural seasonal bloom) | biovolume, cell counts, taxonomic identification | — |
| Grazing rate estimation | heterotrophic nanoflagellates (HNF) and ciliates | none (natural seasonal bloom) | total bacterivory rates | — |
| Physicochemical water analysis | reservoir water column (epilimnion/hypolimnion) | none (natural seasonal bloom) | chlorophyll-a, temperature, total phosphorus, DRP, nitrate, ammonium, DOC, silica concentrations | — |
| Rhodopsin sequence/motif analysis | bacterial MAGs recovered from metagenomes | none | rhodopsin type (type-1/heliorhodopsin), spectral tuning motif (DTE/DTG/ATI), absorption class (green- vs blue-absorbing) | — |
- – 2214 bacterial MAGs (≥40% completeness, ≤5% contamination) recovered, dereplicating to 855 genomes; no archaeal MAGs recovered 2214 MAGs → 855 dereplicated
- ▲ Bacteroidota rose to over half of 16S rRNA reads at bloom onset (days 9–11), with a second peak at day 21 >50% of 16S rRNA reads
- ▲ Total heterotrophic prokaryotic numbers increased during the phytoplankton bloom then stabilized 2 to 6 × 10^6 cells ml⁻¹
- – 615 MAGs encode rhodopsins across 326 clusters, representing ~38% of the dereplicated community 615/855 MAGs (~38%)
- – Perkinsozoa sequence abundance rose sharply in the hypolimnion while decreasing in the epilimnion up to ~40% relative abundance in hypolimnion
- ▲ Aplastidic cryptophytes detected as highly abundant via CARD-FISH despite being underrepresented/undetected by microscopy
- – Ciliate abundance peaked on day 14, preceding an HNF peak by 4 days 4-day lag
- – Giant viral genomic contigs (thousands) recovered and distributed across the entire water column, contrasting with perkinsid depth restriction thousands of contigs
- count 2214 bacterial MAGs (MAGs recovered from 57 metagenomic datasets (≥40% completeness, ≤5% contamination))
- count 855 dereplicated genomes (prokaryotic diversity representation after dereplication)
- fold_change 2 to 6 × 10^6 cells ml⁻¹ (increase in total heterotrophic prokaryotic numbers during bloom)
- other >50% (Bacteroidota share of 16S rRNA reads at bloom onset (days 9–11))
- other ~38% (326/855 MAGs) (proportion of dereplicated MAGs encoding rhodopsins)
- count 615 MAGs / 326 clusters (MAGs encoding rhodopsins; 508 type-1 (223 clusters), 214 heliorhodopsin (103 clusters), 203 encoding both (48 clusters))
- other ~40% (relative abundance of Perkinsozoa sequences in hypolimnion)
- count 1.96 billion reads, ca. 830 Gb (total sequencing output across 57 metagenomic samples)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper is a descriptive, exploratory metagenomic and microscopy-based field study of a freshwater spring bloom, reporting shifts in relative abundances of taxa (16S/18S rRNA reads), counts of metagenome-assembled genomes (MAGs), and physicochemical/microscopy time series across 57 samples spanning 37 days, two depths, and three filter sizes. The provided text does not describe formal inferential statistical hypothesis testing (e.g., group comparisons with p-values); results are presented as time-series trends, proportions, and genome/functional counts.
-
The study reports taxonomic and genomic shifts over time (e.g., relative abundance of Bacteroidota, Actinobacteriota) primarily through descriptive time-series plots and narrative comparison of peaks/troughs.↳ Could also: Formal time-series statistical models (e.g., generalized additive models, ARIMA, or change-point detection) or community-ecology tests (e.g., PERMANOVA on Bray-Curtis dissimilarities) could also be applied — Such approaches can quantify the significance and magnitude of community shifts over time or between depths/filter fractions, complementing the visual/descriptive trend interpretation.
-
Differences in MAG recovery across filter-pore sizes (1288 from 0.22 μm, 808 from 0.8 μm, 112 from 5 μm) are described narratively as a trend.↳ Could also: A statistical comparison (e.g., chi-square or rarefaction-based analysis accounting for sequencing depth) could also be used — This would allow formal assessment of whether recovery differences across filters exceed what might be expected from sequencing effort or depth alone.
-
In situ doubling time estimates and grazing rate estimates are mentioned as derived from combined metagenomic and CARD-FISH/microscopy data.↳ Could also: Reporting these estimates with confidence intervals or bootstrap-derived uncertainty ranges could also be used — This would convey the precision of doubling-time and grazing-rate estimates alongside the point values, which is common practice for rate estimates derived from field abundance data.
-
Co-occurrence patterns between taxa (e.g., giant viruses and perkinsids, ciliates and HNF peaks) are described qualitatively based on timing of peaks.↳ Could also: Formal correlation or cross-correlation analysis (e.g., Spearman correlation with time-lag analysis) could also be used — This would provide a quantitative measure of association strength and lag between the temporal dynamics of different taxonomic groups, in addition to visual inspection of co-occurring peaks.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36698172
Paper: Kavagutti et al. 2023, High-resolution metagenomic reconstruction of the freshwater spring bloom. Microbiome 11:11. DOI 10.1186/s40168-022-01451-4. PMID 36698172 · PMCID PMC9878933.
Designated code (P16, third-party tool): https://github.com/tseemann/barrnap — rRNA gene predictor. Paper Methods: "rRNA sequences were identified using barrnap with default parameters" (16S/23S), then compared to SILVA v138. This is a third-party tool applied to the paper's own data → reproducing it on the paper's deposited MAGs is an equally valid P16 reproduction.
Data: ENA PRJEB52406 (raw reads) + deposited MAG assemblies (GCA_* under the same project). Other deposits: NCLDV phylogenomic tree on figshare (figshare.com/s/5150884f3fbb534c0302).
Pipeline (from Methods)
QC: bbmap suite (reformat/bbduk/bbmerge, Phred 18) → assembly: MEGAHIT v1.1.5 (k=49,69,89,109,129,149) → mapping: bbwrap + jgi_summarize_bam_contig_depths → binning: MetaBAT2 (default) → ORFs: PRODIGAL v2.6.3 → gene taxonomy: MMseqs2 vs GTDB r95 → rRNA: barrnap (default) vs SILVA v138 → bin taxonomy: GTDB-Tk (r95) → QC: CheckM v1.0.18 → dereplication: dRep (-comp 40 -con 5). Viruses: VIBRANT, ViralRecall, ncldv_markersearch. Euk MAGs: BUSCO + metaeuk. Growth: RazerS3 + GRiD + gRodon. Phylogenomics: PASTA/Clustal-O + BMGE + IQ-TREE2.
IN SCOPE (pipeline-derived, tractable, deterministic on deposited outputs)
- D1 Dataset profiling (ENA metadata only): N runs (=57), depth split (39 epilimnion / 18 hypolimnion), total volume (~830 Gb), N deposited MAG assemblies. Already largely done from ENA API. No «our HPC» needed.
- R1 barrnap on deposited MAGs (PRIMARY P16): run
barrnap --kingdom bacwith default parameters on every deposited prokaryotic MAG; tabulate 16S/23S/5S rRNA gene hits per MAG and in total. Reproduces the exact named pipeline step. - R2 CheckM v1.0.18 on deposited MAGs (SECONDARY): reproduce the completeness/contamination distribution reported in Suppl. Fig S3A — 502 HQ (≥90% comp, 22.6%) / 594 MQ (70–89%, 26.8%) / 1118 partial (40–70%, 50%).
STRETCH (heavier, attempt if feasible)
- R3 GTDB-Tk r95 phylum breakdown of MAGs (Bacteroidota 690, Proteobacteria 677, Actinobacteriota 453, Verrucomicrobiota 196, Planctomycetota 92). Needs GTDB r95 reference (~30 GB) — feasible but heavy.
- R4 dRep -comp 40 -con 5 → 855 dereplicated genomes from the full bin set.
OUT OF SCOPE (not attempted — why)
- Full MEGAHIT reassembly of 830 Gb across 57 metagenomes: enormous compute and, more importantly, assembly+binning is not bit-reproducible across tool/version — re-assembly would not reproduce the identities of the 2214 MAGs, so comparing a fresh MAG count to the reported one is not a faithful 1:1 test. We instead run the deterministic downstream tools (barrnap, CheckM) on the authors' own deposited MAGs, which is the faithful reproduction of those pipeline steps.
- Wet-lab: CARD-FISH counts, microscopy, doubling-time observations (Table 1 CARD-FISH columns), nutrient chemistry.
- Detailed NCLDV / phage / eukaryotic phylogenomics & host prediction (manual curation steps, CDD batch, not scriptable end-to-end).
- gRodon/GRiD growth-rate estimates (depend on the full reassembly + read mapping).
Honesty notes / flags to verify
- Paper Results: "1.96 billion reads, ca. 830 Gb". ENA read_count across 57 runs ≈ 6.42 billion reads (≈830 Gb at ~129 bp). The 830 Gb matches; the "1.96 billion reads" does not match the read count in the deposit → flag.
- Deposited MAG assemblies: 1898 distinct GCA / 2898 bin labels vs reported 2214 MAGs (→855 dereplicated) → reconcile which set was deposited.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a solid partial reproduction on the authors' openly deposited data: CheckM v1.0.18 reproduces the Suppl Fig S3A completeness distribution within ~1.5 percentage points on every tier (HQ 21.8% vs 22.6%, MQ 28.3% vs 26.8%, partial 49.9% vs 50.0%) when using the paper's ≥40% denominator, and barrnap 0.9 ran cleanly on all 1898 MAGs. The deviations are on the data-deposit/authors' side, not our method: the deposit (1898 GCA, 1217 ≥40%, incl. 681 sub-40% assemblies) ≠ the reported 2214 recovered / 855 dereplicated set, and the reported 1.96 billion reads is ~3.3x below the ~6.42 billion deposited (while 830 Gb volume matches). These are not fabrication-suspect — the testable scientific substance (quality distribution) holds — but the two count discrepancies are not reconcilable from shared data and need human review, so overall yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.