Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

High-resolution metagenomic reconstruction of the freshwater spring bloom.

Microbiome · 2023
L1 69/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
69/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 35% of all assessed papers rank 745 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction (P16 third-party tool on the authors' deposited data) — paper is described well enough and data are openly deposited; pipeline steps reproduce. DESIGN (D1, from open ENA metadata): 57 metagenomes EXACT, 39 epi/18 hypo split EXACT, ~830 Gb volume within-tol. PRIMARY P16 (R1): barrnap 0.9 (the designated github.com/tseemann/barrnap tool, default params) ran on ALL 1898 deposited MAGs -> rRNA table (159 w/16S, 30 w/full complement); process reproduced (paper gives no headline rRNA number). SECONDARY (R2): CheckM v1.0.18 (exact paper version) lineage_wf on all 1898 deposited MAGs (run in 10 batches to dodge a large-N multiprocessing deadlock) -> completeness DISTRIBUTION reproduces Suppl Fig S3A within ~1.5 percentage points on every tier (HQ 21.8% vs 22.6%, MQ 28.3% vs 26.8%, partial 49.9% vs 50.0%) when using the paper's >=40% denominator. TWO HONEST FLAGS for human review: (a) reported 'ca. 1.96 billion reads' vs ~6.42 billion deposited (read_count ~3.3x; 830 Gb volume matches); (b) deposited MAG set = 1898 GCA (1217 >=40% completeness, 681 below 40%) != reported 2214 recovered / 855 dereplicated -> the public deposit is NOT the QC-filtered set. NOT attempted: full 830 Gb MEGAHIT reassembly (non-deterministic, would not reproduce the 2214 MAG identities -> not a faithful 1:1 test), STRETCH R3 GTDB-Tk phylum split + R4 dRep (heavy ref data, descoped), wet-lab CARD-FISH/microscopy, NCLDV/phage/euk phylogenomics, growth-rate estimates.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-30
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether high-frequency, high-resolution metagenomic sampling of a freshwater reservoir spring bloom (integrated with CARD-FISH and microscopy) can recover genomes of prokaryotes, eukaryotes, and viruses to reveal novel taxa, distributional dynamics, and drivers of bloom collapse that lower-resolution or single-method approaches miss.

Core claims
  • High-frequency metagenomic sampling (57 samples, 3 filter sizes, 2 depths, 37 days) recovers thousands of prokaryotic, eukaryotic, and viral genomes tracking spring bloom succession in fine detail method
  • MAG relative-abundance dynamics combined with CARD-FISH yield concordant in situ doubling time estimates for dominant genome-streamlined microbial lineages finding
  • Discordance between sequence-based and microscopy-based cryptophyte quantitation reveals hidden, highly abundant aplastidic cryptophytes, confirmed by CARD-FISH finding
  • Aplastidic cryptophytes are prevalent throughout the water column and have not previously been included in plankton dynamics models finding
  • First metagenome-assembled genomes of freshwater protists (a diatom and a haptophyte) and thousands of giant viral contigs were recovered resource
  • Contrasting distributions of giant viruses (whole water column) versus parasitic perkinsids (deeper waters) suggest giant viruses, together with nutrient limitation, drive top-down control and collapse of the epilimnetic bloom mechanism
  • 2214 bacterial MAGs (≥40% completeness, ≤5% contamination) were recovered, dereplicating to 855 genomes, with no archaeal MAGs detected resource
  • Rhodopsin genes (mostly proteorhodopsin-type) occur in ~38% of dereplicated bacterial MAGs, indicating widespread light-dependent energy acquisition in the community finding
Experimental setups
Assay System Perturbation Readout Platform
16S rRNA gene sequencing freshwater reservoir microbial community (epilimnion/hypolimnion, 0.22/5/0.8 μm filters) none (natural seasonal bloom) relative abundance of prokaryotic taxonomic groups over time
18S rRNA gene sequencing freshwater reservoir eukaryotic plankton community none (natural seasonal bloom) relative abundance of eukaryotic taxonomic groups over time
Shotgun metagenomic sequencing and MAG binning freshwater reservoir, 3 size fractions (5 μm, 0.8 μm, 0.22 μm), 2 depths, 57 samples none (natural seasonal bloom) genome recovery (completeness/contamination), taxonomic composition, gene content (e.g., rhodopsins)
CARD-FISH freshwater reservoir plankton, specific bacterial and cryptophyte lineages none (natural seasonal bloom) in situ cell abundance and doubling time estimates
Microscopy phytoplankton, ciliates, zooplankton (rotifers, crustaceans) in reservoir water column none (natural seasonal bloom) biovolume, cell counts, taxonomic identification
Grazing rate estimation heterotrophic nanoflagellates (HNF) and ciliates none (natural seasonal bloom) total bacterivory rates
Physicochemical water analysis reservoir water column (epilimnion/hypolimnion) none (natural seasonal bloom) chlorophyll-a, temperature, total phosphorus, DRP, nitrate, ammonium, DOC, silica concentrations
Rhodopsin sequence/motif analysis bacterial MAGs recovered from metagenomes none rhodopsin type (type-1/heliorhodopsin), spectral tuning motif (DTE/DTG/ATI), absorption class (green- vs blue-absorbing)
Key results
  • 2214 bacterial MAGs (≥40% completeness, ≤5% contamination) recovered, dereplicating to 855 genomes; no archaeal MAGs recovered 2214 MAGs → 855 dereplicated
  • Bacteroidota rose to over half of 16S rRNA reads at bloom onset (days 9–11), with a second peak at day 21 >50% of 16S rRNA reads
  • Total heterotrophic prokaryotic numbers increased during the phytoplankton bloom then stabilized 2 to 6 × 10^6 cells ml⁻¹
  • 615 MAGs encode rhodopsins across 326 clusters, representing ~38% of the dereplicated community 615/855 MAGs (~38%)
  • Perkinsozoa sequence abundance rose sharply in the hypolimnion while decreasing in the epilimnion up to ~40% relative abundance in hypolimnion
  • Aplastidic cryptophytes detected as highly abundant via CARD-FISH despite being underrepresented/undetected by microscopy
  • Ciliate abundance peaked on day 14, preceding an HNF peak by 4 days 4-day lag
  • Giant viral genomic contigs (thousands) recovered and distributed across the entire water column, contrasting with perkinsid depth restriction thousands of contigs
Key statistics
  • count 2214 bacterial MAGs (MAGs recovered from 57 metagenomic datasets (≥40% completeness, ≤5% contamination))
  • count 855 dereplicated genomes (prokaryotic diversity representation after dereplication)
  • fold_change 2 to 6 × 10^6 cells ml⁻¹ (increase in total heterotrophic prokaryotic numbers during bloom)
  • other >50% (Bacteroidota share of 16S rRNA reads at bloom onset (days 9–11))
  • other ~38% (326/855 MAGs) (proportion of dereplicated MAGs encoding rhodopsins)
  • count 615 MAGs / 326 clusters (MAGs encoding rhodopsins; 508 type-1 (223 clusters), 214 heliorhodopsin (103 clusters), 203 encoding both (48 clusters))
  • other ~40% (relative abundance of Perkinsozoa sequences in hypolimnion)
  • count 1.96 billion reads, ca. 830 Gb (total sequencing output across 57 metagenomic samples)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper is a descriptive, exploratory metagenomic and microscopy-based field study of a freshwater spring bloom, reporting shifts in relative abundances of taxa (16S/18S rRNA reads), counts of metagenome-assembled genomes (MAGs), and physicochemical/microscopy time series across 57 samples spanning 37 days, two depths, and three filter sizes. The provided text does not describe formal inferential statistical hypothesis testing (e.g., group comparisons with p-values); results are presented as time-series trends, proportions, and genome/functional counts.

Replicationunclear Sample sizeSample sizes are described in terms of number of metagenomic samples (57), reads (1.96 billion, ~830 Gb), filter fractions (0.22 μm, 0.8 μm, 5 μm), and MAG counts (e.g., 2214 bacterial MAGs, 855 dereplicated genomes), but no formal statistical power or replicate-based sample-size justification is stated in this excerpt. GroupsTime points, depths (epilimnion vs. hypolimnion), and filter-pore sizes Pairingna Randomization/blindingnot stated Dispersionunclear
Approaches that could also have been used
  • The study reports taxonomic and genomic shifts over time (e.g., relative abundance of Bacteroidota, Actinobacteriota) primarily through descriptive time-series plots and narrative comparison of peaks/troughs.
    Could also: Formal time-series statistical models (e.g., generalized additive models, ARIMA, or change-point detection) or community-ecology tests (e.g., PERMANOVA on Bray-Curtis dissimilarities) could also be applied — Such approaches can quantify the significance and magnitude of community shifts over time or between depths/filter fractions, complementing the visual/descriptive trend interpretation.
  • Differences in MAG recovery across filter-pore sizes (1288 from 0.22 μm, 808 from 0.8 μm, 112 from 5 μm) are described narratively as a trend.
    Could also: A statistical comparison (e.g., chi-square or rarefaction-based analysis accounting for sequencing depth) could also be used — This would allow formal assessment of whether recovery differences across filters exceed what might be expected from sequencing effort or depth alone.
  • In situ doubling time estimates and grazing rate estimates are mentioned as derived from combined metagenomic and CARD-FISH/microscopy data.
    Could also: Reporting these estimates with confidence intervals or bootstrap-derived uncertainty ranges could also be used — This would convey the precision of doubling-time and grazing-rate estimates alongside the point values, which is common practice for rate estimates derived from field abundance data.
  • Co-occurrence patterns between taxa (e.g., giant viruses and perkinsids, ciliates and HNF peaks) are described qualitatively based on timing of peaks.
    Could also: Formal correlation or cross-correlation analysis (e.g., Spearman correlation with time-lag analysis) could also be used — This would provide a quantitative measure of association strength and lag between the temporal dynamics of different taxonomic groups, in addition to visual inspection of co-occurring peaks.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36698172

Paper: Kavagutti et al. 2023, High-resolution metagenomic reconstruction of the freshwater spring bloom. Microbiome 11:11. DOI 10.1186/s40168-022-01451-4. PMID 36698172 · PMCID PMC9878933.

Designated code (P16, third-party tool): https://github.com/tseemann/barrnap — rRNA gene predictor. Paper Methods: "rRNA sequences were identified using barrnap with default parameters" (16S/23S), then compared to SILVA v138. This is a third-party tool applied to the paper's own data → reproducing it on the paper's deposited MAGs is an equally valid P16 reproduction.

Data: ENA PRJEB52406 (raw reads) + deposited MAG assemblies (GCA_* under the same project). Other deposits: NCLDV phylogenomic tree on figshare (figshare.com/s/5150884f3fbb534c0302).

Pipeline (from Methods)

QC: bbmap suite (reformat/bbduk/bbmerge, Phred 18) → assembly: MEGAHIT v1.1.5 (k=49,69,89,109,129,149) → mapping: bbwrap + jgi_summarize_bam_contig_depths → binning: MetaBAT2 (default) → ORFs: PRODIGAL v2.6.3 → gene taxonomy: MMseqs2 vs GTDB r95 → rRNA: barrnap (default) vs SILVA v138 → bin taxonomy: GTDB-Tk (r95) → QC: CheckM v1.0.18 → dereplication: dRep (-comp 40 -con 5). Viruses: VIBRANT, ViralRecall, ncldv_markersearch. Euk MAGs: BUSCO + metaeuk. Growth: RazerS3 + GRiD + gRodon. Phylogenomics: PASTA/Clustal-O + BMGE + IQ-TREE2.

IN SCOPE (pipeline-derived, tractable, deterministic on deposited outputs)

  • D1 Dataset profiling (ENA metadata only): N runs (=57), depth split (39 epilimnion / 18 hypolimnion), total volume (~830 Gb), N deposited MAG assemblies. Already largely done from ENA API. No «our HPC» needed.
  • R1 barrnap on deposited MAGs (PRIMARY P16): run barrnap --kingdom bac with default parameters on every deposited prokaryotic MAG; tabulate 16S/23S/5S rRNA gene hits per MAG and in total. Reproduces the exact named pipeline step.
  • R2 CheckM v1.0.18 on deposited MAGs (SECONDARY): reproduce the completeness/contamination distribution reported in Suppl. Fig S3A — 502 HQ (≥90% comp, 22.6%) / 594 MQ (70–89%, 26.8%) / 1118 partial (40–70%, 50%).

STRETCH (heavier, attempt if feasible)

  • R3 GTDB-Tk r95 phylum breakdown of MAGs (Bacteroidota 690, Proteobacteria 677, Actinobacteriota 453, Verrucomicrobiota 196, Planctomycetota 92). Needs GTDB r95 reference (~30 GB) — feasible but heavy.
  • R4 dRep -comp 40 -con 5 → 855 dereplicated genomes from the full bin set.

OUT OF SCOPE (not attempted — why)

  • Full MEGAHIT reassembly of 830 Gb across 57 metagenomes: enormous compute and, more importantly, assembly+binning is not bit-reproducible across tool/version — re-assembly would not reproduce the identities of the 2214 MAGs, so comparing a fresh MAG count to the reported one is not a faithful 1:1 test. We instead run the deterministic downstream tools (barrnap, CheckM) on the authors' own deposited MAGs, which is the faithful reproduction of those pipeline steps.
  • Wet-lab: CARD-FISH counts, microscopy, doubling-time observations (Table 1 CARD-FISH columns), nutrient chemistry.
  • Detailed NCLDV / phage / eukaryotic phylogenomics & host prediction (manual curation steps, CDD batch, not scriptable end-to-end).
  • gRodon/GRiD growth-rate estimates (depend on the full reassembly + read mapping).

Honesty notes / flags to verify

  • Paper Results: "1.96 billion reads, ca. 830 Gb". ENA read_count across 57 runs ≈ 6.42 billion reads (≈830 Gb at ~129 bp). The 830 Gb matches; the "1.96 billion reads" does not match the read count in the deposit → flag.
  • Deposited MAG assemblies: 1898 distinct GCA / 2898 bin labels vs reported 2214 MAGs (→855 dereplicated) → reconcile which set was deposited.
Figures / tables: Fig S3A
D1_runs
Reported
57 metagenomes
Reproduced
57 ENA runs
exact
D1_depth
Reported
39 epilimnion / 18 hypolimnion
Reproduced
39 / 18 from library_name depth codes
exact
D1_volume
Reported
ca. 830 Gb
Reproduced
~830 Gb (fastq_bytes sum)
within tolerance
D1_reads
Reported
ca. 1.96 billion reads
Reproduced
~6.42 billion reads (ENA read_count sum)
did not match
D1_mags
Reported
2214 MAGs (855 dereplicated)
Reproduced
1898 distinct GCA deposited & assessed; 1217 are >=40% completeness
partial
R1_barrnap
Reported
rRNA via barrnap (default params)
Reproduced
barrnap 0.9 default run on all 1898 deposited MAGs: 159 w/16S, 68 w/23S, 355 w/5S, 30 w/full 16S+23S+5S
partial
R2_checkm
Reported
Suppl Fig S3A: 502 HQ (22.6%) / 594 MQ (26.8%) / 1118 partial (50%) over 2214 MAGs (>=40% comp)
Reproduced
CheckM v1.0.18 on all 1898 deposited MAGs: 265 HQ / 345 MQ / 607 partial / 681 <40%; of the 1217 >=40%: HQ 21.8%, MQ 28.3%, partial 49.9% -> distribution matches S3A within ~1.5pp
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 69/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is a solid partial reproduction on the authors' openly deposited data: CheckM v1.0.18 reproduces the Suppl Fig S3A completeness distribution within ~1.5 percentage points on every tier (HQ 21.8% vs 22.6%, MQ 28.3% vs 26.8%, partial 49.9% vs 50.0%) when using the paper's ≥40% denominator, and barrnap 0.9 ran cleanly on all 1898 MAGs. The deviations are on the data-deposit/authors' side, not our method: the deposit (1898 GCA, 1217 ≥40%, incl. 681 sub-40% assemblies) ≠ the reported 2214 recovered / 855 dereplicated set, and the reported 1.96 billion reads is ~3.3x below the ~6.42 billion deposited (while 830 Gb volume matches). These are not fabrication-suspect — the testable scientific substance (quality distribution) holds — but the two count discrepancies are not reconcilable from shared data and need human review, so overall yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

90.1 k
tokens (I/O) · 3.4 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.