Taxonomic and functional metagenomic assessment of a Dolichospermum bloom in a large and deep lake south of the Alps.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce, and the headline numbers reproduce 1:1 against the authoritative public deposit. The paper's code artifact is the third-party tool FastQC 0.12.1; the data is BioProject PRJNA1074715 = single run SRR27945399 (the Lake Garda Dolichospermum lemmermannii FEM_B0920 metagenome). The paper's headline sequencing output (R1: 62,967,310 paired-end reads) matches the ENA/SRA read_count for SRR27945399 EXACTLY, and base_count (18,890,193,000) / read_count = 300 confirms R2 (150bp PE) and the NovaSeq-6000 platform from ENA metadata. These ENA values are computed by EBI from the deposited FASTQ, independent of the paper text, so this is a genuine cross-check showing no fabrication signal on the headline numbers. I did NOT re-execute FastQC on «our HPC»: the «infra» VPN requires interactive SAML 2FA approval on the operator's phone, which was not completed during the session despite ~5 clean connection attempts (vpnui driven, single SSO window, fresh 2FA links pushed to Telegram). This is an access/infrastructure constraint on our side, not a reproducibility deficiency of the paper. The FastQC+BBDuk job is fully staged (reproduction/run_fastqc.sbatch, exact pinned env) and is one sbatch away once VPN access is restored; it would also settle R3 (~11% removed). NOT attempted (hard last 20%, out of scope): R4 assembly (16,508 contigs / 76 Mbp), R5 Dolichospermum MAG metrics (4.8 Mbp, N50 40,920 bp, GC 38.1%, 305x, 99.9%/0%), R6 MetaPhlAn relative abundance (28.4%), R7 gene/feature counts (4439 CDS, 40 tRNA) — heavy, DB/param-dependent, and not the named tool.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 83assessed: 2026-06-15 ⛓ 31a0b89d0d6b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a full-shotgun metagenomic analysis of a Dolichospermum lemmermannii surface bloom in Lake Garda resolve the species' taxonomic position within the ADA (Anabaena/Dolichospermum/Aphanizomenon) group and, through genome annotation, reveal the functional and secondary-metabolite traits that promote bloom formation by heterocytous nitrogen-fixing Nostocales in oligotrophic lakes?
- ★ A near-complete metagenome-assembled genome (MAG) of Dolichospermum lemmermannii (FEM_B0920) was recovered from a Lake Garda bloom and clarified the species' taxonomic position within the genus Dolichospermum and the ADA group. finding
- ★ Genome annotation uncovered distinctive adaptive traits explaining how bloom-forming nitrogen-fixing Nostocales persist in oligotrophic lakes. finding
- ★ Biosynthetic gene clusters encoding several secondary metabolites previously unknown in southern Alpine Lake populations were identified, including geosmin, anabaenopeptins, and other bioactive compounds. finding
- ★ Untargeted shotgun metagenomics combined with phylogenomic and KEGG functional analysis is an effective approach to characterize toxigenic cyanobacterial bloom populations and complement conventional monitoring. method
- The reconstructed MAG and its annotation serve as a genomic resource for risk assessment of cyanobacterial metabolites with implications for human health and water resource use. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Shotgun metagenomic sequencing (MAG reconstruction) | Dolichospermum lemmermannii bloom surface sample, Lake Garda | none | metagenome-assembled genome / community taxonomic composition | Illumina NovaSeq-6000, 150 bp paired-end; KAPA HyperPlus library prep; DNeasy PowerWater DNA Isolation Kit |
| Cyanotoxin quantification (LC-MS/MS) | GF/C filter extract from Lake Garda bloom sample | none | concentrations of microcystins, anatoxins, cylindrospermopsin, saxitoxins | Waters Acquity UPLC coupled to SCIEX 4000 QTRAP mass spectrometer |
| Phylogenomic / taxonomic analysis (ANI, phylogenomic trees) | FEM_B0920 MAG vs. reference Dolichospermum genomes (GTDB R220) | none | Average Nucleotide Identity values and phylogenomic tree topology | GTDB-Tk 2.4.0, pyani/OrthoANIu/fastANI, GToTree, IQ-TREE 2 |
| Functional genome annotation (KEGG / gene prediction) | Dolichospermum lemmermannii draft genome | none | protein-coding genes, KEGG orthologs/modules, metabolic pathways, AMR genes | PGAP, Bakta, AMRFinderPlus, ABRicate, GhostKOALA/KEGG Mapper |
| Secondary metabolite biosynthetic gene cluster detection | Dolichospermum genomes | none | presence/identity of BGCs (e.g. geosmin, anabaenopeptins) | antiSMASH 7.1.0 |
| Targeted toxin/geosmin gene marker screening | Dolichospermum genomes and contigs | none | presence of anaC, anaF, mcyB, mcyD, mcyE, cyrJ, sxtA, geoA genes | ISeqDb 0.0.3 |
| Light microscopy phytoplankton analysis | Lake Garda water column (0–20 m), both basins | none | phytoplankton community composition | inverted microscopes |
| Physico-chemical water analysis | Lake Garda water column (multiple depths) | none | temperature, pH, dissolved oxygen, sulfate, nitrogen, phosphorus, transparency | Idronaut Ocean Seven 316 Plus / SBE 19plus SeaCAT probes; Secchi disk; standard methods |
- – Dolichospermum represented 29.5% relative abundance with very high read coverage (1015×) after assembly based on the entire read set. 1015× coverage, 29.5% relative abundance
- – Anatoxin-a (ATX-a) was detected and quantified in the bloom sample, while all other analysed cyanotoxins (microcystins, cylindrospermopsin, saxitoxins, homoATX-a) were not detected. 0.3 µg/L
- – Dolichospermum SB001 reference genome assessed at 87.8% completeness and 0.2% contamination, leading to its exclusion from the main analysis set. 87.8% completeness, 0.2% contamination
- – Secondary metabolite gene clusters for geosmin, anabaenopeptins, and other bioactive compounds were identified in the genome.
- ▼ Epilimnetic phosphorus (SRP and TP) was extremely low (SRP below detection limit, TP < 10 µg/L), confirming oligotrophic conditions during the bloom. TP < 10 µg/L
- other 1015× coverage; 29.5% relative abundance (Dolichospermum read coverage and abundance after assembly)
- count 0.3 µg L−1 (Anatoxin-a (ATX-a) concentration in bloom sample)
- other 87.8% completeness, 0.2% contamination (Dolichospermum SB001 reference genome quality)
- other ANI species boundary 0.95–0.96; different species ANI < 0.90 (ANI thresholds used for species delineation)
- count TP < 10 µg L−1; sulfate 10 mg L−1 (Epilimnion nutrient/chemistry during bloom)
- count water temperature 21.0–23.6°C; Secchi depth 4 m; pH 8.1–8.5; O2 7.7–9.4 mg L−1 (87%–120% saturation) (Environmental conditions in first 20 m during bloom)
- count 251 single-copy genes; 10000 UFBoot replicates; 60 and 98 genomes in two phylogenomic analyses (Phylogenomic analysis parameters)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This metagenomic shotgun study assembled and binned a single environmental sample from a Dolichospermum surface bloom in Lake Garda (September 2020) to produce a near-complete MAG. The primary analytical framework was bioinformatic rather than inferential-statistical: taxonomic placement relied on pairwise Average Nucleotide Identity (ANI) comparisons against curated reference genomes and phylogenomic maximum-likelihood trees (IQ-TREE with ModelFinder) with ultrafast bootstrap support (UFBoot, 10,000 replicates). Functional characterization used KEGG pathway reconstruction and antiSMASH biosynthetic gene cluster detection; environmental and chemical measurements were reported descriptively with no formal hypothesis tests, p-values, or inferential dispersion measures.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Ultrafast bootstrap (UFBoot) branch support via maximum-likelihood phylogenomic inference (IQ-TREE 2.3.4 with ModelFinder substitution model selection) | Phylogenomic trees of Dolichospermum and related Nostocales genomes (two analyses: 58-genome GTDB species-level set and 96-genome full Dolichospermum set, each with one outgroup) | 60 and 98 genomes respectively (including outgroup) | not stated |
| Pairwise Average Nucleotide Identity (ANI) comparison with published species boundary threshold (0.95–0.96) | Taxonomic species delineation of FEM_B0920 MAG against reference Dolichospermum genomes; three ANI implementations used (pyani ANIb, OrthoANIu, fastANI) | null | not stated |
| BLAST-based percentage similarity scoring (antiSMASH 7.1.0) | Identification and characterization of biosynthetic gene clusters (BGCs) in the Dolichospermum MAG and comparator genomes | null | not stated |
| Genome completeness and contamination estimation (marker-gene-based scoring, CheckM 1.2.2 and CheckM2 1.0.2) | Quality assessment of all recovered MAGs; completeness and redundancy thresholds used to filter genomes for downstream phylogenomics (completeness ≥95%, contamination ≤4%) | null | not stated |
-
Taxonomic species delineation used ANI with a 0.95–0.96 threshold computed by three tools (pyani, OrthoANIu, fastANI)↳ Could also: Digital DNA-DNA hybridization (dDDH via GGDC) or MASH genomic distance could also be applied for species-level delineation — dDDH provides an explicit probabilistic species boundary (70% dDDH corresponds to approximately 95–96% ANI) accepted by many nomenclature bodies; MASH scales more efficiently to very large reference sets and is increasingly used for rapid genome-distance screening before detailed ANI computation
-
Phylogenomic trees were built by concatenating alignments of 251 single-copy marker genes and applying maximum-likelihood inference (IQ-TREE with UFBoot support)↳ Could also: A coalescent-based approach such as ASTRAL, using individual per-gene trees as input, could also be used — Coalescent methods explicitly account for incomplete lineage sorting, which can systematically bias concatenation-based topologies in rapidly diversifying lineages; comparing both approaches can reveal whether inferred relationships are robust to the underlying modeling assumption
-
Clade support was quantified with ultrafast bootstrap (UFBoot, 10,000 replicates)↳ Could also: Gene concordance factors (gCF) and site concordance factors (sCF), both available in IQ-TREE 2, could also be reported alongside UFBoot values — Concordance factors measure what fraction of individual gene trees or alignment sites actually support each bipartition, providing a complementary perspective on clade robustness that bootstrap resampling alone does not capture, particularly for nodes with high bootstrap but low gene-level concordance
-
Read assembly used Megahit after error-correction with metaSPAdes; a single final assembler was applied↳ Could also: Running multiple independent assemblers (e.g., full metaSPAdes assembly, IDBA-UD) and comparing or merging their outputs could also be employed — Different assemblers have different sensitivities to coverage depth, repeat structure, and GC bias; a multi-assembler comparison or ensemble strategy (e.g., via metaWRAP) can improve contig recovery and provide an internal consistency check, especially useful for the high-coverage Dolichospermum fraction (1015×)
-
The study was based on a single metagenome from one bloom event with no temporal or spatial replication of the metagenomics↳ Could also: A multi-sample design including bloom versus pre-bloom or post-bloom timepoints, or spatially replicated bloom collections, could also be applied — A single sample fully supports MAG reconstruction and descriptive functional annotation, but multiple samples would allow population-level genomic variation to be assessed, recruitment of reads across time to track bloom dynamics, and more confident inference about which genomic features are bloom-associated versus constitutive
-
Environmental chemical and phytoplankton data were reported as descriptive point values or ranges at discrete depths↳ Could also: If multi-timepoint monitoring data from the LTER network were integrated, ordination or correlation analyses (e.g., RDA or CCA linking environmental variables to community composition) could also contextualize the bloom — Multivariate constrained ordination is a standard tool in lake ecology for partitioning variance in phytoplankton assemblages explained by physical, chemical, and climatic drivers; it would allow the September 2020 bloom conditions to be placed in a broader seasonal or interannual context using existing LTER datasets
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-39227168
Paper: Salmaso N, Cerasino L, Pindo M, Boscaini A. Taxonomic and functional metagenomic assessment of a Dolichospermum bloom in a large and deep lake south of the Alps. FEMS Microbiol Ecol 2024. PMID 39227168 · PMCID PMC11412076 · DOI 10.1093/femsec/fiae117.
Code artifact named in the RU: https://github.com/s-andrews/FastQC (FastQC, a third-party QC tool). Per brief rule P16, applying an existing third-party tool to the paper's own data is a fully valid reproduction.
Data: SRA BioProject PRJNA1074715 → single run SRR27945399 (SAMN39880939), Illumina NovaSeq 6000, PAIRED, the metagenome of the Lake Garda Dolichospermum lemmermannii FEM_B0920 bloom sample.
Pipeline-derived results reported in the paper
| # | Reported result | Tool (paper) | In scope? | Why |
|---|---|---|---|---|
| R1 | "62 967 310 paired-end reads" total sequencing output | sequencer / FastQC count | YES | Directly produced by counting reads — the exact job FastQC does. Low-hanging, clearly specified. |
| R2 | 150 bp paired-end reads, NovaSeq-6000 | FastQC | YES | FastQC reports sequence length + can confirm. |
| R3 | "~11%" of raw reads removed during quality processing | BBDuk (BBMap 39.05) | PARTIAL | Requires running BBDuk with the paper's (loosely specified) params; attempted as a secondary check, not core. |
| R4 | Assembly: 16 508 contigs >1 kb, 76 Mbp | metaSPAdes 3.15.5 / Megahit 1.2.9 | NO (deferred 20%) | Heavy metagenomic assembly; many under-specified params; not the named tool. |
| R5 | Dolichospermum MAG FEM_B0920: 4.8 Mbp, 189 contigs, N50 40 920 bp, GC 38.1%, coverage 305×, completeness 99.9%, contamination 0% | binning + CheckM2 | NO (deferred 20%) | Depends on full assembly+binning chain. |
| R6 | 28.4% Dolichospermum relative abundance | MetaPhlAn 4.0.6 | NO (deferred 20%) | Depends on DB version; heavy. |
| R7 | 4439 protein-coding genes, 40 tRNAs, rRNAs; antiSMASH BGCs | Bakta/PGAP, antiSMASH | NO (deferred 20%) | Depends on the assembled MAG (R5). |
Out of scope (not a pipeline reproduction): wet-lab DNA extraction, sequencing, sampling, chemical/cyanotoxin measurements.
What we attempt (80/20)
Run FastQC 0.12.1 (the named tool, exact version) on the paper's own raw data (SRR27945399 R1+R2 from ENA) and compare the read count (R1) and read length (R2) against the paper. Optionally run BBDuk (BBMap 39.05) to check the ~11% removal (R3). The deep metagenomic assembly/binning/annotation chain (R4–R7) is the hard last 20% and is explicitly not attempted — under-specified params, very heavy compute, and not the named code artifact.
All heavy compute on «our HPC» (SLURM). Data lives on «infra»; only small derived values come back to «host».
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.