Long-read nanopore shotgun metagenomic DNA sequencing for river biodiversity, wildlife, pollution, and environmental health monitoring.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
FINAL. Faithful own-code (P16) reproduction of MetaBioTax on the paper's own ONT SRA reads. The taxonomy-aware NR diamond DB was built from the June-2026 NCBI NR (349G, 707M seqs). The 80/20 FLOOR is reproduced + graded with evidence intact on «host»: C1 metazoan-species count for beach sand IE_Sand_B (SRR27335113) = 85,810 raw rows by the repo metric, which abundance-thresholding shows brackets the paper's 861 (640 at reads>=2000, 1163 at reads>=1000) -> graded 'partial' honestly, since the live NR has grown ~100x since the paper's ~2024 run; C6 per-phylum profile reproduced (Chordata+Arthropoda dominate as reported) -> within-tol; C7 11 samples/12 runs -> exact. The beyond-floor seawater extension C2 (735 species) + C5 (dog top non-human mammal) was genuinely attempted (2 align submissions, ~4h each) but could not be completed: blocked first by the shared cluster CPU-RunMinutes cap, then the entire 349G «infra» work dir was reclaimed by the janitor (verified gone 2026-06-29). C2/C5 graded 'error' (could-not-complete, infra), NOT mismatch — no wrong value was ever obtained. C3/C4 optional, not attempted. HONESTY: absolute species counts cannot match exactly (NR-growth); comparison is order-of-magnitude + ranking + threshold-bracketing. Reproduction conclusion: the pipeline and its qualitative/order-of-magnitude results reproduce; exact counts do not (and provably cannot, given the live reference DB), which is the honest, expected outcome for this method.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 63assessed: 2026-06-16 ⛓ dc6faaaa9375
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-29
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-06-30
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetLong-read shotgun metagenomic sequencing of environmental DNA (eDNA) from a single assay can feasibly monitor species across the tree of life—from viruses to complex multicellular organisms—for river biodiversity, wildlife, pollution, and environmental health monitoring.
- ★ Long-read shotgun metagenomic sequencing of eDNA can simultaneously detect and quantify organismal DNA from viruses to mammals in a single assay finding
- ★ Simultaneously considering viral, microbial, and eukaryotic DNA (rather than siloed analyses) provides deeper biodiversity insights finding
- ★ The single assay can quantify DNA abundance differences across sites and sample types for human, wildlife, plant, and microbial pathogens/parasites of health, agricultural, and economic importance finding
- ★ Environmental genomic data enabled animal phylogeny and transmissible cancer analysis in blue mussel (Mytilus edulis) from natural complex community eDNA finding
- ★ A custom bioinformatics pipeline (MetaBio-Tax) was built to enable metazoan metagenomics, including deuterostomes, without host filtering or read sub-sampling, using DIAMOND alignment against NCBI NR method
- Shotgun whole-mitogenome sequencing is more consistent than metabarcoding, and shotgun metagenomic sequencing of airborne DNA was less biased than metabarcoding of the same samples finding
- The MetaBio-Tax pipeline is publicly available via GitHub and Zenodo resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| eDNA filtration and extraction | river, estuarine, sea water, and beach sand samples (Avoca River, Co. Wicklow, Ireland) | none | recovered total DNA for downstream analyses | Qiagen DNeasy Blood and Tissue Kit (modified), Millipore Sterivex-GP 0.22 µm filters |
| DNA concentration measurement (spectrophotometry) | 37 eDNA/negative field control water and sand samples | none | total eDNA concentration (ng/µl) per 100 ml filtered volume | Thermo Scientific Nanodrop 2000 Spectrophotometer |
| quantitative PCR (qPCR) | 31 eDNA water/sand samples (2022-2024), IMR32 gDNA standard curves | none | pan-eukaryotic and human eDNA concentration | qPCR with IMR32-derived standard curves |
| long-read shotgun metagenomic DNA sequencing | 11 eDNA samples from river/estuarine/sea water and beach sand | none | taxonomic identification and relative abundance (nt_bpm) of viruses, microbes, and eukaryotes | Oxford Nanopore PromethION 48, FLO-PRO002 (R9.4.1) flow cells, SQK-LSK109 ligation kit |
| DNA quality/integrity assessment | extracted eDNA from water/sand samples | none | DNA concentration, purity, and fragment integrity | Nanodrop, Qubit, 0.35% agarose gel electrophoresis |
| cloud-based metagenomic bioinformatic classification | raw nanopore reads from 11 eDNA samples | none | non-deuterostome (microbial/viral) taxonomic abundance, nt_bpm | Chan Zuckerberg (CZ) ID nanopore microbial metagenomic pipeline |
| custom metazoan metagenomic bioinformatic classification | raw nanopore reads from 11 eDNA samples, including deuterostomes | none | taxonomic classification of metazoan (protostome and deuterostome) reads without host filtering/sub-sampling | DIAMOND v2.1.7 blastx (very-sensitive) against NCBI NR; Porechop v0.2.4 for pre-processing; MetaBio-Tax v1.0 pipeline |
- – DNA concentration of the 11 sequenced eDNA samples ranged from 25.5 ng/µl to 474.4 ng/µl, compared to 1.9 ng/µl for the negative field control and 3.4 ng/µl for private well water 25.5-474.4 ng/µl
- – Sequencing of 11 pooled libraries (including negative field control) generated 50.02 Gb of delivered clean data after adapter, low-quality, and short-read (<1000 bp) removal 50.02 Gb
- – Average sequencing yield was 4.7 million bases and 1.3 million reads per sample, with an average N50 read length of 5434 bp N50=5434 bp
- – Single-assay long-read shotgun metagenomic sequencing detected biodiversity spanning viruses to vertebrates across river, estuarine, sea, and sand eDNA samples
- – Environmental genomic data allowed phylogenetic and transmissible cancer analysis of blue mussel (Mytilus edulis) directly from complex community eDNA
- count 37 (total eDNA/negative field control samples generated (2022-2024))
- count 31 (samples with DNA concentration measured by spectrometry (2022/2023))
- count 31 (samples assessed by qPCR (2022-2024))
- count 11 (samples shotgun long-read sequenced (2022/2023))
- fold_change 25.5 ng/µl to 474.4 ng/µl (DNA concentration range of sequenced samples)
- mean 50.02 Gb clean data (total delivered clean sequencing data across pooled libraries)
- mean 4.7 million bases, 1.3 million reads, N50 5434 bp (average per-sample sequencing output)
- count PW A 0.97M, RAMP10 A 0.80M, HAR C 0.68M reads (samples with fewer than 1 million reads generated)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a feasibility/methods paper demonstrating long-read ONT shotgun metagenomic sequencing of aquatic eDNA from 11 sites along an Irish river system. The analytical approach is primarily bioinformatic (taxonomic read classification via DIAMOND/CZ ID, qPCR quantification) rather than hypothesis-driven inferential statistics. Results are reported descriptively as read abundances normalized to nucleotides per million base pairs (nt_bpm), with site-to-site patterns visualized rather than formally tested. The paper text as provided is truncated, so any statistical sections appearing later in the manuscript are not assessable here.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| qPCR quantification (pan-eukaryotic and human eDNA standard curves) | All 31 2022/2023 samples; used to assess eDNA concentration and guide sequencing sample selection | 31 samples | not stated |
| Nanodrop spectrophotometry (DNA concentration) | 31 2022/2023 samples; used to measure total eDNA yield per 100 ml filtered volume | 31 samples | na |
| DIAMOND blastx alignment (taxonomic classification, best-hit, max-target-seqs 1) | Metazoan metagenomics pipeline; all reads from 11 sequenced samples aligned to NCBI NR protein database | 11 sequenced samples; average ~1.3 million reads per sample | not stated |
| CZ ID nanopore microbial metagenomic pipeline (read-count threshold: ≥2 reads per genus) | Pan-biodiversity (excl. deuterostomes) genera assessment across 11 samples | 11 sequenced samples; sub-sampled to 1M reads where applicable | not stated |
-
One sample per site was selected for sequencing based on highest DNA concentration and qPCR signal, yielding n=11 sequenced samples with no within-site replication↳ Could also: Multiple biological replicates per site could be sequenced and analyzed, with site-level estimates averaged or modeled using a mixed-effects framework — Within-site replication would allow estimation of sampling variability and support formal statistical comparison of community composition across sites; it is widely used in microbial ecology to distinguish true site differences from stochastic sampling noise
-
Site-to-site differences in taxonomic read abundances are presented descriptively (visual comparison of nt_bpm across sites) without formal statistical testing↳ Could also: Community-level multivariate methods such as PERMANOVA (permutational MANOVA on Bray-Curtis or UniFrac distances) or ordination (PCoA, NMDS) are standard in microbiome and eDNA studies for formally testing whether community composition differs across site types — These approaches provide a structured framework for assessing whether observed differences in community composition exceed what would be expected by chance, and are the norm in comparable aquatic eDNA and microbiome studies
-
A fixed read-count threshold of ≥2 reads per genus was applied as the sole detection criterion↳ Could also: Rarefaction to equal sequencing depth followed by prevalence filtering, or probabilistic contamination models (e.g., decontam), could also be applied to set detection thresholds — A fixed count threshold can behave differently across samples with varying total read counts; rarefaction or depth-normalized thresholds ensure comparability across samples, and contamination modeling (especially relevant given the negative field controls collected) can systematically distinguish signal from background
-
nt_bpm (nucleotides per million base pairs) was used as the primary abundance normalization metric within CZ ID↳ Could also: Relative abundance (proportion of reads assigned to each taxon), reads per million (RPM), or TMM/DESeq2-style library-size normalization used in differential abundance tools (e.g., MaAsLin2, ALDEx2) are also widely applied in metagenomic studies — Different normalization strategies make different assumptions about total community load and compositionality; reporting and comparing multiple normalization methods, or applying compositional data analysis (e.g., centered log-ratio transformation), is common in studies aiming to compare across highly variable sample types such as water versus sand
-
Best-hit alignment (max-target-seqs 1 in DIAMOND blastx) was used to assign each read to a single taxon↳ Could also: Lowest Common Ancestor (LCA) assignment (as implemented in tools such as MEGAN or Kraken2/Bracken) considers all significant alignments and assigns reads to the most specific node in the taxonomy tree supported by the evidence — Best-hit assignment can over-attribute reads to a single species when multiple equally good hits exist across related taxa; LCA approaches are designed to handle this ambiguity and are standard in taxonomic profiling of environmental metagenomes
-
DNA concentrations are summarized with a range and a single median value (storage duration) without accompanying dispersion measures for the main sequencing yield metrics↳ Could also: Reporting mean ± SD (or median with IQR for skewed distributions) for read counts, bases generated, and N50 per sample would also convey within-study variability — Dispersion measures allow readers to judge the consistency of the sequencing approach across samples and to contextualize any outlier samples; they are standard in methods-validation studies assessing protocol reproducibility
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-42038409
Paper: Nousias O, Duffy FG, Duffy IJ, McCauley M, Whilde J, Duffy DJ. Long-read nanopore shotgun metagenomic DNA sequencing for river biodiversity, wildlife, pollution, and environmental health monitoring. NAR Genom Bioinform 2026. DOI 10.1093/nargab/lqag040 · PMCID PMC13107125.
Code: https://github.com/nousiaso/MetaBioTax @ d58c1b2b8be0e5a7b818e7ee1ea63d11e0cc5f0e
(this is the authors' own pipeline → P16 own-code reproduction).
Data: SRA BioProject PRJNA1044147 — 11 samples (12 runs; RAMP23C has both
ONT + Illumina). Public, ENA fastq directly downloadable.
The pipeline (MetaBioTax pipeline.sh)
A single SLURM bash script:
- Download NCBI NR protein db (
nr.gz),taxdump(nodes.dmp/names.dmp), andprot.accession2taxid. diamond makedb→ taxonomy-aware NR database.diamond blastxof reads vs NR,--outfmt 6(15 cols incl.staxids,stitle). ONT reads:--max-target-seqs 10 --more-sensitive -F 15 --range-culling --top 10. Illumina:--max-target-seqs 1 --more-sensitive.- Three embedded Python scripts over the tabular alignment:
- summary.txt — per-target-taxon read counts: walks each hit's
staxidsup the NCBI tree until it hits one of 18 metazoan phyla (Arthropoda, Mollusca, Chordata, Rotifera, …). - mammal_counts.txt — counts hits whose
stitle[Organism]is in the shippedmammal_species_list.txt. - metazoa_count.txt — same against
metazoa_species_list.txt(shipped split as_part_aa+_part_ab); the number of non-zero rows = "metazoan species detected" in the paper.
- summary.txt — per-target-taxon read counts: walks each hit's
DIAMOND version: paper Methods say v2.1.7 "very-sensitive"; the repo
pipeline.sh pins module diamond/2.1.8 and uses --more-sensitive for the
actual commands. We follow the repo (the runnable artifact) and pin
diamond 2.1.8, noting this minor paper-vs-repo discrepancy.
IN SCOPE (pipeline-derived, attempted)
These are direct outputs of MetaBioTax run on the paper's own SRA reads:
| id | reported result | paper loc | sample / SRA run |
|---|---|---|---|
| C1 | 861 metazoan species detected in 10 g beach sand (SOUTHB) | Fig 3b / Abstract | IE_Sand_B = SRR27335113 |
| C2 | 735 metazoan species (NORTHC seawater) | Fig 3b | NorthBeachC = SRR27298840 |
| C3 | 24 metazoan species (negative field control) | Fig 3b | NFC Qia_A = SRR27331235 |
| C4 | 281 metazoan species (private well PWA) | Fig 3b | Pri_Well_A = SRR27332129 |
| C5 | Dog (Canis lupus familiaris) = highest-abundance non-human mammal in seawater | Fig 4a | NorthBeachC = SRR27298840 |
| C6 | per-phylum read-abundance profile (Arthropoda/Chordata/Mollusca/Rotifera…) | Fig 2/3a, summary.txt | per sample |
| C7 | 11 sequenced samples under PRJNA1044147 | Suppl. Table S1 | (metadata, exact) |
Primary attempt this run = C1 (beach sand, 861 metazoan species) + its mammal profile (C5-style) + phylum summary (C6). C2–C4 are additional samples attempted if compute budget within the 12 h walltime allows.
OUT OF SCOPE / NOT ATTEMPTED (and why)
- Exact count matching is not achievable in principle. The pipeline aligns
against the live NCBI NR, which grows continuously; the paper's run used NR
as of ~2024, ours uses NR as of June 2026. More reference proteins → generally
more species pass the regex/taxonomy match. So absolute counts are an
honest "different DB snapshot" comparison, gradable at best
partial/within-tolon order-of-magnitude + ranking, notexact. Stated up front. - Mussel mitogenome read (11 741 bp, 70.14 % of mt genome), Fig 5 — a separate
read-mapping/assembly-to-reference analysis, not produced by
pipeline.sh(no mapping/assembly code in repo). Out of scope (non_pipeline for this repo). - nt_bpm normalisation values (e.g. 398 356 nt_bpm Bradyrhizobium): a derived normalisation layered on top of the counts; the
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The reproduction faithfully set up the authors' own MetaBioTax pipeline (DIAMOND blastx vs taxonomy-aware NCBI NR) on the paper's public reads, but the heavy compute had not produced any numbers by the finalization cutoff, so only the sample-count metadata (C7, 11 samples/12 runs) is verified exact; the central claims C1/C5/C6 (861 metazoan spp., dog as top mammal, per-phylum profile) are NOT REACHED. The shortfall is on our side (unfinished run) compounded by expected NR version drift (2024 vs June-2026), not an authors' defect — values remain derivable in principle and no fabrication signal appeared. Overall a solid-but-incomplete partial reproduction with no critical discrepancy.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.