eDNAmap: A Metabarcoding Web Tool for Comparing Marine Biodiversity, With Special Reference to Teleost Fish.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL. eDNAmap (Inoue et al. 2025) is a two-layer paper. (1) DOWNSTREAM: the GitHub repo is a Flask+R(vegan) visualization/stats web tool that ships only a 6-sample/~140-species DEMO Excel (Miya22.xlsx) and R scripts (permanova/nMDS/hclust/pheatmap.R); the demo is NOT the paper's 220-sample dataset, and the PERMANOVA grouping columns for Fig 4 (WataseHB/OsumiHB) are in no shipped file -> the downstream figure numbers cannot be reproduced from shipped artifacts. (2) UPSTREAM: the headline 4847 ASVs comes from a QIIME2 2024.10.1 + DADA2 + cutadapt + BLAST/MIDORI2 pipeline whose code is NOT shipped and whose parameters are NOT reported anywhere (no DADA2 trunc/maxEE, no primer sequences, no BLAST cutoffs, no read-count stats). So a parameter-matched 1:1 reproduction of 4847 is impossible by construction; we did NOT tune to chase it (forbidden 20%). Instead, per P16, we ran a STANDARD MiFish DADA2 reanalysis on the real public SRA data (PRJNA1241902): cutadapt MiFish-U primer removal -> per-cruise DADA2 error models -> mergePairs -> consensus chimera removal -> per-sample singleton/doubleton/tripleton filter; this executed cleanly through download+trim and into DADA2 on «our HPC» («job»). The taxonomy-free ASV count this yields is expected to differ in magnitude from the fish-only, manually-curated 4847 and characterizes the regime rather than passing/failing the paper. SOLIDLY VERIFIED (C4): 220 SRA runs / ~104.1M reads / cruise split KH-20-9=133, KH-22-5=87, all from public metadata. NOT ATTEMPTED (documented 20%): exact 4847 match (params unspecified), BLAST/MIDORI2 fish-only taxonomy + manual curation, bottle->station aggregation (Table S1 paywalled at Wiley), Fig 4 PERMANOVA p-values (grouping vars unshipped). KEY AUDIT POINT: the paper's quantitative claims are not independently checkable from the deposited code+demo alone -- only by re-running an under-specified upstream pipeline on the SRA data; flagged for the human reviewer (not fabrication, but a reproducibility gap). NOTE: finalized on operator instruction while DADA2 «job» was still computing the final ASV integer; that value, when the job lands, goes to «infra» result_out/asv_result.json (it does not change this partial verdict).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 59assessed: 2026-06-16 ⛓ be5b9c720396
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a web-based platform (eDNAmap) effectively store, visualise and compare marine eDNA metabarcoding species/sequence composition data across locations, and be used to verify the existence of biogeographic boundaries such as the Watase line/Tokara Gap for teleost fish?
- ★ eDNAmap is a web-based platform that maps sampling locations, generates heatmaps to evaluate batch effects, and performs nMDS and cluster analyses using similarity indices on uploaded eDNA composition data resource
- ★ eDNAmap supports cross-study comparison of ASV tables derived from different genetic markers by comparing sequence composition via ASV IDs without species identification method
- ★ eDNAmap can detect and visualise potential analytical/batch effects arising from integrating data processed on different sequencing platforms method
- ★ eDNAmap can be used to verify biogeographic boundaries; teleost fish ASV compositions showed a biogeographic distinction between southern and northern regions across the Osumi and Watase hypothetical boundaries finding
- ★ eDNAmap is flexible enough to analyse non-fish taxa (e.g. dinoflagellates, corals), enabling detection of concordant biogeographic patterns across groups finding
- Version 1 of the eDNAmap database consists primarily of teleost fish data from the Northwestern Pacific compiled from 12 published papers including three research cruises resource
- Community similarity indices are calculated more accurately using ASV IDs than OTU IDs or species names due to resolution of single-nucleotide differences method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| eDNA metabarcoding (12S rRNA, MiFish primers) for teleost fish | Marine seawater, Tokara Gap/Kuroshio region (KH20-9 and KH22-5 cruises) | none | ASV-sample matrix / ASV composition per station | NextSeq500 (KH20-9) or HiSeq X (KH22-5), 2 × 150-bp paired-end; Sterivex-GP cartridge filters 0.45 μm |
| eDNA metabarcoding (18S rRNA V4 region) for dinoflagellates | Marine seawater, KH20-9 cruise | none | ASV/species composition per station | — |
| Bioinformatic preprocessing (QIIME2/DADA2) producing ASV tables | fastq sequence data from cruises | none | ASV-sample matrix (chimera/singleton removal) | Qiime2 v2024.10.1, DADA2, cutadapt |
| BLAST species identification | Teleost ASV sequences | none | assigned species names / ASV counts | MIDORI2 long database (57,969 sequences) |
| Community composition analysis (nMDS, cluster, PERMANOVA, heatmap) | KH22-5 and KH20-9 teleost and dinoflagellate ASV/species tables | none | similarity of compositions / p-values / ASV detection counts | R vegan (metaMDS, adonis), pheatmap, hclust |
- – KH22-5 and KH20-9 samples formed distinct clusters in nMDS and cluster analysis, indicating batch effects from different sequencing platforms
- – Average number of ASVs detected per station was much higher for KH22-5 than KH20-9 205.2 vs 33.8 ASVs/station
- – In KH22-5, fish composition differed significantly across Watase and Osumi HBs p1=0.016 (Watase), p1=0.049 (Osumi)
- – In KH22-5, after excluding Kuroshio-axis stations, differences across both boundaries became stronger p2=0.001 (Watase), p2=0.017 (Osumi)
- – In KH20-9, Osumi HB division was significant but Watase HB was not Osumi p1=0.001; Watase p1=0.133
- – Dinoflagellate composition differed significantly across Osumi HB but not Watase HB Osumi p1=0.002; Watase p1=0.075
- – Uploaded Monterey Bay example Excel file (OTU_Closek19) was analysed and output generated quickly, demonstrating global plotting ~14 s
- mean 205.2 vs 33.8 ASVs per station (Average ASVs per station, KH22-5 vs KH20-9 (batch effect))
- pvalue p1=0.016; p1=0.049 (PERMANOVA KH22-5 Watase and Osumi HBs (with Kuroshio stations))
- pvalue p2=0.001; p2=0.017 (PERMANOVA KH22-5 Watase and Osumi after excluding Kuroshio axis)
- pvalue Osumi p1=0.001; Watase p1=0.133 (PERMANOVA KH20-9 teleost; Osumi p2=0.004, Watase p2=0.161)
- pvalue Osumi p1=0.002; Watase p1=0.075 (PERMANOVA KH20-9 dinoflagellate; Osumi p2=0.004, Watase p2=0.139)
- count 4847 ASVs (Final total teleost ASVs after filtering)
- count 54 stations and 220 samples (Used for analysis including negative controls (KH20-9 + KH22-5))
- count 57,969 sequences (MIDORI2 long 12S rRNA reference database size)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper presents eDNAmap, a web tool for marine eDNA metabarcoding data; the primary statistical analyses are ordination and hypothesis testing of community composition. Nonmetric multidimensional scaling (nMDS) was used to visualise pairwise community dissimilarities (Jaccard or Bray-Curtis) among sampling stations, and PERMANOVA (vegan::adonis) was used to test whether stations on opposite sides of two proposed biogeographic boundaries (Watase and Osumi lines) differed in ASV composition. Analyses were performed separately per cruise (KH22-5; KH20-9) and taxon (teleost fish; dinoflagellates), and were repeated with and without stations located on the Kuroshio axis. Exact p-values are reported; no effect sizes or dispersion measures accompany them.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| PERMANOVA (vegan::adonis, permutational multivariate analysis of variance) | ASV composition comparison north vs. south of Watase and Osumi hypothetical boundaries in KH22-5 teleost fish, KH20-9 teleost fish, and KH20-9 dinoflagellates; repeated with and without Kuroshio-axis stations | KH22-5: 21 stations; KH20-9: 33 stations (exact n per boundary comparison not stated) | not stated |
| Nonmetric multidimensional scaling (nMDS, vegan::metaMDS) | Visualisation of pairwise community similarity among all sampling stations (Figures 3B, 4A-C) | 54 stations combined; split per cruise for Case Study 2 | na |
| Hierarchical clustering (hclust) | Cluster dendrogram of sampling stations by ASV composition (Figure 3C) | 54 stations (combined cruise analysis for batch-effect check) | na |
-
PERMANOVA was used to test group differences in community composition, and only p-values are reported for each comparison↳ Could also: Report the PERMANOVA R² (partial eta-squared) alongside p-values, which vegan::adonis returns by default — R² quantifies the proportion of total compositional variance attributable to the boundary grouping; with p-values alone it is not possible to gauge whether a statistically supported boundary accounts for 5% or 50% of variation, which is especially informative when sample sizes differ between cruises
-
Approximately 12 PERMANOVA tests were conducted across combinations of boundaries, cruises, taxa, and Kuroshio-exclusion conditions without multiple-testing correction↳ Could also: Apply a Bonferroni or Benjamini-Hochberg FDR correction across the family of PERMANOVA tests, or pre-specify a reduced set of primary comparisons — Correcting for multiplicity controls the probability of false positives when many related tests are performed; this would not change the most strongly supported results (e.g., p = 0.001) but would clarify the status of borderline ones (e.g., p = 0.049)
-
PERMANOVA was applied to compare group centroids without testing the assumption of equal within-group dispersion↳ Could also: Accompany PERMANOVA with a test of homogeneity of multivariate dispersion (vegan::betadisper / PERMDISP2) — PERMANOVA is sensitive to differences in within-group spread as well as location; a significant PERMANOVA result could reflect unequal dispersion rather than a shift in community centroid, and betadisper can distinguish these scenarios
-
Community similarity was calculated on raw ASV count tables after removal of singletons, doubletons, and tripletons, without explicit rarefaction or library-size normalisation before ordination↳ Could also: Rarefy samples to a common sequencing depth, or apply a variance-stabilising transformation (e.g., Hellinger, CLR), prior to nMDS and PERMANOVA — Differences in sequencing depth between cruises (average 205 ASVs per station in KH22-5 vs. 33 in KH20-9) can inflate dissimilarity estimates independently of biological composition; normalisation is a common approach to reduce this artefact
-
The Jaccard index (presence/absence) was used for fish eDNA and Bray-Curtis (read-count weighted) for dinoflagellates, with the choice stated but not compared within each taxon↳ Could also: Present ordinations under both distance metrics for at least one dataset, or test sensitivity of PERMANOVA results to metric choice — The two indices can yield different groupings when rare vs. abundant ASVs drive community patterns; a side-by-side comparison or sensitivity analysis makes the influence of this methodological choice explicit
-
Geographic structure was assessed by visual inspection of nMDS plots and a binary north/south grouping in PERMANOVA↳ Could also: Apply a Mantel test or distance-based redundancy analysis (db-RDA) to directly model the relationship between community dissimilarity and geographic distance or environmental gradients — These approaches can quantify how much of compositional turnover is explained by geographic position continuously, rather than by a dichotomous boundary assignment, which complements the boundary-hypothesis framing and addresses whether the pattern is gradient-like or step-like
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41189540 (eDNAmap)
Paper: Inoue J. et al. eDNAmap: A Metabarcoding Web Tool for Comparing Marine Biodiversity, With Special Reference to Teleost Fish. Mol Ecol Resour 2025. DOI 10.1111/1755-0998.70066 · PMID 41189540 · PMCID PMC12627913. Repo: https://github.com/jun-inoue/eDNAmap (v1.0.0). Data: SRA BioProject PRJNA1241902 (220 runs SRR32859190–SRR32859409, ~104.1M reads, ~11 GB fastq.gz; 12S MiFish amplicons, 2×150 bp).
Two-layer pipeline structure (key finding)
The paper's computational results come from two distinct layers:
-
UPSTREAM (sequence processing → ASV table). QIIME2 v2024.10.1 + DADA2 (ASV inference), cutadapt (MiFish primer removal), BLAST vs MIDORI2 long DB (taxonomy), then filtering (remove singletons/doubletons/tripletons per sample, drop negative-control and non-teleost/freshwater ASVs). Produces the headline number: "a final total of 4847 ASVs".
- No code for this layer is shipped (the GitHub repo is the downstream tool only). The paper reports no parameters: no DADA2 truncation lengths / maxEE, no cutadapt primer sequences, no BLAST %identity/e-value/coverage.
-
DOWNSTREAM (eDNAmap tool). Flask + R(vegan) app that takes an already-built ASV table (Excel) and makes maps (GMT), heatmaps (pheatmap), NMDS (metaMDS), hclust, and PERMANOVA (adonis). Scripts:
scripts/{permanova,nMDS,hclust,pheatmap}.R.- Shipped example data = a 6-sample / ~140-species DEMO (
static/ASVtables/ Miya22.xlsx: reads sheet 6 samples Samp1–6 × ~140 species; environments sheet SampleID/Cruise/Station/Lat/Lon/Depth/Day, no WataseHB/OsumiHB columns). - This demo is NOT the paper's 54-station/220-sample/4847-ASV dataset.
permanova.Rneeds200_envis.csvwithWataseHB,OsumiHB,Zone,WaterPropgrouping columns — these biogeographic boundary assignments are not in any shipped file (they were added manually for Fig 4).
- Shipped example data = a 6-sample / ~140-species DEMO (
In scope (attempted)
- C1 — total ASV count. Reproduce the upstream pipeline on the public SRA data
with the standard MiFish DADA2 workflow (cutadapt MiFish-U primers → DADA2 in
R → bimera removal → per-sample singleton/doubleton/tripleton removal) and compare
the resulting ASV count to the paper's 4847 ASVs.
- Parameters are MY documented choices (paper specifies none); this is a best-effort standard reanalysis, not a parameter-matched reproduction.
- Note: paper's 4847 is fish-only after BLAST curation; my count is taxonomy-free, so an exact match is neither expected nor pursued (the BLAST/ MIDORI2 fish-filtering + unstated thresholds + manual curation = the hard 20%).
Out of scope / not attempted (documented)
- Per-cruise per-station ASV averages (205.2 KH22-5 / 33.8 KH20-9, Fig 3D/Case Study 1): needs SRR→station→cruise mapping (Table S1, Wiley supplementary) + taxonomy filtering. Attempted only if SRA/BioSample metadata yields the cruise split cheaply; otherwise dropped as 20%.
- PERMANOVA p-values (Fig 4): grouping variables WataseHB/OsumiHB not shipped; not derivable without authors' boundary assignments → not faithfully reproducible.
- BLAST/MIDORI2 taxonomy + fish curation: thresholds unspecified, manual steps.
- Wet-lab / sampling / map cartography: not pipeline-derived.
Reproducibility verdict on the headline number
The exact value 4847 ASVs is under-specified in the paper (no DADA2/cutadapt/ BLAST parameters, no shipped upstream code, shipped demo ≠ paper data). We run an honest standard reanalysis to place the number in regime and document the gap, rather than tuning parameters to hit 4847.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a two-layer reproducibility gap, not a discrepancy or fabrication. The raw SRA data (PRJNA1241902, 220 runs) is public and was verified 1:1 for C4, but the headline 4847 ASVs comes from an upstream QIIME2/DADA2/BLAST pipeline whose code is not shipped and whose parameters are nowhere reported, while the GitHub repo ships only a 6-sample demo. The deviation therefore sits on the authors' side (under-specification) and the comparison is a taxonomy-free vs fish-only metric mismatch, so the reported values are only partly derivable. The core conclusions (per-station contrast, tool figures) could not be confirmed — the DADA2 job was still running at finalization — so the result is an explainable, solid-but-incomplete partial, criticality yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.