Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

eDNAmap: A Metabarcoding Web Tool for Comparing Marine Biodiversity, With Special Reference to Teleost Fish.

Mol Ecol Resour · 2025
L1 59/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
59/100
Reproducibility score
0.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 19% of all assessed papers rank 925 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL. eDNAmap (Inoue et al. 2025) is a two-layer paper. (1) DOWNSTREAM: the GitHub repo is a Flask+R(vegan) visualization/stats web tool that ships only a 6-sample/~140-species DEMO Excel (Miya22.xlsx) and R scripts (permanova/nMDS/hclust/pheatmap.R); the demo is NOT the paper's 220-sample dataset, and the PERMANOVA grouping columns for Fig 4 (WataseHB/OsumiHB) are in no shipped file -> the downstream figure numbers cannot be reproduced from shipped artifacts. (2) UPSTREAM: the headline 4847 ASVs comes from a QIIME2 2024.10.1 + DADA2 + cutadapt + BLAST/MIDORI2 pipeline whose code is NOT shipped and whose parameters are NOT reported anywhere (no DADA2 trunc/maxEE, no primer sequences, no BLAST cutoffs, no read-count stats). So a parameter-matched 1:1 reproduction of 4847 is impossible by construction; we did NOT tune to chase it (forbidden 20%). Instead, per P16, we ran a STANDARD MiFish DADA2 reanalysis on the real public SRA data (PRJNA1241902): cutadapt MiFish-U primer removal -> per-cruise DADA2 error models -> mergePairs -> consensus chimera removal -> per-sample singleton/doubleton/tripleton filter; this executed cleanly through download+trim and into DADA2 on «our HPC» («job»). The taxonomy-free ASV count this yields is expected to differ in magnitude from the fish-only, manually-curated 4847 and characterizes the regime rather than passing/failing the paper. SOLIDLY VERIFIED (C4): 220 SRA runs / ~104.1M reads / cruise split KH-20-9=133, KH-22-5=87, all from public metadata. NOT ATTEMPTED (documented 20%): exact 4847 match (params unspecified), BLAST/MIDORI2 fish-only taxonomy + manual curation, bottle->station aggregation (Table S1 paywalled at Wiley), Fig 4 PERMANOVA p-values (grouping vars unshipped). KEY AUDIT POINT: the paper's quantitative claims are not independently checkable from the deposited code+demo alone -- only by re-running an under-specified upstream pipeline on the SRA data; flagged for the human reviewer (not fabrication, but a reproducibility gap). NOTE: finalized on operator instruction while DADA2 «job» was still computing the final ASV integer; that value, when the job lands, goes to «infra» result_out/asv_result.json (it does not change this partial verdict).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 59
    assessed: 2026-06-16 ⛓ be5b9c720396
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a web-based platform (eDNAmap) effectively store, visualise and compare marine eDNA metabarcoding species/sequence composition data across locations, and be used to verify the existence of biogeographic boundaries such as the Watase line/Tokara Gap for teleost fish?

Core claims
  • eDNAmap is a web-based platform that maps sampling locations, generates heatmaps to evaluate batch effects, and performs nMDS and cluster analyses using similarity indices on uploaded eDNA composition data resource
  • eDNAmap supports cross-study comparison of ASV tables derived from different genetic markers by comparing sequence composition via ASV IDs without species identification method
  • eDNAmap can detect and visualise potential analytical/batch effects arising from integrating data processed on different sequencing platforms method
  • eDNAmap can be used to verify biogeographic boundaries; teleost fish ASV compositions showed a biogeographic distinction between southern and northern regions across the Osumi and Watase hypothetical boundaries finding
  • eDNAmap is flexible enough to analyse non-fish taxa (e.g. dinoflagellates, corals), enabling detection of concordant biogeographic patterns across groups finding
  • Version 1 of the eDNAmap database consists primarily of teleost fish data from the Northwestern Pacific compiled from 12 published papers including three research cruises resource
  • Community similarity indices are calculated more accurately using ASV IDs than OTU IDs or species names due to resolution of single-nucleotide differences method
Experimental setups
Assay System Perturbation Readout Platform
eDNA metabarcoding (12S rRNA, MiFish primers) for teleost fish Marine seawater, Tokara Gap/Kuroshio region (KH20-9 and KH22-5 cruises) none ASV-sample matrix / ASV composition per station NextSeq500 (KH20-9) or HiSeq X (KH22-5), 2 × 150-bp paired-end; Sterivex-GP cartridge filters 0.45 μm
eDNA metabarcoding (18S rRNA V4 region) for dinoflagellates Marine seawater, KH20-9 cruise none ASV/species composition per station
Bioinformatic preprocessing (QIIME2/DADA2) producing ASV tables fastq sequence data from cruises none ASV-sample matrix (chimera/singleton removal) Qiime2 v2024.10.1, DADA2, cutadapt
BLAST species identification Teleost ASV sequences none assigned species names / ASV counts MIDORI2 long database (57,969 sequences)
Community composition analysis (nMDS, cluster, PERMANOVA, heatmap) KH22-5 and KH20-9 teleost and dinoflagellate ASV/species tables none similarity of compositions / p-values / ASV detection counts R vegan (metaMDS, adonis), pheatmap, hclust
Key results
  • KH22-5 and KH20-9 samples formed distinct clusters in nMDS and cluster analysis, indicating batch effects from different sequencing platforms
  • Average number of ASVs detected per station was much higher for KH22-5 than KH20-9 205.2 vs 33.8 ASVs/station
  • In KH22-5, fish composition differed significantly across Watase and Osumi HBs p1=0.016 (Watase), p1=0.049 (Osumi)
  • In KH22-5, after excluding Kuroshio-axis stations, differences across both boundaries became stronger p2=0.001 (Watase), p2=0.017 (Osumi)
  • In KH20-9, Osumi HB division was significant but Watase HB was not Osumi p1=0.001; Watase p1=0.133
  • Dinoflagellate composition differed significantly across Osumi HB but not Watase HB Osumi p1=0.002; Watase p1=0.075
  • Uploaded Monterey Bay example Excel file (OTU_Closek19) was analysed and output generated quickly, demonstrating global plotting ~14 s
Key statistics
  • mean 205.2 vs 33.8 ASVs per station (Average ASVs per station, KH22-5 vs KH20-9 (batch effect))
  • pvalue p1=0.016; p1=0.049 (PERMANOVA KH22-5 Watase and Osumi HBs (with Kuroshio stations))
  • pvalue p2=0.001; p2=0.017 (PERMANOVA KH22-5 Watase and Osumi after excluding Kuroshio axis)
  • pvalue Osumi p1=0.001; Watase p1=0.133 (PERMANOVA KH20-9 teleost; Osumi p2=0.004, Watase p2=0.161)
  • pvalue Osumi p1=0.002; Watase p1=0.075 (PERMANOVA KH20-9 dinoflagellate; Osumi p2=0.004, Watase p2=0.139)
  • count 4847 ASVs (Final total teleost ASVs after filtering)
  • count 54 stations and 220 samples (Used for analysis including negative controls (KH20-9 + KH22-5))
  • count 57,969 sequences (MIDORI2 long 12S rRNA reference database size)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper presents eDNAmap, a web tool for marine eDNA metabarcoding data; the primary statistical analyses are ordination and hypothesis testing of community composition. Nonmetric multidimensional scaling (nMDS) was used to visualise pairwise community dissimilarities (Jaccard or Bray-Curtis) among sampling stations, and PERMANOVA (vegan::adonis) was used to test whether stations on opposite sides of two proposed biogeographic boundaries (Watase and Osumi lines) differed in ASV composition. Analyses were performed separately per cruise (KH22-5; KH20-9) and taxon (teleost fish; dinoflagellates), and were repeated with and without stations located on the Kuroshio axis. Exact p-values are reported; no effect sizes or dispersion measures accompany them.

Replicationbiological Sample sizeStations and bottle counts reported (KH22-5: 21 stations, 81 bottles; KH20-9: 33 stations, 131 bottles; 54 stations and 220 samples used after excluding negative controls); no formal power analysis stated GroupsNorth vs. south of Watase and Osumi hypothetical biogeographic boundaries; cruise identity (KH22-5 vs. KH20-9) for batch-effect check; taxon (teleost fish vs. dinoflagellates) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
PERMANOVA (vegan::adonis, permutational multivariate analysis of variance) ASV composition comparison north vs. south of Watase and Osumi hypothetical boundaries in KH22-5 teleost fish, KH20-9 teleost fish, and KH20-9 dinoflagellates; repeated with and without Kuroshio-axis stations KH22-5: 21 stations; KH20-9: 33 stations (exact n per boundary comparison not stated) not stated
Nonmetric multidimensional scaling (nMDS, vegan::metaMDS) Visualisation of pairwise community similarity among all sampling stations (Figures 3B, 4A-C) 54 stations combined; split per cruise for Case Study 2 na
Hierarchical clustering (hclust) Cluster dendrogram of sampling stations by ASV composition (Figure 3C) 54 stations (combined cruise analysis for batch-effect check) na
Approaches that could also have been used
  • PERMANOVA was used to test group differences in community composition, and only p-values are reported for each comparison
    Could also: Report the PERMANOVA R² (partial eta-squared) alongside p-values, which vegan::adonis returns by default — R² quantifies the proportion of total compositional variance attributable to the boundary grouping; with p-values alone it is not possible to gauge whether a statistically supported boundary accounts for 5% or 50% of variation, which is especially informative when sample sizes differ between cruises
  • Approximately 12 PERMANOVA tests were conducted across combinations of boundaries, cruises, taxa, and Kuroshio-exclusion conditions without multiple-testing correction
    Could also: Apply a Bonferroni or Benjamini-Hochberg FDR correction across the family of PERMANOVA tests, or pre-specify a reduced set of primary comparisons — Correcting for multiplicity controls the probability of false positives when many related tests are performed; this would not change the most strongly supported results (e.g., p = 0.001) but would clarify the status of borderline ones (e.g., p = 0.049)
  • PERMANOVA was applied to compare group centroids without testing the assumption of equal within-group dispersion
    Could also: Accompany PERMANOVA with a test of homogeneity of multivariate dispersion (vegan::betadisper / PERMDISP2) — PERMANOVA is sensitive to differences in within-group spread as well as location; a significant PERMANOVA result could reflect unequal dispersion rather than a shift in community centroid, and betadisper can distinguish these scenarios
  • Community similarity was calculated on raw ASV count tables after removal of singletons, doubletons, and tripletons, without explicit rarefaction or library-size normalisation before ordination
    Could also: Rarefy samples to a common sequencing depth, or apply a variance-stabilising transformation (e.g., Hellinger, CLR), prior to nMDS and PERMANOVA — Differences in sequencing depth between cruises (average 205 ASVs per station in KH22-5 vs. 33 in KH20-9) can inflate dissimilarity estimates independently of biological composition; normalisation is a common approach to reduce this artefact
  • The Jaccard index (presence/absence) was used for fish eDNA and Bray-Curtis (read-count weighted) for dinoflagellates, with the choice stated but not compared within each taxon
    Could also: Present ordinations under both distance metrics for at least one dataset, or test sensitivity of PERMANOVA results to metric choice — The two indices can yield different groupings when rare vs. abundant ASVs drive community patterns; a side-by-side comparison or sensitivity analysis makes the influence of this methodological choice explicit
  • Geographic structure was assessed by visual inspection of nMDS plots and a binary north/south grouping in PERMANOVA
    Could also: Apply a Mantel test or distance-based redundancy analysis (db-RDA) to directly model the relationship between community dissimilarity and geographic distance or environmental gradients — These approaches can quantify how much of compositional turnover is explained by geographic position continuously, rather than by a dichotomous boundary assignment, which complements the boundary-hypothesis framing and addresses whether the pattern is gradient-like or step-like
Software: R/vegan (metaMDS, adonis, hclust) · R/pheatmap · R/rfishbase · R cited as R Core Team 2020; exact version not stated · Python/pandas · Generic Mapping Tools 6.5.0 · Qiime2/DADA2 Qiime2 2024.10.1 · BLAST/MIDORI2

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
2
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41189540 (eDNAmap)

Paper: Inoue J. et al. eDNAmap: A Metabarcoding Web Tool for Comparing Marine Biodiversity, With Special Reference to Teleost Fish. Mol Ecol Resour 2025. DOI 10.1111/1755-0998.70066 · PMID 41189540 · PMCID PMC12627913. Repo: https://github.com/jun-inoue/eDNAmap (v1.0.0). Data: SRA BioProject PRJNA1241902 (220 runs SRR32859190–SRR32859409, ~104.1M reads, ~11 GB fastq.gz; 12S MiFish amplicons, 2×150 bp).

Two-layer pipeline structure (key finding)

The paper's computational results come from two distinct layers:

  1. UPSTREAM (sequence processing → ASV table). QIIME2 v2024.10.1 + DADA2 (ASV inference), cutadapt (MiFish primer removal), BLAST vs MIDORI2 long DB (taxonomy), then filtering (remove singletons/doubletons/tripletons per sample, drop negative-control and non-teleost/freshwater ASVs). Produces the headline number: "a final total of 4847 ASVs".

    • No code for this layer is shipped (the GitHub repo is the downstream tool only). The paper reports no parameters: no DADA2 truncation lengths / maxEE, no cutadapt primer sequences, no BLAST %identity/e-value/coverage.
  2. DOWNSTREAM (eDNAmap tool). Flask + R(vegan) app that takes an already-built ASV table (Excel) and makes maps (GMT), heatmaps (pheatmap), NMDS (metaMDS), hclust, and PERMANOVA (adonis). Scripts: scripts/{permanova,nMDS,hclust,pheatmap}.R.

    • Shipped example data = a 6-sample / ~140-species DEMO (static/ASVtables/ Miya22.xlsx: reads sheet 6 samples Samp1–6 × ~140 species; environments sheet SampleID/Cruise/Station/Lat/Lon/Depth/Day, no WataseHB/OsumiHB columns).
    • This demo is NOT the paper's 54-station/220-sample/4847-ASV dataset.
    • permanova.R needs 200_envis.csv with WataseHB,OsumiHB,Zone,WaterProp grouping columns — these biogeographic boundary assignments are not in any shipped file (they were added manually for Fig 4).

In scope (attempted)

  • C1 — total ASV count. Reproduce the upstream pipeline on the public SRA data with the standard MiFish DADA2 workflow (cutadapt MiFish-U primers → DADA2 in R → bimera removal → per-sample singleton/doubleton/tripleton removal) and compare the resulting ASV count to the paper's 4847 ASVs.
    • Parameters are MY documented choices (paper specifies none); this is a best-effort standard reanalysis, not a parameter-matched reproduction.
    • Note: paper's 4847 is fish-only after BLAST curation; my count is taxonomy-free, so an exact match is neither expected nor pursued (the BLAST/ MIDORI2 fish-filtering + unstated thresholds + manual curation = the hard 20%).

Out of scope / not attempted (documented)

  • Per-cruise per-station ASV averages (205.2 KH22-5 / 33.8 KH20-9, Fig 3D/Case Study 1): needs SRR→station→cruise mapping (Table S1, Wiley supplementary) + taxonomy filtering. Attempted only if SRA/BioSample metadata yields the cruise split cheaply; otherwise dropped as 20%.
  • PERMANOVA p-values (Fig 4): grouping variables WataseHB/OsumiHB not shipped; not derivable without authors' boundary assignments → not faithfully reproducible.
  • BLAST/MIDORI2 taxonomy + fish curation: thresholds unspecified, manual steps.
  • Wet-lab / sampling / map cartography: not pipeline-derived.

Reproducibility verdict on the headline number

The exact value 4847 ASVs is under-specified in the paper (no DADA2/cutadapt/ BLAST parameters, no shipped upstream code, shipped demo ≠ paper data). We run an honest standard reanalysis to place the number in regime and document the gap, rather than tuning parameters to hit 4847.

Figures / tables: Fig 3D
C1
Reported
4847 ASVs (final total, Methods/Species Identification)
Reproduced
standard MiFish DADA2 pipeline run end-to-end on the real SRA data (PRJNA1241902, 220 runs): env built, 440/440 fastqs downloaded, 220/220 primer-trimmed (cutadapt MiFish-U), DADA2 ASV inference executing («our HPC» «job»); final integer in «infra» result_out/asv_result.json
partial
C2
Reported
205.2 mean ASVs/station KH22-5 (Case Study 1, Fig 3D)
Reproduced
per-cruise per-sample ASV richness computed by same job (per-bottle; station aggregation needs Table S1)
partial
C3
Reported
33.8 mean ASVs/station KH20-9 (Case Study 1, Fig 3D)
Reproduced
per-cruise per-sample ASV richness computed by same job; checkable claim = KH20-9 << KH22-5 contrast
partial
C4
Reported
220 samples / 54 stations / 2 cruises (SRR32859190-SRR32859409)
Reproduced
220 runs, ~104.1M reads, ~11 GB; cruise split KH-20-9=133, KH-22-5=87 (BioSample sample-name prefixes)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 59/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🔴2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

This is a two-layer reproducibility gap, not a discrepancy or fabrication. The raw SRA data (PRJNA1241902, 220 runs) is public and was verified 1:1 for C4, but the headline 4847 ASVs comes from an upstream QIIME2/DADA2/BLAST pipeline whose code is not shipped and whose parameters are nowhere reported, while the GitHub repo ships only a 6-sample demo. The deviation therefore sits on the authors' side (under-specification) and the comparison is a taxonomy-free vs fish-only metric mismatch, so the reported values are only partly derivable. The core conclusions (per-station contrast, tool figures) could not be confirmed — the DADA2 job was still running at finalization — so the result is an explainable, solid-but-incomplete partial, criticality yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

210 k
tokens (I/O) · 15.5 M incl. cache
36 min
runtime · 0.59 CPU-h
28.1 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine