Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Fungal metabarcoding data integration framework for the MycoDiversity DataBase (MDDB).

J Integr Bioinform · 2020
L1 78/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 51% of all assessed papers rank 533 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce. The repo (naturalis/mycodiversity, authors' own, P16) ships the PROFUNGIS Snakemake pipeline plus a concrete reference output (test_zotu/SRR1502226_zotus_final.fa = 34 fungal ZOTUs). Re-ran the full pipeline end-to-end on the documented example run SRR1502226 (454, single-end, 7188 reads, ITS2 primers fITS7/ITS4) on «our HPC». Result: 32/34 final fungal ZOTUs recovered, and ALL 32 are 100% sequence-identical to a distinct reference ZOTU once a single uniform ~12bp 5' truncation offset (cutadapt/usearch version drift in primer trimming before the truncate-to-250 step) is accounted for. Naive byte-diff misleadingly shows 0 matches because of that offset; offset-aware comparison shows an essentially exact 1:1. The 2 unrecovered ZOTUs sit at the 0.5% abundance / contamination threshold. No fabrication signal: the shipped reference is fully regenerable from shipped code + public data. NOT attempted (the hard ~20%): the paper's database-scale aggregates (172463/110910 ZOTUs, 511 runs, Table 3/4/6) which require the full corpus and the authors' MySQL MDDB, neither of which ships per-run intermediates to check. Key unblocker: usearch11 (license-removed from repo) was obtained via R. Edgar's now-open-sourced usearch11.0.667 binary; UNITE BLAST db ships in the repo.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 78
    assessed: 2026-06-15 ⛓ 4b9c2a484dd6
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Public fungal metabarcoding data in sequence read archives are valuable but heterogeneously annotated and not uniformly processed; the paper proposes that a curated integration framework (MycoDiversity DataBase) applying uniform curation, mapping, and processing can make these data comparable for large-scale studies of fungal biodiversity.

Core claims
  • The MycoDiversity DataBase (MDDB) is a curated repository integrating public fungal metabarcoding data of environmental samples to study fungal biodiversity patterns in space and time. resource
  • Data integration is achieved through three methodologies: extensive curation of annotations, generation of mappings across repositories, and application of a uniform processing method to raw DNA data. method
  • A data acquisition pipeline retrieves SRA metadata and raw HTS data starting from a publication DOI, enabling automated linkage of literature, study, sample, and sequence data. method
  • Heterogeneous metadata (e.g., geographic location and coordinate formats) is enriched and standardized using controlled vocabularies and tools such as Geocoder/GeoNames to enable cross-study comparison. method
  • Sequence read accession numbers can be detected automatically from publication PDFs using a defined set of regular expressions for SRA, ENA, and BioProject prefixes. method
  • The ITS region is the primary fungal DNA barcode marker and OTUs/species hypotheses are used as species-level proxies for taxonomic identification against reference databases such as UNITE. mechanism
Experimental setups
Assay System Perturbation Readout Platform
DNA metabarcoding (ITS amplicon HTS) fungal communities in environmental samples (e.g., soil) none OTUs / species hypotheses (fungal community composition) various HTS platforms (not specified)
Data acquisition / metadata retrieval pipeline PubMed and NCBI SRA records for 25 fungal metabarcoding publications none PMID–SRA mappings, BioSamples, experiments, sequence run files Biopython v.1.73, NCBI Entrez E-utilities, Python v.2.7.15, SRAUtils
Regular-expression accession-number extraction from PDFs PDF text of 25 publications none detected sequence archive accession numbers (SRA/ENA/BioProject prefixes) PyPDF2 v.1.26.0
Raw sequence data retrieval and format conversion SRA SRR run objects none Sanger FASTQ sequence reads NCBI SRA Toolkit (prefetch, fastq-dump v.2.9.0)
Geographic metadata standardization/enrichment SRA sample location metadata none standardized country/continent terms and decimal latitude/longitude Geocoder v.1.38.1, GeoNames API
Key results
  • From 25 selected publications, accession numbers were detected for 22 (88%) via regular-expression search of PDFs 22 of 25 (88%)
  • Running the acquisition pipeline on one publication retrieved metadata including 166 BioSamples and 166 experiments 166 BioSamples; 166 experiments
  • Pipeline run on a single publication with PMID cross-reference to SRA took ~3:35 min and produced five CSV files totaling 393 KB 00:03:35 min; 393 KB
  • Two of the detected articles contained GENBANK identifiers (processed data) and were excluded as out of scope 2 articles
Key statistics
  • count 22 accession numbers from 25 publications (88%) (regex detection success rate from publication PDFs)
  • count 166 BioSamples and 166 experiments (metadata retrieved for one publication run)
  • other 00:03:35 min (pipeline runtime for one publication with SRA cross-reference)
  • count 458,797 species hypotheses (SHs) as of August 2018 (UNITE reference ITS database)
  • count 120,000–144,000 fungal species described; ~5.1 million predicted (estimated fungal diversity)
  • count >10 petabasepairs of open-access HTS data; ~200,000 studies; >4000 environmental metabarcoding SRA records (SRA contents as of end April 2019)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a bioinformatics database and framework paper, not an inferential-statistics study. The authors describe a pipeline for acquiring, curating, and integrating public fungal metabarcoding data from NCBI SRA into the MycoDiversity DataBase (MDDB). Quantitative reporting is limited to descriptive counts and one proportion (22 of 25 publications yielding detectable accession numbers, 88%). No formal hypothesis tests or statistical models are applied; analytical components are sequence-similarity searches (BLAST), OTU clustering, and regular-expression matching.

Replicationunclear Sample size25 publications selected as the starting corpus via keyword search; per-study sample counts noted illustratively (e.g., 166 BioSamples for one exemplar publication); no power calculation or formal sample-size justification described GroupsNo comparison groups; paper describes a database construction framework Pairingna Randomization/blindingnot stated Dispersionnone
Statistical tests used
Test Applied to n Assumptions
BLAST sequence similarity search Taxonomic annotation of OTUs against the UNITE ITS reference database not stated
OTU clustering at sequence similarity thresholds Grouping of amplicon reads into species-proxy units prior to database ingestion not stated
Regular-expression matching (proportion: 22/25 = 88%) Detection of SRA accession numbers in PDF text of 25 selected publications 25 publications na
Approaches that could also have been used
  • OTU clustering at a fixed sequence-similarity threshold was used to create species-proxy units for ingestion into MDDB
    Could also: Amplicon Sequence Variant (ASV) methods such as DADA2 or DEBLUR could also be applied, resolving reads into exact biological sequences without an arbitrary similarity cutoff — ASV units are reproducible exact sequences that can be compared directly across studies without re-clustering, which aligns particularly well with the paper's cross-study integration objective
  • Regular expressions parsed from PDF text were used to extract SRA accession numbers, achieving an 88% detection rate; one publication required manual retrieval because its accession numbers were in a table
    Could also: Named-entity recognition, structured data extraction via Europe PMC or CrossRef full-text APIs, or journal-supplementary-material parsers could also capture accession numbers in tables and figures — API- or NLP-based approaches can handle accession numbers in non-body locations (tables, supplements), potentially raising the automated retrieval rate above 88%
  • BLAST was described as the primary tool for sequence similarity searches against UNITE for OTU taxonomic annotation
    Could also: VSEARCH, USEARCH, or probabilistic classifiers such as QIIME2's q2-feature-classifier (naive Bayes trained on UNITE) could also be used for taxonomy assignment — These alternatives can provide faster runtime at the scales envisioned for MDDB, and probabilistic classifiers additionally output confidence scores alongside taxonomic assignments
  • Geographic metadata was standardized by mapping country names to GeoNames identifiers and converting coordinates to decimal format
    Could also: Mapping to established biodiversity spatial vocabularies such as Darwin Core (dwc:decimalLatitude / dwc:decimalLongitude) and the TDWG World Geographic Scheme could also be applied — Widely adopted biodiversity informatics standards would facilitate direct interoperability with aggregators such as GBIF or iDigBio without requiring additional transformation
  • The 88% accession-number detection success rate is reported as a single proportion with no uncertainty measure (22/25 publications)
    Could also: A Wilson score or Clopper-Pearson exact 95% confidence interval around this proportion could also be reported — With n = 25, the binomial uncertainty around an 88% rate is non-negligible; an interval would quantify the precision of the pipeline's performance for readers evaluating its reliability
  • The 25 initial publications were selected via keyword search ('ITS', 'fungi', 'metabarcoding') without a pre-specified systematic review protocol
    Could also: A PRISMA-style systematic literature search with pre-registered inclusion and exclusion criteria could also be used to define the corpus — A formal protocol makes corpus selection reproducible and allows readers to assess which studies are represented in the database and which might be missing
Software: Python 2.7.15 · Biopython 1.73 · PyPDF2 1.26.0 · NCBI SRA Toolkit (prefetch / fastq-dump) 2.9.0 · Geocoder 1.38.1 · BLAST · UNITE ITS reference database · SRAUtils (bootstrapper for SRA Run Info CGI)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
10
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

10.15156/BIO/587475 DOI in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
SRP026207 ENA in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
SRP043706 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
SRP066844 ENA in Introduction (http://purl.org/orb/Introduction)
no other assessed paper uses this yet
SRP067281 ENA in Introduction (http://purl.org/orb/Introduction)
no other assessed paper uses this yet
SRP070752 ENA in Introduction (http://purl.org/orb/Introduction)
no other assessed paper uses this yet
SRP075244 ENA in Introduction (http://purl.org/orb/Introduction)
no other assessed paper uses this yet
SRP087758 ENA in Introduction (http://purl.org/orb/Introduction)
no other assessed paper uses this yet
SRR1502225 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
SRR1502736 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
SRX642180 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
SRX642691 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-32463383 (MycoDiversity DataBase / PROFUNGIS)

Paper: Martorelli et al. 2020, J Integr Bioinform, "Fungal metabarcoding data integration framework for the MycoDiversity DataBase (MDDB)." Repo: https://github.com/naturalis/mycodiversity (MIT, public, Python/Snakemake).

The repo bundles three components:

  1. ncbi_data_acquisition — text-mines NCBI for SRA metadata of fungal ITS studies (lit→SRA mapper).
  2. PROFUNGIS — Snakemake pipeline: raw SRA ITS reads → denoised, abundance- & contamination-filtered ZOTUs (fungal Zero-radius OTUs).
  3. PROFUNGIS_post_processing — loads ZOTU FASTAs into the MDDB reference tables.

Pipeline-derived results in the paper (Table 4, database aggregate)

  • 511 SRR files processed (2.67 Gb); 3,037,390 raw sequences
  • 172,463 ZOTUs generated total; 110,910 assigned to Fungi
  • 25 articles / 21 SRP studies / 4,470 SRR files surveyed These are whole-database aggregates requiring the full 511-run processing AND the authors' MySQL MDDB instance. → OUT OF SCOPE (the hard ~20%): no shipped DB, no per-run intermediates, ~hundreds of GB + many runs. Not attempted; noted.

IN SCOPE — the clean, shipped 1:1 target (the 80%)

The repo ships a concrete reference output: PROFUNGIS_post_processing/test_zotu/SRR1502226_zotus_final.fa documented as "raw file processed via startPROFUNGIS.py, primary source RUN:SRR1502226".

  • SRR1502226: 454 GS FLX Titanium, SINGLE-end, AMPLICON, 7,188 reads, 3.38 Mb, study PRJNA252425 (soil metagenome). Tiny → fast, deterministic.
  • Shipped reference = 34 fungal ZOTUs, each exactly 250 bp (confirms the 454 truncate-to-250 branch of the Snakefile).

Claim C1 (reproduced): Run PROFUNGIS end-to-end on SRR1502226 with the documented ITS2 primers (fITS7 GTGARTCATCGAATCTTTG / ITS4 TCCTCCGCTTATTGATATGC, platform 454, default params maxEE=1, minLen=100, abundance 0.5%, UNITE 70%-identity contamination filter) and compare the FINAL ZOTU set to the shipped SRR1502226_zotus_final.fa:

  • C1a: number of final fungal ZOTUs (reported/shipped = 34)
  • C1b: ZOTU sequence identity (exact-sequence overlap vs shipped reference)

Pipeline per result

result pipeline tool chain
C1 final ZOTUs for SRR1502226 PROFUNGIS (Snakemake) cutadapt → usearch11 truncate/filter → vsearch derep → usearch11 unoise3 → usearch11 otutab → abundance_filter.py (0.5%) → blastn vs shipped UNITE → faSomeRecords

Dependencies / unblockers

  • usearch11 is license-removed from the repo, BUT usearch11.0.667 was open-sourced (R. Edgar) and is freely downloadable (github rcedgar/usearch_old_binaries, 64-bit) → fetched in-job.
  • UNITE BLAST DB is shipped in the repo (deps/Unite/unite.{nhr,nin,nsq}).
  • Other deps via conda (snakemake, cutadapt, vsearch, blast, sra-tools); shipped faSomeRecords replaced by a small portable Python shim (same CLI) for reliability.

Determinism caveat

Authors' exact tool versions (cutadapt/usearch/blast) are unpinned, so byte-exact match is the hard 20%. Target: ZOTU count and sequence-set overlap vs the 34 shipped ZOTUs. usearch UNOISE3 is deterministic per version/input.

Figures / tables: Table
C1a
Reported
34 final fungal ZOTUs (SRR1502226, shipped reference)
Reproduced
32 final fungal ZOTUs
within tolerance
C1b
Reported
34 reference ZOTU sequences (250 bp)
Reproduced
32/32 reproduced ZOTUs are 100% identical (offset-aware, >=120bp overlap) to a distinct reference ZOTU
exact
DB1
Reported
172463 total / 110910 fungal ZOTUs; 511 SRR runs; 3037390 raw seqs (Table 4, whole-DB aggregate)
Reproduced
NOT ATTEMPTED (out of scope: needs all 511 runs + authors' MySQL MDDB)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

For the checkable scope (repo's shipped reference for run SRR1502226), this is effectively a 1:1 reproduction: 32/34 final fungal ZOTUs recovered and all 32 are 100% sequence-identical to a distinct reference ZOTU once a uniform ~12bp 5' primer-trim offset is accounted for. The only deviations — 2 ZOTUs at the 0.5% abundance/contamination threshold and the byte-level offset — sit on our methodology/tooling side (unpinned tool versions), not the authors' or data side, with no fabrication signal. The downgrade to yellow reflects that the paper's actual headline claims are whole-DB aggregates (Table 4: 172463/110910 ZOTUs, 511 runs) that are not independently checkable (full corpus + authors' MySQL MDDB not shipped), so the central conclusion is only partially confirmed.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

187.1 k
tokens (I/O) · 13.3 M incl. cache
18 min
runtime · 0.02 CPU-h
1.3 GB
peak RAM
5 (2 failed)
HPC jobs
hummel
machine