Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Evaluation of taxonomic classification and profiling methods for long-read shotgun metagenomic sequencing datasets.

BMC Bioinformatics · 2022
L1 99/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
99/100
Reproducibility score
1.4 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 55 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (1:1). Portik et al. 2022 (BMC Bioinformatics 23:541) benchmarks 14 taxonomic classifier/profiler variants on long-read mock-community metagenomes. In-scope, fully-specified pipeline result = Stage 2: the metric-computation step (per-method kreport -> species/genus precision/recall/F1/F0.5 + TP/FP/FN at 0.001%/0.1%/1% thresholds = Table 4) for the HiFi ATCC MSA-1003 dataset (= raw accession PRJNA546278). The authors deposited all 14 per-method kreports AND the exact Jupyter notebook on OSF (bqtdu / component g8fph). We downloaded the deposit, patched the notebook's hardcoded macOS indir path, and independently re-executed the authors' notebook on a «our HPC» compute node (SLURM «job»; conda python 3.10.20, pandas 2.3.3). The regenerated species-level metrics match the paper's Table 4 for all 14 methods: TP/FP/FN integer-identical and precision/recall to the printed 2 decimals; 13/14 species rows exact, 1 within-tol (MEGAN-LR-Prot F1 computed 0.857 vs paper-printed 0.85, a 2-dp display rounding; counts/P/R identical). Both Fig-2 read-utilization values checked also reproduce exactly (MEGAN-LR-Nuc-HiFi 90.0%, BugSeq-V2 82.8%), as do the two headline qualitative claims (long-read methods reach perfect species precision FP=0; short-read methods Kraken2/Bracken give 96/112 FP species). Total read count cross-checks from the deposited Kraken2 kreport (14 unclassified + 2,419,023 root = 2,419,037 = paper N). Overall: 15/16 claims exact, 1 within-tol, 0 mismatch -> a clean computational reproduction with no fabrication indicators. NOT attempted (out of heavy, DB-bound scope): rerunning the 14 upstream classifiers from raw reads to regenerate the kreports; and the other 3 mock datasets' tables (same notebook/pipeline, deposited identically) beyond the primary HiFi ATCC MSA-1003.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 60
    assessed: 2026-06-19 ⛓ 336e0f798001
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-29
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests which of 11 taxonomic classification/profiling methods (including five long-read-specific methods) perform best on long-read (PacBio HiFi and ONT) shotgun metagenomic mock community datasets, and whether long-read sequencing provides more accurate taxonomic classification than short-read (Illumina) sequencing.

Core claims
  • Long-read classifiers generally performed best among the 11 methods tested finding
  • Several short-read classification and profiling methods produced many false positives (especially at lower abundances) and required heavy filtering, reducing recall, and produced inaccurate abundance estimates finding
  • BugSeq, MEGAN-LR & DIAMOND, and sourmash displayed high precision and recall without any filtering required finding
  • In PacBio HiFi datasets, the top-performing methods detected all species down to the 0.1% abundance level with high precision finding
  • MetaMaps and MMseqs2 required moderate filtering to achieve precision/recall comparable to top-performing methods finding
  • Read quality affected performance of methods relying on protein prediction or exact k-mer matching, which performed better with PacBio HiFi than ONT finding
  • Long-read datasets with a large proportion of shorter reads (<2 kb) produced lower precision and worse abundance estimates relative to length-filtered datasets finding
  • Long-read datasets produced significantly better classification results than short-read datasets finding
Experimental setups
Assay System Perturbation Readout Platform
Taxonomic classification/profiling (11 methods) ATCC MSA-1003 mock community (20 bacteria species, staggered abundance) none (mock community) read classification, detection metrics (precision/recall), relative abundance estimates PacBio Sequel II (HiFi)
Taxonomic classification/profiling (11 methods) ZymoBIOMICS Gut Microbiome Standard D6331 (17 species: 14 bacteria, 1 archaea, 2 yeasts, staggered abundance) none (mock community) read classification, detection metrics, relative abundance estimates PacBio Sequel II (HiFi)
Taxonomic classification/profiling (11 methods) ZymoBIOMICS D6300 (10 species, even abundance) length filtering (<2kb and >50kb removed) vs. secondary short-read-enriched dataset read classification, detection metrics, relative abundance estimates ONT GridION, R10.3 chemistry
Taxonomic classification/profiling (11 methods) ZymoBIOMICS D6300 (10 species, even abundance) length filtering and subsampling vs. secondary short-read-enriched dataset read classification, detection metrics, relative abundance estimates ONT PromethION, Q20 chemistry
Short-read taxonomic classification/profiling (short-read methods only) ATCC MSA-1003 mock community none (mock community) read classification, detection metrics, relative abundance estimates Illumina HiSeq2500 (125 bp PE, pre-trimmed)
Short-read taxonomic classification/profiling (short-read methods only) ZymoBIOMICS D6300 mock community subsampling to 20 million reads read classification, detection metrics, relative abundance estimates Illumina NovaSeq 6000 (150 bp PE)
Short-read taxonomic classification/profiling on simulated short reads derived from long reads (SR-Sim) ATCC MSA-1003 and Zymo D6300 mock communities in silico fragmentation of long reads into 150 bp non-overlapping segments (10 random segments per read) read classification, detection metrics, relative abundance estimates
Key results
  • BugSeq, MEGAN-LR & DIAMOND, and sourmash achieved high precision and recall without filtering
  • Top-performing methods detected all species down to 0.1% abundance with high precision in PacBio HiFi datasets 0.1% abundance level
  • ONT R10 length filtering removed the majority of reads as too short 873,079 short reads removed (75% of total)
  • ONT R10 length filtering removed ultra-long reads 12,129 reads removed (0.01% of total)
  • ONT Q20 length filtering removed short reads before subsampling 2.13 million reads removed (39% of total)
  • mOTUs2 'long read' mode fragmentation greatly increased effective read counts 25-35x more reads than initial long-read datasets
  • Long-read datasets produced significantly better classification results than short-read datasets
Key statistics
  • count 2,419,037 reads (HiFi ATCC MSA-1003 dataset read count)
  • other median length 8,310 bp, mean 8,492 bp, 20.54 Gb total (HiFi ATCC MSA-1003 read length/data summary)
  • count 1,978,852 reads (HiFi Zymo D6331 dataset read count)
  • other median length 8,077 bp, mean 9,092 bp, 17.99 Gb total (HiFi Zymo D6331 read length/data summary)
  • count 275,318 reads retained (23% of total) (ONT R10 Zymo D6300 after length filtering)
  • count 2,000,000 reads (subsampled) (ONT Q20 Zymo D6300 after length filtering and subsampling)
  • count ~10,038,314 reads (Illumina ATCC MSA-1003 paired-end reads)
  • count 20,000,000 reads (subsampled from ~103 million) (Illumina Zymo D6300 subsampled read count)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational benchmarking study comparing 11 taxonomic classification/profiling methods applied to long-read (PacBio HiFi, ONT) and short-read (Illumina) mock community metagenomic datasets. Performance was assessed using detection metrics (precision, recall, F-scores) and relative abundance estimates compared against known mock community compositions, rather than through classical inferential hypothesis testing. The abstract states long-read datasets produced 'significantly better results' than short-read datasets for classification, but the provided text does not describe a specific formal statistical test, p-value, or software used to support this comparison.

Replicationunclear Sample sizePerformance metrics were computed from single publicly available sequencing datasets per mock community/platform combination (e.g., ~2.4 million reads for HiFi ATCC MSA-1003, ~1.9 million reads for HiFi Zymo D6331); no biological or technical replicate sequencing runs per condition are described in this text Groups11 taxonomic classification/profiling methods applied across long-read (PacBio HiFi, ONT) and short-read (Illumina, simulated short-read) mock community datasets Pairingunclear Randomization/blindingnot stated Dispersionunclear
Approaches that could also have been used
  • The abstract characterizes long-read datasets as producing 'significantly better' classification results than short-read datasets without a formal statistical test described in this excerpt
    Could also: A paired comparison method such as a Wilcoxon signed-rank test or paired t-test across matched mock communities, or a permutation/bootstrap test on precision-recall-F-score differences — Formal hypothesis testing with a reported p-value or effect size would let readers gauge the magnitude and reliability of the long-read versus short-read performance difference beyond a qualitative descriptor
  • Method performance (precision, recall, F-score, abundance estimates) is presented as point estimates derived from single sequencing datasets per platform/community combination
    Could also: Bootstrap resampling of reads (subsampling with replacement) to generate confidence intervals around precision, recall, and F-score estimates — Resampling-based intervals would convey the uncertainty of performance metrics that arise from finite read sampling, complementing the single-dataset point estimates
  • Relative abundance estimates from each method are compared descriptively to the known true abundances of the mock communities
    Could also: A quantitative agreement statistic such as Pearson or Spearman correlation, root-mean-square error, or Bland-Altman analysis between estimated and true abundances — These approaches would numerically summarize how closely each method's abundance estimates track true composition and allow direct statistical comparison across methods
  • Eleven methods are compared across multiple datasets and metrics without a stated correction for multiple comparisons
    Could also: A multiple-testing correction such as Benjamini-Hochberg FDR or Bonferroni, applied if pairwise statistical tests among methods were performed — When many pairwise method comparisons are made across several datasets, a multiplicity correction helps control the overall false-positive rate among the comparisons
  • Read-length filtering thresholds (e.g., removing reads <2 kb or >50 kb) were applied uniformly based on qualitative observations about compatibility and performance
    Could also: A sensitivity analysis or regression modeling read length as a continuous predictor of classification performance — Modeling length effects continuously (e.g., via regression) could complement fixed-threshold filtering by quantifying how performance changes gradually with read length
Software: NanoFilt · Kraken2 · Bracken · Centrifuge · MetaPhlAn3 · mOTUs2

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36513983

Paper: Portik, Brown, Pierce-Ward (2022). Evaluation of taxonomic classification and profiling methods for long-read shotgun metagenomic sequencing datasets. BMC Bioinformatics. PMCID PMC9749362.

This is a benchmarking paper: it runs 14 taxonomic classifier/profiler variants on mock-community metagenomes (known ground truth) and reports detection + abundance accuracy metrics.

Pipeline stages

  1. Raw reads -> per-method taxonomic profile (kreport). Run each of 14 tools (Kraken2, Bracken, Centrifuge-h22/h500, MetaPhlAn3, mOTUs, Sourmash-k31/k51, MetaMaps, MMseqs2, MEGAN-LR-Prot, MEGAN-LR-Nuc-HiFi/ONT, BugSeq-V2) on each mock dataset; standardize output to kraken-report (kreport) format.
  2. kreport -> detection & abundance metrics. A pandas Jupyter notebook reads the kreports, compares to the known mock-community truth set, and computes TP/FP/FN, precision, recall, F1, F0.5 (species + genus) at 0.001%/0.1%/1% thresholds, plus relative-abundance estimates, L1/goodness-of-fit, and read utilization. This step produces Table 4 and Figs 2-8.

In scope (attempted)

  • Stage 2 — the metric-computation pipeline. Fully specified and deposited: the authors put the per-method kreports and the exact Jupyter notebooks on OSF (osf.io/bqtdu, component g8fph). We independently re-run the notebook algorithm on the deposited kreports and check we regenerate Table 4 / Figs 2-8. Primary dataset: HiFi ATCC MSA-1003 (= raw-read accession PRJNA546278, run SRR9328980). Extension: the other deposited datasets (HiFi Zymo D6331, ONT R10/Q20 Zymo D6300, Illumina, SR-Sim).
  • Dataset profiling of PRJNA546278 (raw reads) and the OSF deposit.

Out of scope / hard (not the quick floor)

  • Stage 1 — rerunning the 14 classifiers from raw reads to regenerate the kreports. Each tool requires large reference databases (Kraken2/Bracken/ Centrifuge indices, NCBI nr for DIAMOND/MEGAN-LR-Prot, nt for minimap2/ MEGAN-LR-Nuc, GTDB/GenBank for sourmash) and database versions are a known irreproducibility source. This is the genuinely heavy, DB-bound part. We do NOT claim to reproduce it wholesale. We may attempt the single lightest case (sourmash on the HiFi reads with a prebuilt GTDB database) on «our HPC» as a stretch validation that the upstream pipeline regenerates a comparable kreport.
  • Wet-lab steps (DNA extraction, library prep, sequencing): out of scope.

Why stage 2 is a valid, faithful reproduction

The brief (P16) explicitly counts running a third-party/existing tool on the paper's own data as equally valid. Here we run the authors' own deposited code on the authors' own deposited intermediate data and check it reproduces the published tables/figures — a direct computational reproduction of the results as published. Confirmed pre-check: the deposited executed notebook output already matches the printed Table 4 exactly, so the remaining question is whether the code+data deterministically regenerate it (independent execution).

Code / data pointers

Figures / tables: TableFig 2
sp_kraken2
Reported
P=0.17 R=1.00 F1=0.29 (TP20 FP96 FN0)
Reproduced
P=0.172 R=1.00 F1=0.294 (TP20 FP96 FN0)
exact
sp_bracken
Reported
P=0.15 R=1.00 F1=0.26 (TP20 FP112 FN0)
Reproduced
P=0.152 R=1.00 F1=0.264 (TP20 FP112 FN0)
exact
sp_centrifuge_h22
Reported
P=0.56 R=1.00 F1=0.71 (TP20 FP16 FN0)
Reproduced
P=0.556 R=1.00 F1=0.715 (TP20 FP16 FN0)
exact
sp_centrifuge_h500
Reported
P=0.68 R=0.95 F1=0.79 (TP19 FP9 FN1)
Reproduced
P=0.679 R=0.95 F1=0.792 (TP19 FP9 FN1)
exact
sp_metaphlan3
Reported
P=0.26 R=0.65 F1=0.37 (TP13 FP38 FN7)
Reproduced
P=0.255 R=0.65 F1=0.366 (TP13 FP38 FN7)
exact
sp_motus
Reported
P=0.92 R=0.60 F1=0.73 (TP12 FP1 FN8)
Reproduced
P=0.923 R=0.60 F1=0.727 (TP12 FP1 FN8)
exact
sp_sourmash_k31
Reported
P=0.80 R=1.00 F1=0.89 (TP20 FP5 FN0)
Reproduced
P=0.80 R=1.00 F1=0.889 (TP20 FP5 FN0)
exact
sp_sourmash_k51
Reported
P=0.87 R=1.00 F1=0.93 (TP20 FP3 FN0)
Reproduced
P=0.87 R=1.00 F1=0.930 (TP20 FP3 FN0)
exact
sp_metamaps
Reported
P=0.66 R=0.95 F1=0.78 (TP19 FP10 FN1)
Reproduced
P=0.655 R=0.95 F1=0.775 (TP19 FP10 FN1)
exact
sp_mmseqs2
Reported
P=0.83 R=0.95 F1=0.88 (TP19 FP4 FN1)
Reproduced
P=0.826 R=0.95 F1=0.884 (TP19 FP4 FN1)
exact
sp_megan_prot
Reported
P=1.00 R=0.75 F1=0.85 (TP15 FP0 FN5)
Reproduced
P=1.00 R=0.75 F1=0.857 (TP15 FP0 FN5)
within tolerance
sp_megan_nuc_hifi
Reported
P=1.00 R=0.70 F1=0.82 (TP14 FP0 FN6)
Reproduced
P=1.00 R=0.70 F1=0.824 (TP14 FP0 FN6)
exact
sp_megan_nuc_ont
Reported
P=1.00 R=0.70 F1=0.82 (TP14 FP0 FN6)
Reproduced
P=1.00 R=0.70 F1=0.824 (TP14 FP0 FN6)
exact
sp_bugseq
Reported
P=1.00 R=0.80 F1=0.89 (TP16 FP0 FN4)
Reproduced
P=1.00 R=0.80 F1=0.889 (TP16 FP0 FN4)
exact
util_megan_nuc
Reported
MEGAN-LR-Nuc-HiFi read utilization 90.0%
Reproduced
90.0%
exact
util_bugseq
Reported
BugSeq-V2 read utilization 82.8%
Reproduced
82.8%
exact
headline_lr_precision
Reported
Long-read methods (MEGAN-LR-Prot/Nuc, BugSeq) perfect species precision 1.00, FP=0
Reproduced
MEGAN-LR-Prot/Nuc-HiFi/Nuc-ONT & BugSeq-V2 all P=1.00 FP=0
exact
headline_sr_fp
Reported
Short-read methods many FP species: Kraken2 96 FP, Bracken 112 FP
Reproduced
Kraken2 FP=96, Bracken FP=112
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 99/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

164.7 k
tokens (I/O) · 8.9 M incl. cache
16 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.