The selection of software and database for metagenomics sequence analysis impacts the outcome of microbial profiling and pathogen detection.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡A deviation arose in the data or preprocessing
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (1:1 on the primary pipeline). Paper is well-described; the deposited dataset (PRJNA717669, 12 rat-tissue metagenomes) delivers exactly what the Methods promise. We ran the paper's host-filtering pipeline (KneadData 0.12 = Trimmomatic 0.40 + Bowtie2 --very-sensitive against the EXACT reference assemblies named in Methods: R.norvegicus GCF_015227675.2, R.rattus GCF_011064425.1, Human GCA_000001405.28) on all 12 samples on «our HPC». Every reported host-filtering statistic reproduces to 4-5 significant figures: mean post-filter 164,620.8 vs 164,610; SD 283,710.4 vs 283,715; R22.K 7,023 vs 7,012 (both 99.97%); R22.L 842,783 vs 842,789 (both 96.13%); overall 99.30% host removed vs '>99%'. Metadata claims (D1-D3) are exact, including the SD 3,003,069 and R22.K raw 27,171,645. No fabrication concern -- reported values are fully derivable from the shipped reads + described pipeline. Kraken2 (K1) is a PARTIAL: using k2_standard_20210517 (closest prebuilt to the paper's Mar-2021 data) the Standard-DB total classified = 284,280, ~10% above the paper's across-4-DB max of 256,822, consistent with expected Kraken2 standard-DB version drift (the paper built its own Standard DB ~2021). NOT attempted: K2 distinct-taxa matrix across all 9-10 profilers (out of 80% floor), CAMI-simulated precision/recall (reads not deposited findably), wet-lab Leptospira confirmation (non-computational).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether the choice of direct-read shotgun metagenomics taxonomic profiling software and reference databases affects the accuracy of microbial community profiling and pathogen (Leptospira) detection in both simulated and biological samples.
- ★ Obtaining an accurate species-level microbial profile using current direct-read metagenomics profiling software is still a challenging task. finding
- ★ Discrepancies in results from different software/database combinations can lead to significant variations in the distinct microbial taxa classified and in community characterizations. finding
- ★ Differences in database contents and read profiling algorithms are the main contributors to discrepancies between software. mechanism
- ★ Inclusion of host genomes and genomes of taxa of interest in the databases increases profiling accuracy. finding
- ★ Software included in the study differed in their ability to detect Leptospira, a major zoonotic pathogen, especially at species-level resolution. finding
- ★ Software and database selection for metagenomics analysis must be based on the purpose of the study. method
- Ten widely used direct-read metagenomics profiling software (BLASTN, DIAMOND+MEGAN, Kraken2, Bracken, Centrifuge, CLARK, CLARK-s, Metaphlan3, Kaiju) and four databases were compared. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| shotgun metagenomics taxonomic profiling (10 software: BLASTN, DIAMOND, MEGAN, Kraken2, Bracken, Centrifuge, CLARK, CLARK-s, Metaphlan3, Kaiju) | simulated mice gut microbiome samples (CAMI dataset, Illumina paired-end 150bp) | none (in silico simulated reads) | precision and recall of taxonomic classification vs. known true profile | — |
| shotgun metagenomic sequencing | kidney, spleen, and lung tissue from wild rats (Rattus rattus, Rattus norvegicus) | none | microbial taxonomic community composition | Illumina HiSeq (NEBNext Ultra DNA Library Prep Kit) |
| taxonomic profiling with varying reference databases | biological rat tissue samples | database composition (Kraken2_std, Kraken2_mini, Kraken2_max, Kraken2_cus) | differences in taxonomic profile depending on database used | Kraken2 |
| pathogen detection (Leptospira) comparison across software vs. traditional methods | rat tissue samples | none | presence/absence and species-level resolution of Leptospira | — |
| DNA quality/purity assessment | extracted rat tissue DNA | none | DNA purity (OD260/OD280) and concentration | Nanodrop, Qubit 2.0, agarose gel |
| read pre-processing/quality filtering and host-read removal | sequenced reads from rat samples | none | quality-filtered, host-depleted reads | KneadData with Trimmomatic v0.33 and Bowtie2 v2.3; assessed with FastQC |
| rarefaction analysis of read depth needed for community characterization | profiling output of each software | subsampling at different read depths | expected number of reads required to characterize microbial community | R package 'vegan' (rarefy function) |
- – Current direct-read metagenomics profiling software still struggle to produce an accurate species-level microbial profile.
- – Different software and database combinations produced significant variations in microbial taxa classified and in differentially abundant taxa identified.
- – Software differed in ability to detect Leptospira, particularly at the species level.
- count 64 simulated mice gut microbiome samples from 12 mice (CAMI benchmarking dataset)
- count 3 simulated Illumina samples used (Sim.0, Sim.1, Sim.2), ~5GB per sample (used for precision/recall evaluation of software)
- count 4 rats from 2 species (R. rattus: R28; R. norvegicus: R22, R26, R27) (biological samples collected from Saint Kitts)
- other 10 profiling software and 4 databases evaluated (overall study design)
- other Kraken2 database sizes: miniKrakenV2 8GB, standard 53GB, maxikraken2 150GB, custom 60GB (comparison of Kraken2 databases)
- other BLASTN nt database 172GB; DIAMOND+MEGAN nr database 218GB; Kaiju RefSeq database 234GB; CLARK bacteria+archaea+viruses+human database 168GB (database sizes used by other software)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a benchmarking/comparative study evaluating ten direct-read metagenomic taxonomic profiling software packages and up to four databases, applied to three simulated mouse gut microbiome samples and 12 biological rodent tissue samples (4 rats × 3 tissues). Software accuracy on the simulated data was quantified using precision and recall rates (based on true/false positive and false negative taxa) at each taxonomic level, calculated in R. A rarefaction analysis (the 'rarefy' function in the R package 'vegan') was also used to estimate the sequencing depth needed to characterize the microbial community for each software's output. The abstract references identification of 'differentially abundant taxa' and 'significant variations' across software/database combinations, but the specific statistical test(s) underlying those comparisons are not included in the provided text excerpt, which ends partway through the Methods section.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Precision and recall rate calculation (true/false positive and false negative taxa counts) at each taxonomic level | comparison of software-classified taxonomic profiles vs. the known ('true') taxonomic profiles of simulated mice gut microbiome samples | 3 simulated Illumina samples (Sim.0, Sim.1, Sim.2) | not stated |
| Rarefaction analysis (rarefy function, R package vegan) | estimating the expected number of reads required to characterize the microbial community for each software's profiling result | not stated | not stated |
-
Precision and recall were reported as single point estimates per software/database combination, computed from three simulated samples.↳ Could also: Reporting these metrics with a measure of spread (e.g., mean ± SD or a bootstrap-derived confidence interval across the simulated replicates) would also be a standard way to convey estimate uncertainty. — A spread or interval estimate helps readers judge how consistent a software's precision/recall is across samples, rather than relying on a single value per sample.
-
Differences in precision/recall and community composition across ten software and four databases were compared primarily in a descriptive fashion in the excerpt provided.↳ Could also: A formal statistical model, such as a repeated-measures or mixed-effects ANOVA treating sample as a blocking/random factor, could also be used to test whether observed differences between software/database combinations exceed expected sampling variation. — This framework would let the many pairwise software/database comparisons be evaluated against an estimate of variability rather than compared as raw descriptive numbers.
-
Only three simulated samples were used to establish the precision/recall benchmark.↳ Could also: Increasing the number of simulated replicates, or applying resampling/bootstrap approaches to the existing samples, would also support variance estimation around the precision/recall metrics. — Additional replicates or resampling can narrow the uncertainty around performance estimates for each software/database combination.
-
Community diversity/read-depth characterization used a rarefaction analysis (vegan::rarefy) without a described formal significance test between software or database outputs in this excerpt.↳ Could also: Comparing community composition or diversity between software/database outputs with a permutation-based approach, such as PERMANOVA on a Bray-Curtis or UniFrac dissimilarity matrix, is also a standard method for testing whether community profiles differ significantly. — This would extend the descriptive rarefaction/diversity comparison into a formal hypothesis-testing framework for community-level differences.
-
The study involves many comparisons across ten software, up to four databases, and multiple taxonomic levels.↳ Could also: Where formal significance tests are applied across this many comparisons, a multiple-comparison correction (e.g., Benjamini-Hochberg FDR or Bonferroni) could also be used. — Such corrections help control the false discovery or family-wise error rate when many comparisons are evaluated simultaneously.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37027361
Paper: Xu R, Rajeev S, Salvador LCM. The selection of software and database for metagenomics sequence analysis impacts the outcome of microbial profiling and pathogen detection. PLoS One 2023. DOI 10.1371/journal.pone.0284031.
Code: https://github.com/rx32940/Metagenomics_tools (shell + R scripts, third-party tools orchestrated; author's own glue code). Data: SRA BioProject PRJNA717669 (12 rat tissue metagenome runs) + CAMI-simulated mouse-gut samples (Sim.0/1/2, generated by authors, not deposited as an accession).
Study design (for context)
Benchmark of 9–10 taxonomic profilers (BLASTN, Kraken2, Bracken, CLARK, CLARK-s, Centrifuge, MetaPhlAn3, Diamond+MEGAN, Kaiju) × 4 Kraken2 databases (Standard, Mini, Maxi, Custom) on (a) 3 CAMI-simulated samples and (b) 12 real rat tissue metagenomes, plus diversity/DA/pathogen-detection downstream analyses.
In scope (pipeline-derived, reproducible)
| id | result | pipeline | tractability |
|---|---|---|---|
| D1 | Dataset N: 12 samples (4 rats × 3 tissues), PRJNA717669 | SRA metadata | done (metadata) |
| D2 | Raw reads: mean ~23M, SD 3,003,069 per sample | read counting | done (metadata) |
| D3 | R22.K raw = 27,171,645 reads | read counting | done (metadata) |
| H1 | Host filtering: >99% of reads removed as host DNA | KneadData (Trimmomatic SLIDINGWINDOW:4:20 MINLEN:50 + Bowtie2 --very-sensitive vs R.norv GCF_015227675.2, R.rattus GCF_011064425.1, Human GCA_000001405.28) | needs «our HPC» — primary compute target |
| H2 | Post-filter reads: avg 164,610 (SD 283,715); R22.K→7,012 (99.97%); R22.L→842,789 (96.13%) | KneadData | needs «our HPC» |
| K1 | Kraken2 classified-read counts (S1 Table): 129,061–256,822 across 4 DBs | Kraken2 standard/custom DB | needs «our HPC» (DB build/download heavy) |
| K2 | Distinct taxa per tool (S1 Table): 18 (MetaPhlAn3) – 4816 (Kaiju) | each profiler | needs «our HPC», many tools — stretch |
Out of scope / lower priority
- Wet-lab confirmation of Leptospira (PCR / DFA / culture, Table 3) — not computational.
- CAMI-simulated precision/recall (Fig 2) — simulated reads were generated by the authors and (as far as located) not deposited under a stable accession; reproducing the exact simulation + all 10 tools × DBs is very heavy. Attempt only if «our HPC» time allows.
- Full 10-tool × 4-DB matrix + DESeq2 DA / PERMANOVA / rarefaction — depends on having every tool's full profile; out of the 80% floor. Attempt incrementally after H1/H2/K1.
Strategy
Floor target (quick): D1–D3 (done from metadata) + H1/H2 (host filtering — one clean pipeline with several exact reported numbers). Then push to K1 (Kraken2 standard DB) and beyond as «our HPC» capacity permits. Honest partials recorded; nothing forced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a strong 1:1 reproduction: the deposited PRJNA717669 (12 rat-tissue metagenomes) delivers exactly what Methods promise, and all metadata (D1-D3) and host-filtering claims (H1-H2) reproduce to 4-5 significant figures from the documented KneadData pipeline and the exact named reference assemblies — no fabrication concern. The single non-rounding deviation is K1, where the Kraken2 Standard-DB total (284,280) runs ~10% above the paper's across-4-DB max (256,822); this is expected database-version drift on our side (k2_standard_20210517 vs the authors' self-built ~2021 DB), not an authors' defect. The cross-tool taxa matrix (K2) and CAMI-simulated precision/recall were out of the reproduction floor / not findably deposited, so the paper's core software-choice thesis was only partially exercised — but everything reproduced confirms the reported values.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.