Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

The selection of software and database for metagenomics sequence analysis impacts the outcome of microbial profiling and pathogen detection.

PLoS One · 2023
L1 86/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
How its reproducibility compares
86/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 70% of all assessed papers rank 334 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (1:1 on the primary pipeline). Paper is well-described; the deposited dataset (PRJNA717669, 12 rat-tissue metagenomes) delivers exactly what the Methods promise. We ran the paper's host-filtering pipeline (KneadData 0.12 = Trimmomatic 0.40 + Bowtie2 --very-sensitive against the EXACT reference assemblies named in Methods: R.norvegicus GCF_015227675.2, R.rattus GCF_011064425.1, Human GCA_000001405.28) on all 12 samples on «our HPC». Every reported host-filtering statistic reproduces to 4-5 significant figures: mean post-filter 164,620.8 vs 164,610; SD 283,710.4 vs 283,715; R22.K 7,023 vs 7,012 (both 99.97%); R22.L 842,783 vs 842,789 (both 96.13%); overall 99.30% host removed vs '>99%'. Metadata claims (D1-D3) are exact, including the SD 3,003,069 and R22.K raw 27,171,645. No fabrication concern -- reported values are fully derivable from the shipped reads + described pipeline. Kraken2 (K1) is a PARTIAL: using k2_standard_20210517 (closest prebuilt to the paper's Mar-2021 data) the Standard-DB total classified = 284,280, ~10% above the paper's across-4-DB max of 256,822, consistent with expected Kraken2 standard-DB version drift (the paper built its own Standard DB ~2021). NOT attempted: K2 distinct-taxa matrix across all 9-10 profilers (out of 80% floor), CAMI-simulated precision/recall (reads not deposited findably), wet-lab Leptospira confirmation (non-computational).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether the choice of direct-read shotgun metagenomics taxonomic profiling software and reference databases affects the accuracy of microbial community profiling and pathogen (Leptospira) detection in both simulated and biological samples.

Core claims
  • Obtaining an accurate species-level microbial profile using current direct-read metagenomics profiling software is still a challenging task. finding
  • Discrepancies in results from different software/database combinations can lead to significant variations in the distinct microbial taxa classified and in community characterizations. finding
  • Differences in database contents and read profiling algorithms are the main contributors to discrepancies between software. mechanism
  • Inclusion of host genomes and genomes of taxa of interest in the databases increases profiling accuracy. finding
  • Software included in the study differed in their ability to detect Leptospira, a major zoonotic pathogen, especially at species-level resolution. finding
  • Software and database selection for metagenomics analysis must be based on the purpose of the study. method
  • Ten widely used direct-read metagenomics profiling software (BLASTN, DIAMOND+MEGAN, Kraken2, Bracken, Centrifuge, CLARK, CLARK-s, Metaphlan3, Kaiju) and four databases were compared. resource
Experimental setups
Assay System Perturbation Readout Platform
shotgun metagenomics taxonomic profiling (10 software: BLASTN, DIAMOND, MEGAN, Kraken2, Bracken, Centrifuge, CLARK, CLARK-s, Metaphlan3, Kaiju) simulated mice gut microbiome samples (CAMI dataset, Illumina paired-end 150bp) none (in silico simulated reads) precision and recall of taxonomic classification vs. known true profile
shotgun metagenomic sequencing kidney, spleen, and lung tissue from wild rats (Rattus rattus, Rattus norvegicus) none microbial taxonomic community composition Illumina HiSeq (NEBNext Ultra DNA Library Prep Kit)
taxonomic profiling with varying reference databases biological rat tissue samples database composition (Kraken2_std, Kraken2_mini, Kraken2_max, Kraken2_cus) differences in taxonomic profile depending on database used Kraken2
pathogen detection (Leptospira) comparison across software vs. traditional methods rat tissue samples none presence/absence and species-level resolution of Leptospira
DNA quality/purity assessment extracted rat tissue DNA none DNA purity (OD260/OD280) and concentration Nanodrop, Qubit 2.0, agarose gel
read pre-processing/quality filtering and host-read removal sequenced reads from rat samples none quality-filtered, host-depleted reads KneadData with Trimmomatic v0.33 and Bowtie2 v2.3; assessed with FastQC
rarefaction analysis of read depth needed for community characterization profiling output of each software subsampling at different read depths expected number of reads required to characterize microbial community R package 'vegan' (rarefy function)
Key results
  • Current direct-read metagenomics profiling software still struggle to produce an accurate species-level microbial profile.
  • Different software and database combinations produced significant variations in microbial taxa classified and in differentially abundant taxa identified.
  • Software differed in ability to detect Leptospira, particularly at the species level.
Key statistics
  • count 64 simulated mice gut microbiome samples from 12 mice (CAMI benchmarking dataset)
  • count 3 simulated Illumina samples used (Sim.0, Sim.1, Sim.2), ~5GB per sample (used for precision/recall evaluation of software)
  • count 4 rats from 2 species (R. rattus: R28; R. norvegicus: R22, R26, R27) (biological samples collected from Saint Kitts)
  • other 10 profiling software and 4 databases evaluated (overall study design)
  • other Kraken2 database sizes: miniKrakenV2 8GB, standard 53GB, maxikraken2 150GB, custom 60GB (comparison of Kraken2 databases)
  • other BLASTN nt database 172GB; DIAMOND+MEGAN nr database 218GB; Kaiju RefSeq database 234GB; CLARK bacteria+archaea+viruses+human database 168GB (database sizes used by other software)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a benchmarking/comparative study evaluating ten direct-read metagenomic taxonomic profiling software packages and up to four databases, applied to three simulated mouse gut microbiome samples and 12 biological rodent tissue samples (4 rats × 3 tissues). Software accuracy on the simulated data was quantified using precision and recall rates (based on true/false positive and false negative taxa) at each taxonomic level, calculated in R. A rarefaction analysis (the 'rarefy' function in the R package 'vegan') was also used to estimate the sequencing depth needed to characterize the microbial community for each software's output. The abstract references identification of 'differentially abundant taxa' and 'significant variations' across software/database combinations, but the specific statistical test(s) underlying those comparisons are not included in the provided text excerpt, which ends partway through the Methods section.

Replicationmixed Sample size3 simulated Illumina mice gut microbiome samples (Sim.0–Sim.2) used for precision/recall benchmarking; 12 biological tissue samples (kidney, spleen, lung) from 4 wild rats (1 R. rattus, 3 R. norvegicus) used for the biological comparison. No formal power/sample-size calculation described. Groupsten metagenomic profiling software x up to four databases, applied to the same simulated and biological samples Pairingunclear Randomization/blindingna Dispersionunclear
Statistical tests used
Test Applied to n Assumptions
Precision and recall rate calculation (true/false positive and false negative taxa counts) at each taxonomic level comparison of software-classified taxonomic profiles vs. the known ('true') taxonomic profiles of simulated mice gut microbiome samples 3 simulated Illumina samples (Sim.0, Sim.1, Sim.2) not stated
Rarefaction analysis (rarefy function, R package vegan) estimating the expected number of reads required to characterize the microbial community for each software's profiling result not stated not stated
Approaches that could also have been used
  • Precision and recall were reported as single point estimates per software/database combination, computed from three simulated samples.
    Could also: Reporting these metrics with a measure of spread (e.g., mean ± SD or a bootstrap-derived confidence interval across the simulated replicates) would also be a standard way to convey estimate uncertainty. — A spread or interval estimate helps readers judge how consistent a software's precision/recall is across samples, rather than relying on a single value per sample.
  • Differences in precision/recall and community composition across ten software and four databases were compared primarily in a descriptive fashion in the excerpt provided.
    Could also: A formal statistical model, such as a repeated-measures or mixed-effects ANOVA treating sample as a blocking/random factor, could also be used to test whether observed differences between software/database combinations exceed expected sampling variation. — This framework would let the many pairwise software/database comparisons be evaluated against an estimate of variability rather than compared as raw descriptive numbers.
  • Only three simulated samples were used to establish the precision/recall benchmark.
    Could also: Increasing the number of simulated replicates, or applying resampling/bootstrap approaches to the existing samples, would also support variance estimation around the precision/recall metrics. — Additional replicates or resampling can narrow the uncertainty around performance estimates for each software/database combination.
  • Community diversity/read-depth characterization used a rarefaction analysis (vegan::rarefy) without a described formal significance test between software or database outputs in this excerpt.
    Could also: Comparing community composition or diversity between software/database outputs with a permutation-based approach, such as PERMANOVA on a Bray-Curtis or UniFrac dissimilarity matrix, is also a standard method for testing whether community profiles differ significantly. — This would extend the descriptive rarefaction/diversity comparison into a formal hypothesis-testing framework for community-level differences.
  • The study involves many comparisons across ten software, up to four databases, and multiple taxonomic levels.
    Could also: Where formal significance tests are applied across this many comparisons, a multiple-comparison correction (e.g., Benjamini-Hochberg FDR or Bonferroni) could also be used. — Such corrections help control the false discovery or family-wise error rate when many comparisons are evaluated simultaneously.
Software: R · R package vegan (rarefy function) · KneadData (with Trimmomatic and Bowtie2) Trimmomatic 0.33; Bowtie2 2.3 · FastQC

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37027361

Paper: Xu R, Rajeev S, Salvador LCM. The selection of software and database for metagenomics sequence analysis impacts the outcome of microbial profiling and pathogen detection. PLoS One 2023. DOI 10.1371/journal.pone.0284031.

Code: https://github.com/rx32940/Metagenomics_tools (shell + R scripts, third-party tools orchestrated; author's own glue code). Data: SRA BioProject PRJNA717669 (12 rat tissue metagenome runs) + CAMI-simulated mouse-gut samples (Sim.0/1/2, generated by authors, not deposited as an accession).

Study design (for context)

Benchmark of 9–10 taxonomic profilers (BLASTN, Kraken2, Bracken, CLARK, CLARK-s, Centrifuge, MetaPhlAn3, Diamond+MEGAN, Kaiju) × 4 Kraken2 databases (Standard, Mini, Maxi, Custom) on (a) 3 CAMI-simulated samples and (b) 12 real rat tissue metagenomes, plus diversity/DA/pathogen-detection downstream analyses.

In scope (pipeline-derived, reproducible)

id result pipeline tractability
D1 Dataset N: 12 samples (4 rats × 3 tissues), PRJNA717669 SRA metadata done (metadata)
D2 Raw reads: mean ~23M, SD 3,003,069 per sample read counting done (metadata)
D3 R22.K raw = 27,171,645 reads read counting done (metadata)
H1 Host filtering: >99% of reads removed as host DNA KneadData (Trimmomatic SLIDINGWINDOW:4:20 MINLEN:50 + Bowtie2 --very-sensitive vs R.norv GCF_015227675.2, R.rattus GCF_011064425.1, Human GCA_000001405.28) needs «our HPC» — primary compute target
H2 Post-filter reads: avg 164,610 (SD 283,715); R22.K→7,012 (99.97%); R22.L→842,789 (96.13%) KneadData needs «our HPC»
K1 Kraken2 classified-read counts (S1 Table): 129,061–256,822 across 4 DBs Kraken2 standard/custom DB needs «our HPC» (DB build/download heavy)
K2 Distinct taxa per tool (S1 Table): 18 (MetaPhlAn3) – 4816 (Kaiju) each profiler needs «our HPC», many tools — stretch

Out of scope / lower priority

  • Wet-lab confirmation of Leptospira (PCR / DFA / culture, Table 3) — not computational.
  • CAMI-simulated precision/recall (Fig 2) — simulated reads were generated by the authors and (as far as located) not deposited under a stable accession; reproducing the exact simulation + all 10 tools × DBs is very heavy. Attempt only if «our HPC» time allows.
  • Full 10-tool × 4-DB matrix + DESeq2 DA / PERMANOVA / rarefaction — depends on having every tool's full profile; out of the 80% floor. Attempt incrementally after H1/H2/K1.

Strategy

Floor target (quick): D1–D3 (done from metadata) + H1/H2 (host filtering — one clean pipeline with several exact reported numbers). Then push to K1 (Kraken2 standard DB) and beyond as «our HPC» capacity permits. Honest partials recorded; nothing forced.

Figures / tables: S1 Table
D1
Reported
12 samples (4 rats x 3 tissues)
Reproduced
12 ENA runs (R22/R26/R27/R28 x K/S/L)
exact
D2a
Reported
mean raw reads ~23 million
Reproduced
23,589,865
within tolerance
D2b
Reported
raw read SD 3,003,069
Reproduced
3,003,069
exact
D3
Reported
R22.K raw = 27,171,645 reads
Reproduced
27,171,645
exact
H1
Reported
>99% reads removed as host DNA (overall)
Reproduced
99.30% overall (per-sample 96.13-99.99%)
within tolerance
H2a
Reported
mean post-filter 164,610 (SD 283,715)
Reproduced
164,620.8 (SD 283,710.4)
within tolerance
H2b
Reported
R22.K post-filter = 7,012 (99.97%)
Reproduced
7,023 (99.97%)
within tolerance
H2c
Reported
R22.L post-filter = 842,789 (96.13%)
Reproduced
842,783 (96.13%)
within tolerance
K1
Reported
Kraken2 classified-read counts 129,061-256,822 across 4 DBs
Reproduced
284,280 (Standard DB total; per-sample 150-81,118)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 86/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5

This is a strong 1:1 reproduction: the deposited PRJNA717669 (12 rat-tissue metagenomes) delivers exactly what Methods promise, and all metadata (D1-D3) and host-filtering claims (H1-H2) reproduce to 4-5 significant figures from the documented KneadData pipeline and the exact named reference assemblies — no fabrication concern. The single non-rounding deviation is K1, where the Kraken2 Standard-DB total (284,280) runs ~10% above the paper's across-4-DB max (256,822); this is expected database-version drift on our side (k2_standard_20210517 vs the authors' self-built ~2021 DB), not an authors' defect. The cross-tool taxa matrix (K2) and CAMI-simulated precision/recall were out of the reproduction floor / not findably deposited, so the paper's core software-choice thesis was only partially exercised — but everything reproduced confirms the reported values.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

75 k
tokens (I/O) · 3.1 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.