Comparison of Metagenomics and Metatranscriptomics Tools: A Guide to Making the Right Choice.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓The central claim held under reproduction
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL reproduction (1:1 on dataset identity, partial on taxa counts). Tool-comparison paper; in-scope reproducible target = the authors' own QIIME2 16S pipeline (BioinfoIPBLN/16S-Metatranscriptomic-Analysis) on their case-study data PRJNA750303. Dataset identity reproduces EXACTLY: 40 amplicon samples, mean 80,360 reads, 250 nt, 30 EC/10 HC (groups from SRA aliases). Ran import->DADA2 denoise-paired (repo defaults trim/trunc=0, chimera consensus)->classify-sklearn(SILVA-138)->collapse L6/L7 on «our HPC» («job», 27 min; QC: 81% merged, 56% non-chimeric, 17,279 ASVs). Genus counts reproduce within ~20-25% (HC 483 vs 408, EC 799 vs 640); species counts run ~2.4-2.7x high (EC 569 vs 214, HC 274 vs 114). The QUALITATIVE result is fully reproduced: EC>HC for every metric, genera>species, and EC/HC ratios match (genera 1.57 vs 1.65, species 1.88 vs 2.08). Gaps are fully explained by documented deltas: SILVA 138 vs paper's SILVA 132, the authors' exact classifier.qza not deposited (internal cluster path), QIIME2/DADA2 version drift, and under-specified abundance/prevalence filtering before counting. No fabrication indicated. NOT attempted (out of scope): Kraken2/Bracken RNA-Seq metatranscriptome half (data not in this accession, no accession given -> data_unavailable), Mende/Almeida external simulated mocks (not deposited), runtime/speed comparisons (hardware-dependent).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 67assessed: 2026-06-22 ⛓ c450c4d1ac3d
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper aims to compare 16S/marker-gene metagenomics, shotgun metagenomics, and metatranscriptomics technologies and their bioinformatic tools, and to provide two easy-to-use Nextflow pipelines (QIIME2 for marker-gene metagenomics; Kraken2/Bracken for metatranscriptomics) to guide tool selection for microbiome studies.
- ★ 16S rRNA gene sequencing enables taxonomic identification of bacteria/archaea via hypervariable regions without amplifying human DNA, but is limited by short-read biases (GC bias, sequencing errors) and poor species-level resolution finding
- ★ Shotgun metagenomics sequencing profiles all taxonomic domains and predicted biological functions of a microbial community but does not reveal which genes are actively expressed finding
- ★ Metatranscriptomics identifies microbial community mRNAs, quantifying gene expression levels and active biological pathways, and can characterize host-microbiome symbiotic interactions finding
- ★ The authors developed two Nextflow pipelines: one using QIIME2 for marker-gene metagenomics and one using Kraken2/Bracken for metatranscriptomics, available on GitHub resource
- Approximately 99% of genes found in the human tissue gene pool are derived from microorganisms finding
- Functional redundancy exists among related bacterial taxa, so functions can be conserved despite perturbations disrupting bacterial population balance finding
- QIIME2, together with its previous version, has accumulated approximately 29,000 citations, reflecting its prominence for marker-gene analysis finding
- Metagenomics taxonomic classifiers were developed as faster alternatives to BLAST-based comparison against GenBank, trading some sensitivity for speed method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Pipeline benchmarking (marker-gene metagenomics via QIIME2; metatranscriptomics via Kraken2/Bracken, implemented in Nextflow) | simulated and experimental sequencing datasets | none | pipeline performance | Nextflow |
- ▼ 16S rRNA-based methods fail to detect more than 50% of species within the phylum Radiation, which represents 15% of the entire bacterial domain 50%
- – Approximately 99% of genes in the human tissue gene pool are derived from microorganisms rather than the host 99%
- ▲ QIIME2 and its predecessor have together accumulated around 29,000 citations 29,000 citations
- ▲ More than 4300 articles on gut microbiota were published in the last 5 years according to PubMed 4300 articles
- count ~40 trillion eukaryotic cells (estimated number of human body cells)
- count ~22,000 genes (number of genes in human genome)
- count ~100 trillion microbial cells (estimated size of human microbiota)
- count ~2 million genes (number of genes in human microbiome)
- other 99% (proportion of human tissue-pool genes derived from microorganisms)
- count >4300 articles (PubMed articles on gut microbiota in the last 5 years)
- count ~29,000 citations (combined citations of QIIME2 and its previous version)
- other >50% of species undetected; phylum Radiation = 15% of bacterial domain (limitation of 16S rRNA sequencing at species level)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a narrative review and methods-comparison article on metagenomics and metatranscriptomics tools, rather than a study reporting inferential statistical hypothesis tests. The provided text describes the biological background, sequencing technologies, and bioinformatic tools (e.g., QIIME2, Kraken2/Bracken, assemblers, taxonomic classifiers) and states that the authors developed two Nextflow pipelines evaluated on simulated and experimental datasets, but the excerpt does not describe a formal statistical testing framework (e.g., hypothesis tests, p-values, or effect sizes) for that evaluation.
-
The article describes tool/pipeline performance in narrative terms (e.g., speed, sensitivity, accuracy trade-offs) based on simulated and experimental datasets without reporting formal quantitative comparison statistics in this excerpt↳ Could also: A benchmarking framework with paired quantitative metrics (e.g., precision/recall/F1 per tool, computed on the same simulated datasets) summarized with dispersion measures across replicate simulations — Reporting quantitative benchmark metrics with variability across simulated replicates would allow readers to gauge how consistently one tool outperforms another rather than relying on qualitative description
-
Multiple bioinformatic tools/pipelines are compared narratively across categories (assembly, taxonomic classification, pre-processing)↳ Could also: A formal multi-tool comparison design (e.g., repeated-measures ANOVA or Friedman test across tools applied to the same benchmark datasets) with post-hoc correction for multiple pairwise comparisons — Such a design would let the family-wise error rate be controlled when many tools are compared simultaneously on shared benchmark data
-
Pipeline evaluation is described as using both simulated and experimental datasets, but the sample sizes/number of datasets used are not stated in this excerpt↳ Could also: Explicit power/sample-size justification or a stated number of simulated replicates and experimental samples used for evaluation — Stating the number of datasets or replicates used in benchmarking would let readers assess the precision of any reported performance differences
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36553546
Paper: Terrón-Camero LC, Gordillo-González F, Salas-Espejo E, Andrés-León E. "Comparison of Metagenomics and Metatranscriptomics Tools: A Guide to Making the Right Choice." Genes (Basel) 2022;13(12):2280. PMID 36553546 · PMC9777648 · DOI 10.3390/genes13122280.
Code: https://github.com/BioinfoIPBLN/16S-Metatranscriptomic-Analysis (own repo of the authors' group, IPBLN). Two Nextflow (DSL1) pipelines:
Qiime-pipeline/— QIIME2 16S amplicon: import → DADA2 denoise → SILVA classify-sklearn → collapse to taxonomic level → export count matrix (+ phylogeny, α/β diversity, optional metagenomeSeq differential abundance).Kraken-bracken-pipeline/— Kraken2 + Bracken classification (+ Krona), for shotgun / metatranscriptomic reads (host already removed upstream).
Type of paper: a TOOL-COMPARISON / benchmark ("a guide to making the right choice") that runs QIIME2 vs Kraken2/Bracken on (a) simulated mocks, (b) synthetic gut mocks and (c) one real case-study dataset, reporting detection accuracy (true/false positives), correlations and runtimes per tool/database.
Datasets the paper relies on
| ref | what | accession | obtainable? |
|---|---|---|---|
| Li et al. (case study) | 16S rRNA amplicon, endometrial tissue | PRJNA750303 (given) | YES — 40 AMPLICON runs SRR15276323–SRR15276362 |
| Li et al. (case study) | RNA-Seq metatranscriptome, 60 paired | NOT in PRJNA750303 (0 RNA-Seq runs there) | accession not provided → out of reach |
| Mende et al. 2012 | simulated 10/100/400-species shotgun mocks | not given in brief | external, not provided |
| Almeida et al. 2018 | synthetic gut 16S mocks A100/A500 | not given in brief | external, not provided |
The brief pins only PRJNA750303. ENA confirms it contains exactly 40 paired-end AMPLICON (16S, genomic) runs (~78k–82k read pairs each), and zero RNA-Seq/transcriptomic runs. So the metatranscriptomic (RNA-Seq) half of the case study cannot be reproduced from the provided accession.
IN SCOPE (pipeline-derived, attempted)
QIIME2 16S pipeline on PRJNA750303 (the paper's own case-study DNA data), counting
detected taxa per group (EC tumor vs HC healthy), matching the repo's
Qiime-pipeline exactly:
- C1 QIIME 16S genera detected — tumor (EC): reported 640
- C2 QIIME 16S genera detected — healthy (HC): reported 408
- C3 QIIME 16S species detected — tumor (EC): reported 214
- C4 QIIME 16S species detected — healthy (HC): reported 114
- C5 dataset N: 40 16S samples, avg 80,360 paired-reads, 250 nt (Results; used as a dataset-profiling cross-check, computed from ENA read_count).
Group labels come from the SRA sample aliases: N1–N10 = healthy (10), C1–C35 = cancer/tumor (30) — matches the paper's 30 EC / 10 HC split.
Pipeline parameters (from repo Qiime-pipeline/main.nf + nextflow.config):
qiime tools importCasavaOneEightSingleLanePerSampleDirFmt, PairedEnd, demultiplexedqiime dada2 denoise-paired— defaults: trim-left-f/r = 0, trunc-len-f/r = 0,--p-chimera-method consensusqiime feature-classifier classify-sklearnagainst a SILVA classifier.qza (paper text: SILVA132; we use a standard pre-trained SILVA Naive-Bayes classifier and DOCUMENT the version difference)qiime taxa collapse --p-level 6(genus) and--p-level 7(species), export biom→tsv; count rows with non-zero abundance per group.
OUT OF SCOPE / NOT ATTEMPTED (reason)
- RNA-Seq / metatranscriptomic Kraken/RefSeq case-study results (215/253 genera,
224/298 species; 24.2M→1.03M reads): RNA-Seq data is NOT in PRJNA750303 and no
other accession is provided →
data_unavailablefor that half. - Mende simulated 10/100/400-species results (corr 0.97/0.73/0.86; FP rates): simulated read sets not provided in the brief; external.
- Almeida A100/A500 synthetic gut results (QIIME 88 genera, corr 0.67; Kraken/SILVA 0 species; etc.): mock read sets not provide
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Dataset identity reproduces exactly (40 samples, mean 80,360 reads, 250 nt) and the central qualitative conclusion is fully confirmed — EC>HC on every metric, genera>species, and matching EC/HC ratios (genera 1.65 vs 1.57; species 2.08 vs 1.88). The quantitative gap (genera ~1.2x, species ~2.4-2.7x high) sits on the input/reference side: SILVA 138 vs the paper's SILVA 132, the authors' exact classifier.qza not deposited, version drift, and under-specified count filtering — a mix of our method choices and authors' under-specification, not fabrication. The RNA-Seq metatranscriptomic half and external mocks are correctly out of scope (data not deposited), so this is an honest, well-documented partial reproduction.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.