Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comparison of Metagenomics and Metatranscriptomics Tools: A Guide to Making the Right Choice.

Genes (Basel) · 2022
L1 67/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Same input data as the authors
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
67/100
Reproducibility score
0.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 29% of all assessed papers rank 795 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction (1:1 on dataset identity, partial on taxa counts). Tool-comparison paper; in-scope reproducible target = the authors' own QIIME2 16S pipeline (BioinfoIPBLN/16S-Metatranscriptomic-Analysis) on their case-study data PRJNA750303. Dataset identity reproduces EXACTLY: 40 amplicon samples, mean 80,360 reads, 250 nt, 30 EC/10 HC (groups from SRA aliases). Ran import->DADA2 denoise-paired (repo defaults trim/trunc=0, chimera consensus)->classify-sklearn(SILVA-138)->collapse L6/L7 on «our HPC» («job», 27 min; QC: 81% merged, 56% non-chimeric, 17,279 ASVs). Genus counts reproduce within ~20-25% (HC 483 vs 408, EC 799 vs 640); species counts run ~2.4-2.7x high (EC 569 vs 214, HC 274 vs 114). The QUALITATIVE result is fully reproduced: EC>HC for every metric, genera>species, and EC/HC ratios match (genera 1.57 vs 1.65, species 1.88 vs 2.08). Gaps are fully explained by documented deltas: SILVA 138 vs paper's SILVA 132, the authors' exact classifier.qza not deposited (internal cluster path), QIIME2/DADA2 version drift, and under-specified abundance/prevalence filtering before counting. No fabrication indicated. NOT attempted (out of scope): Kraken2/Bracken RNA-Seq metatranscriptome half (data not in this accession, no accession given -> data_unavailable), Mende/Almeida external simulated mocks (not deposited), runtime/speed comparisons (hardware-dependent).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 67
    assessed: 2026-06-22 ⛓ c450c4d1ac3d
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper aims to compare 16S/marker-gene metagenomics, shotgun metagenomics, and metatranscriptomics technologies and their bioinformatic tools, and to provide two easy-to-use Nextflow pipelines (QIIME2 for marker-gene metagenomics; Kraken2/Bracken for metatranscriptomics) to guide tool selection for microbiome studies.

Core claims
  • 16S rRNA gene sequencing enables taxonomic identification of bacteria/archaea via hypervariable regions without amplifying human DNA, but is limited by short-read biases (GC bias, sequencing errors) and poor species-level resolution finding
  • Shotgun metagenomics sequencing profiles all taxonomic domains and predicted biological functions of a microbial community but does not reveal which genes are actively expressed finding
  • Metatranscriptomics identifies microbial community mRNAs, quantifying gene expression levels and active biological pathways, and can characterize host-microbiome symbiotic interactions finding
  • The authors developed two Nextflow pipelines: one using QIIME2 for marker-gene metagenomics and one using Kraken2/Bracken for metatranscriptomics, available on GitHub resource
  • Approximately 99% of genes found in the human tissue gene pool are derived from microorganisms finding
  • Functional redundancy exists among related bacterial taxa, so functions can be conserved despite perturbations disrupting bacterial population balance finding
  • QIIME2, together with its previous version, has accumulated approximately 29,000 citations, reflecting its prominence for marker-gene analysis finding
  • Metagenomics taxonomic classifiers were developed as faster alternatives to BLAST-based comparison against GenBank, trading some sensitivity for speed method
Experimental setups
Assay System Perturbation Readout Platform
Pipeline benchmarking (marker-gene metagenomics via QIIME2; metatranscriptomics via Kraken2/Bracken, implemented in Nextflow) simulated and experimental sequencing datasets none pipeline performance Nextflow
Key results
  • 16S rRNA-based methods fail to detect more than 50% of species within the phylum Radiation, which represents 15% of the entire bacterial domain 50%
  • Approximately 99% of genes in the human tissue gene pool are derived from microorganisms rather than the host 99%
  • QIIME2 and its predecessor have together accumulated around 29,000 citations 29,000 citations
  • More than 4300 articles on gut microbiota were published in the last 5 years according to PubMed 4300 articles
Key statistics
  • count ~40 trillion eukaryotic cells (estimated number of human body cells)
  • count ~22,000 genes (number of genes in human genome)
  • count ~100 trillion microbial cells (estimated size of human microbiota)
  • count ~2 million genes (number of genes in human microbiome)
  • other 99% (proportion of human tissue-pool genes derived from microorganisms)
  • count >4300 articles (PubMed articles on gut microbiota in the last 5 years)
  • count ~29,000 citations (combined citations of QIIME2 and its previous version)
  • other >50% of species undetected; phylum Radiation = 15% of bacterial domain (limitation of 16S rRNA sequencing at species level)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a narrative review and methods-comparison article on metagenomics and metatranscriptomics tools, rather than a study reporting inferential statistical hypothesis tests. The provided text describes the biological background, sequencing technologies, and bioinformatic tools (e.g., QIIME2, Kraken2/Bracken, assemblers, taxonomic classifiers) and states that the authors developed two Nextflow pipelines evaluated on simulated and experimental datasets, but the excerpt does not describe a formal statistical testing framework (e.g., hypothesis tests, p-values, or effect sizes) for that evaluation.

Replicationunclear GroupsPerformance/behavior of different bioinformatic tools and pipelines (e.g., QIIME2 vs. Kraken2/Bracken) on simulated and experimental datasets, as described narratively Pairingna Randomization/blindingnot stated Dispersionnone
Approaches that could also have been used
  • The article describes tool/pipeline performance in narrative terms (e.g., speed, sensitivity, accuracy trade-offs) based on simulated and experimental datasets without reporting formal quantitative comparison statistics in this excerpt
    Could also: A benchmarking framework with paired quantitative metrics (e.g., precision/recall/F1 per tool, computed on the same simulated datasets) summarized with dispersion measures across replicate simulations — Reporting quantitative benchmark metrics with variability across simulated replicates would allow readers to gauge how consistently one tool outperforms another rather than relying on qualitative description
  • Multiple bioinformatic tools/pipelines are compared narratively across categories (assembly, taxonomic classification, pre-processing)
    Could also: A formal multi-tool comparison design (e.g., repeated-measures ANOVA or Friedman test across tools applied to the same benchmark datasets) with post-hoc correction for multiple pairwise comparisons — Such a design would let the family-wise error rate be controlled when many tools are compared simultaneously on shared benchmark data
  • Pipeline evaluation is described as using both simulated and experimental datasets, but the sample sizes/number of datasets used are not stated in this excerpt
    Could also: Explicit power/sample-size justification or a stated number of simulated replicates and experimental samples used for evaluation — Stating the number of datasets or replicates used in benchmarking would let readers assess the precision of any reported performance differences
Software: QIIME 2 · Kraken2/Bracken · Nextflow · FastQC · MultiQC · cutadapt

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36553546

Paper: Terrón-Camero LC, Gordillo-González F, Salas-Espejo E, Andrés-León E. "Comparison of Metagenomics and Metatranscriptomics Tools: A Guide to Making the Right Choice." Genes (Basel) 2022;13(12):2280. PMID 36553546 · PMC9777648 · DOI 10.3390/genes13122280.

Code: https://github.com/BioinfoIPBLN/16S-Metatranscriptomic-Analysis (own repo of the authors' group, IPBLN). Two Nextflow (DSL1) pipelines:

  • Qiime-pipeline/ — QIIME2 16S amplicon: import → DADA2 denoise → SILVA classify-sklearn → collapse to taxonomic level → export count matrix (+ phylogeny, α/β diversity, optional metagenomeSeq differential abundance).
  • Kraken-bracken-pipeline/ — Kraken2 + Bracken classification (+ Krona), for shotgun / metatranscriptomic reads (host already removed upstream).

Type of paper: a TOOL-COMPARISON / benchmark ("a guide to making the right choice") that runs QIIME2 vs Kraken2/Bracken on (a) simulated mocks, (b) synthetic gut mocks and (c) one real case-study dataset, reporting detection accuracy (true/false positives), correlations and runtimes per tool/database.


Datasets the paper relies on

ref what accession obtainable?
Li et al. (case study) 16S rRNA amplicon, endometrial tissue PRJNA750303 (given) YES — 40 AMPLICON runs SRR15276323–SRR15276362
Li et al. (case study) RNA-Seq metatranscriptome, 60 paired NOT in PRJNA750303 (0 RNA-Seq runs there) accession not provided → out of reach
Mende et al. 2012 simulated 10/100/400-species shotgun mocks not given in brief external, not provided
Almeida et al. 2018 synthetic gut 16S mocks A100/A500 not given in brief external, not provided

The brief pins only PRJNA750303. ENA confirms it contains exactly 40 paired-end AMPLICON (16S, genomic) runs (~78k–82k read pairs each), and zero RNA-Seq/transcriptomic runs. So the metatranscriptomic (RNA-Seq) half of the case study cannot be reproduced from the provided accession.


IN SCOPE (pipeline-derived, attempted)

QIIME2 16S pipeline on PRJNA750303 (the paper's own case-study DNA data), counting detected taxa per group (EC tumor vs HC healthy), matching the repo's Qiime-pipeline exactly:

  • C1 QIIME 16S genera detected — tumor (EC): reported 640
  • C2 QIIME 16S genera detected — healthy (HC): reported 408
  • C3 QIIME 16S species detected — tumor (EC): reported 214
  • C4 QIIME 16S species detected — healthy (HC): reported 114
  • C5 dataset N: 40 16S samples, avg 80,360 paired-reads, 250 nt (Results; used as a dataset-profiling cross-check, computed from ENA read_count).

Group labels come from the SRA sample aliases: N1–N10 = healthy (10), C1–C35 = cancer/tumor (30) — matches the paper's 30 EC / 10 HC split.

Pipeline parameters (from repo Qiime-pipeline/main.nf + nextflow.config):

  • qiime tools import CasavaOneEightSingleLanePerSampleDirFmt, PairedEnd, demultiplexed
  • qiime dada2 denoise-paired — defaults: trim-left-f/r = 0, trunc-len-f/r = 0, --p-chimera-method consensus
  • qiime feature-classifier classify-sklearn against a SILVA classifier.qza (paper text: SILVA132; we use a standard pre-trained SILVA Naive-Bayes classifier and DOCUMENT the version difference)
  • qiime taxa collapse --p-level 6 (genus) and --p-level 7 (species), export biom→tsv; count rows with non-zero abundance per group.

OUT OF SCOPE / NOT ATTEMPTED (reason)

  • RNA-Seq / metatranscriptomic Kraken/RefSeq case-study results (215/253 genera, 224/298 species; 24.2M→1.03M reads): RNA-Seq data is NOT in PRJNA750303 and no other accession is provided → data_unavailable for that half.
  • Mende simulated 10/100/400-species results (corr 0.97/0.73/0.86; FP rates): simulated read sets not provided in the brief; external.
  • Almeida A100/A500 synthetic gut results (QIIME 88 genera, corr 0.67; Kraken/SILVA 0 species; etc.): mock read sets not provide
C5
Reported
40 16S samples (30 EC, 10 HC)
Reproduced
40 amplicon runs in PRJNA750303 (30 C, 10 N)
exact
C5b
Reported
avg 80,360 paired-reads/sample
Reproduced
mean 80,360 (DADA2 input) / 80,359.7 (ENA)
exact
C5c
Reported
250 nt reads
Reproduced
250 nt
exact
C1
Reported
640 genera (tumor/EC, QIIME 16S)
Reproduced
799 named (1055 all L6); ratio 1.25x
partial
C2
Reported
408 genera (healthy/HC, QIIME 16S)
Reproduced
483 named (637 all L6); ratio 1.18x
partial
C1s
Reported
340 genera shared EC&HC
Reproduced
408 named (542 all); ratio 1.20x
partial
C3
Reported
214 species (tumor/EC, QIIME 16S)
Reproduced
569 named (1780 all L7); ratio 2.66x
partial
C4
Reported
114 species (healthy/HC, QIIME 16S)
Reproduced
274 named (975 all L7); ratio 2.40x
partial
C3s
Reported
72 species shared EC&HC
Reproduced
158 named (733 all); ratio 2.19x
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 67/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

Dataset identity reproduces exactly (40 samples, mean 80,360 reads, 250 nt) and the central qualitative conclusion is fully confirmed — EC>HC on every metric, genera>species, and matching EC/HC ratios (genera 1.65 vs 1.57; species 2.08 vs 1.88). The quantitative gap (genera ~1.2x, species ~2.4-2.7x high) sits on the input/reference side: SILVA 138 vs the paper's SILVA 132, the authors' exact classifier.qza not deposited, version drift, and under-specified count filtering — a mix of our method choices and authors' under-specification, not fabrication. The RNA-Seq metatranscriptomic half and external mocks are correctly out of scope (data not deposited), so this is an honest, well-documented partial reproduction.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

250.5 k
tokens (I/O) · 20.2 M incl. cache
70 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.