Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

MetaMap: an atlas of metatranscriptomic reads in human disease-related RNA-seq data.

Gigascience · 2018
L1 95/100 PQI 96
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score 0
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
95/100
Reproducibility score
1.2 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 89% of all assessed papers rank 105 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the DOWNSTREAM results 1:1. The theislab/MetaMap repo is not a runnable pipeline; it ships the finished composite metafeature OTU count table (MetaMapData.RData, 3436 metafeatures x 17,278 libraries, sha256 0a62182...6e9e) plus a deterministic R tutorial. We re-ran the shipped tutorial and recomputed the paper's downstream claims on «our HPC» (R 4.5.2, DESeq2 1.50.2). RESULTS: HPV is the #1 differentially-abundant metafeature in study SRP066090 (padj 8.5e-20) -- exact match to Fig2/tutorial; SRP041338 (this RU's named accession) phiX+EBV read fraction = 94.1% vs reported ~95% -- within-tol; SRP091453 = 97.6% vs ~97% -- exact; relative classified bacteria:virus:archaea composition 32:64:4% matches Fig1's 32:63:5%; atlas dimensions 436 projects / 17,278 libraries match exactly. NOT ATTEMPTED (out of scope, the irreproducible 99%): the upstream LRZ-scale read pipeline -- STAR hg38 alignment of ~150TB / 17,278 libraries (the 90.7% human claim), CLARK-S 16,551-genome DB build + classification of >500 billion reads, BLASTN E<=1e-10 validation, and the throughput numbers (hardware-dependent). The three infection validation studies other than HPV (Salmonella P<1e-75, HSV-1, Rhinovirus A) were not attempted because their SRP accessions are not stated in the paper (only HPV's SRP066090 is pinned via the shipped Rmd). No fabrication signal: every checked value is derivable from the shipped data and matches. status=partial because the heavy upstream pipeline is unreproducible by construction, while all shipped-data-derived results reproduced.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 95
    assessed: 2026-06-16 ⛓ 509b7ba59192
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The authors hypothesize that reads in archived human RNA-seq datasets that fail to map to the human reference genome contain valuable information about the presence of microbes and viruses in body niches and/or under defined disease conditions, which can be systematically mined to correlate microbial/viral detection with human diseases.

Core claims
  • A two-step 'omni' RNA-seq pipeline (MetaMap) combining STAR human alignment with CLARK-S metagenomic classification can quantify archaeal, bacterial, and viral reads from the non-human read fraction of human RNA-seq data method
  • MetaMap recapitulates ground-truth infection agents from independent controlled dual RNA-seq experiments, detecting the correct pathogen as the most differentially abundant metafeature finding
  • The MetaMap database, derived from re-analysis of >17,000 samples across >400 disease-related studies, provides a resource for hypothesis generation about the microbiome's role in human disease resource
  • MetaMap abundance estimates correlate with an alternative BLAST-based classification while running more than three orders of magnitude faster finding
  • MetaMap minimizes false positives, providing abundance and significance measures that allow users to identify and counterselect contaminant/spike-in species (e.g., phiX, EBV) finding
  • Intraproject comparisons testing one metafeature at a time are recommended to avoid technical confounders (tissue sterility, sequencing depth, RNA-vs-DNA abundance) method
Experimental setups
Assay System Perturbation Readout Platform
Bulk RNA-seq re-analysis (computational metatranscriptomic pipeline) Human primary/clinical samples from >400 SRA studies (>17,000 cDNA libraries) none OTU/metafeature count matrix (reads per million classified to archaeal, bacterial, viral species) plus human gene expression counts STAR v2.5.2 (hg38) + CLARK-S v1.2.3; Illumina source data; LRZ HPC (CoolMUC2, Teramem)
Dual RNA-seq validation (differential metafeature abundance) HeLa cells Salmonella enterica serovar Typhimurium infection vs mock-treated controls (Westermann et al.) Differential metafeature abundance between infected and control samples DESeq2
Dual RNA-seq validation Human tumor/cell samples (Zhang et al.) Human papillomavirus-positive vs HPV-negative Most differentially abundant metafeature
Dual RNA-seq validation Human cells (Rutkowski et al.) Herpes simplex virus infection vs control Most differentially abundant metafeature
Dual RNA-seq validation Human cells (Bai et al.) Rhinovirus infection vs control Most differentially abundant metafeature
Negative/specificity control re-analysis B lymphoblast cell line (projects SRP041338, SRP091453) none (EBV-transformed, noninfectious) Proportion of reads classified to EBV, phiX, and bacterial metafeatures
BLAST-based metatranscriptomic classification (technical comparison) Westermann et al. study (42 samples) none Average metafeature abundance and classification speed vs CLARK-S/MetaMap BLASTN (E-value 1e-10); LRZ CoolMUC3
Key results
  • 90.7% of all RNA-seq reads mapped to the human genome; 0.03% archaeal, 0.20% bacterial, 0.39% viral metafeatures, and 8.6% unclassified 90.7%/0.03%/0.20%/0.39%/8.6%
  • S. enterica was the most differentially abundant metafeature between Salmonella-infected and control HeLa samples (Westermann et al.) P<1e-75
  • Correct infection agents recovered as top differential metafeatures in all four validation studies (alphapapillomavirus 9, human alphaherpesvirus 1, rhinovirus A)
  • Related off-target species also enriched: Salmonella bongori (P<1e-67) and Panine alphaherpesvirus 3 (P<1e-9) P<1e-67; P<1e-9
  • CLARK-S/MetaMap and BLAST average metafeature abundances correlated significantly across 42 samples Spearman Rho=0.16, P=3.1e-10
  • MetaMap processed reads more than three orders of magnitude faster than BLAST >1000-fold
  • In lymphoblast control projects, on average 95% (SRP041338) and 97% (SRP091453) of metafeature reads were phiX or EBV, with bacterial reads in the bottom percentile 95%; 97%
  • Pipeline processing throughput reached median 25 and 21 million reads per hour per core per run for STAR and CLARK-S steps respectively 25 and 21 million reads/hour/core
Key statistics
  • pvalue P<1e-75 (S. enterica differential abundance, infected vs control, Westermann et al.)
  • correlation Spearman Rho=0.16, P=3.1e-10 (Correlation of average metafeature abundance between CLARK-S and BLAST)
  • pvalue P<1e-67 (Salmonella bongori differential abundance, Westermann et al.)
  • pvalue P<1e-9 (Panine alphaherpesvirus 3 differential abundance, Rutkowski et al.)
  • count >17,000 samples / >500 billion reads / ~150 terabytes / >400 studies (Scale of public RNA-seq data re-analyzed by MetaMap)
  • count 484 SRPs containing 21,659 RNA-seq runs (Result of SRA query before download)
  • count 16,551 genome sequences corresponding to 6,979 unique species (CLARK-S metagenomic reference database content)
  • other 95% and 97% (Average proportion of metafeature reads classified as phiX or EBV in lymphoblast projects SRP041338 and SRP091453)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

MetaMap is a bioinformatics pipeline paper describing large-scale re-analysis of human RNA-seq data for microbial and viral reads using a two-step approach (STAR alignment then CLARK-S k-mer classification). Validation was performed by applying DESeq2 differential abundance analysis to four controlled infection datasets, testing whether the known pathogen ranked as the most significantly differentially abundant metafeature between infected and control samples. Agreement between the CLARK-S pipeline and an alternative BLAST-based approach was assessed with Spearman rank correlation; results were reported primarily as volcano plots, box plots (median and IQR), and ranked P-value thresholds.

Replicationunclear Sample sizeSample sizes per validation study are not stated in this paper; the paper reanalyzes existing published datasets using their provided annotations; 42 samples noted for one study in the method-comparison context; >17,000 samples from >400 studies described for the full MetaMap resource GroupsInfected vs. mock/control within each of four validation studies; lymphoblast cell line projects assessed as negative controls against the full project collection Pairingunpaired Randomization/blindingnot stated DispersionIQR Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionNot explicitly stated in the text; DESeq2 applies Benjamini-Hochberg FDR adjustment internally by default, but the paper does not discuss adjusted vs. unadjusted P-values
Statistical tests used
Test Applied to n Assumptions
DESeq2 negative binomial Wald test (differential abundance) Comparison of metafeature abundance between infected and control/mock samples across four dual RNA-seq validation studies (Westermann/Salmonella, Zhang/HPV, Rutkowski/HSV, Bai/rhinovirus); applied across all detected metafeatures within each study Not explicitly stated per study for the DESeq2 models; 42 samples noted for the Westermann study in the context of the BLAST validation not stated
Spearman rank correlation Comparison of average metafeature abundances (reads per million) estimated by CLARK-S vs. BLAST across samples from the Westermann study (Fig. 4A; rho = 0.16, P = 3.1e-10) 42 samples (Westermann study, stated in text) not stated
Approaches that could also have been used
  • Agreement between CLARK-S and BLAST metafeature abundance estimates was quantified with Spearman rank correlation (rho = 0.16)
    Could also: Bland-Altman analysis or the concordance correlation coefficient (CCC) could also be used to assess method agreement — Spearman correlation measures monotonic association but does not distinguish systematic bias from random error; Bland-Altman plots and CCC are specifically designed for method-comparison contexts and can reveal proportional bias or limits of agreement across the abundance range
  • Differential metafeature abundance was tested using DESeq2 with a negative binomial model fitted to raw counts
    Could also: edgeR (quasi-likelihood F-test or exact test) or limma-voom could also be applied to RNA-seq-derived count data for differential abundance testing — edgeR and limma-voom are well-validated alternatives for overdispersed count data; comparing results across methods is a common sensitivity check in RNA-seq analysis and can indicate whether findings are method-dependent
  • Metafeature abundance was normalized as reads per million (RPM) total reads for cross-sample display and for estimating library size factors in DESeq2
    Could also: Trimmed mean of M-values (TMM) normalization or DESeq2 median-of-ratios size factors applied to the metafeature count matrix directly could also be used — TMM and median-of-ratios account for compositional variation in the detected metafeature profile itself, rather than scaling to total (mostly human) reads; this distinction can matter when the non-human fraction varies substantially across samples
  • Multiple testing across all detected metafeatures within each study was handled implicitly by DESeq2, with P-values reported as thresholds without explicit discussion of adjustment
    Could also: Explicitly reporting Benjamini-Hochberg adjusted P-values (padj, which DESeq2 computes by default) alongside or instead of raw P-values could also be done — Explicitly reporting adjusted P-values makes the multiple-testing correction transparent to readers, particularly given that the number of metafeatures tested simultaneously is not stated and could be large
  • Top-hit metafeature abundance across conditions was displayed with box plots showing median and IQR
    Could also: Individual data points overlaid on box plots (strip charts or beeswarm plots) could also be used, especially given that sample sizes per group are small in at least some validation studies — Overlaying individual observations is commonly recommended when n is small, as it preserves information about sample size and distribution shape that summary boxes alone can obscure
  • Validation success was defined as the known pathogen being the single top-ranked (most significant) metafeature in each study
    Could also: A receiver operating characteristic (ROC) curve or precision-recall analysis across the full ranked metafeature list could also quantify detection performance — ROC and precision-recall frameworks provide a continuous summary of sensitivity and specificity across all possible ranking thresholds, rather than evaluating only rank-1 performance, which can give a more complete picture of pipeline discrimination
Software: DESeq2 (R package) · STAR 2.5.2 · CLARK-S 1.2.3 · BLASTN · SRAdb (R package)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
27
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

10.5524/100456 DOI in References (http://purl.org/orb/References)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-29901703 (MetaMap)

Paper: Simon et al. MetaMap: an atlas of metatranscriptomic reads in human disease-related RNA-seq data. GigaScience 2018, 7(6):giy070. Repo: https://github.com/theislab/MetaMap (HEAD 8b961b9, pushed 2018-11-28, not archived, GPL). Named data accession (this RU): SRA SRP041338.

What the pipeline is

Two-stage subtractive metatranscriptomics applied to ~150 TB of public human RNA-seq:

  1. STAR v2.5.2 align to hg38, --quantMode GeneCounts --outReadsUnmapped Fastx → human counts + unmapped reads.
  2. CLARK-S v1.2.3 classify unmapped reads against a custom 16,551-genome / 6,979-species bacteria+virus DB (--spaced -m 2 -n 32) → composite metafeature OTU count table.
  3. BLASTN (E ≤ 1e-10) orthogonal validation. DESeq2 differential abundance per case study. Final atlas: 436 SRA projects / 17,278 cDNA libraries.

Key reproducibility fact

The repo is not a runnable pipeline (no Snakemake/Nextflow/shell; pipeline lives on protocols.io + LRZ-scale compute). It ships the finished result object data/MetaMapData.RData (≈11.4 MB): meta_count (metafeature × sample sparse counts), sample_info (study, sample_attribute, spots), meta_info (TaxID, Species). It also ships Tutorial_statistical_analysis.Rmd + a knitted .html reference output. → The paper's downstream, pipeline-derived results are deterministically recomputable from the shipped count table, without re-running any alignment/classification.

IN SCOPE (attempted — deterministic, low-compute, from shipped MetaMapData.RData)

  • C1 — Tutorial / Fig 2 (HPV): DESeq2 differential abundance for study SRP066090 (positive vs negative), as coded verbatim in the shipped Rmd. Expected: Human Alphapapillomavirus 9 (TaxID 337041) is the top/most-significant differentially-abundant metafeature; regenerate volcano plot + HPV abundance boxplot. Reference = committed .html.
  • C2 — SRP041338 (THIS RU's accession) / Fig 3C: "On average, 95% of all metafeature reads were classified as phiX or EBV" for SRP041338. Directly computable: per-sample fraction of metafeature reads belonging to phiX + EBV metafeatures, averaged.
  • C3 — SRP091453 / Fig 3C: same, expected 97%.
  • C4 (opportunistic) — global metafeature composition (Fig 1): relative bacteria : virus : archaea proportions among classified reads, if meta_info carries a superkingdom/domain column. NOTE: the absolute Fig-1 percentages (90.7% human, 8.6% unclassified, 0.20% bacterial, 0.39% viral, 0.03% archaeal) include the human + unclassified read totals, which are NOT in the shipped object → only the relative classified composition is checkable here.

OUT OF SCOPE (not attempted — the heavy/irreproducible pipeline)

  • STAR alignment of 17,278 libraries to hg38 (≈150 TB, LRZ-scale): the 90.7% human claim.
  • CLARK-S DB build (16,551 genomes) + classification of >500 billion reads.
  • BLASTN validation pass.
  • Throughput numbers (25 / 21 M reads/hr/core) — hardware-dependent.
  • The four infection validation studies' DESeq2 top-hits (Salmonella P<1e-75, HSV-1, Rhinovirus A): their SRP accessions are not stated in the paper; only HPV (SRP066090) is pinned via the Rmd, so only C1 is attempted from the validation set.

Environment

«our HPC» SLURM, conda prefix env salmon-deseq2 (R 4.5.2, DESeq2 1.50.2, Matrix 1.7.5, ggplot2 4.0.3; ggrepel optional). Repo cloned inside the compute job at pinned commit.

Figures / tables: Fig 2Fig 3CFig 1
C1
Reported
HPV Alphapapillomavirus 9 (TaxID 337041) dominant in SRP066090 (Fig2/tutorial)
Reproduced
rank #1 by padj, padj=8.53e-20, log2FC=11.64, CPM pos 101 vs neg 0.07
exact
C2
Reported
~95% phiX+EBV reads in SRP041338 (this RU accession, Fig3C)
Reproduced
mean 94.12% / median 94.91% / pooled 94.86% (n=17)
within tolerance
C3
Reported
~97% phiX+EBV reads in SRP091453 (Fig3C)
Reproduced
mean 97.56% / pooled 97.79% (n=6)
exact
C4
Reported
Fig1 relative classified composition bact:viral:arch 0.20:0.39:0.03 (~32:63:5%)
Reproduced
Bacteria 32.06% : Viruses 63.81% : Archaea 4.13%
within tolerance
S1
Reported
436 SRA projects
Reproduced
436 unique studies
exact
S2
Reported
17,278 cDNA libraries
Reproduced
17,278 sample columns
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 95/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score 0

Every result that can be deterministically derived from the shipped composite OTU table reproduced cleanly: HPV is the #1 hit in SRP066090 (padj 8.53e-20), contamination fractions (94.12%/97.56% vs ~95%/~97%), relative classified composition (32:64:4 vs 32:63:5), and atlas dimensions (436/17,278) all match with only rounding-level deviations and no fabrication signal. The limitation is on the data-availability/scope side, not the authors' integrity: the repo ships the pipeline output rather than a runnable pipeline, so the upstream ~150 TB read-processing (90.7% human, throughput, 3 of 4 infection studies) was unreproducible-by-construction or lacked stated accessions. Overall a strong-but-partial reproduction → q8 yellow, criticality yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

115.8 k
tokens (I/O) · 6.1 M incl. cache
16 min
runtime · 0 CPU-h
0.9 GB
peak RAM
2
HPC jobs
hummel
machine