Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Improved eukaryotic detection compatible with large-scale automated analysis of metagenomes.

Microbiome · 2023
L1 89/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
89/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 77% of all assessed papers rank 246 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED. CORRAL (P16 third-party Nextflow tool, commit 2c990f6, default params verified vs paper Methods) run on all 35 ZymoBIOMICS D6300 mock samples of PRJEB38036, with the EukDetect v1 marker DB (521824 markers). C1 EXACT: both fungi detected in 35/35, S. cerevisiae unambiguous and C. neoformans flagged ambiguous, exactly as Fig 4C. C2 within-tol: S. cerevisiae 182 markers reproduces EXACTLY; deepest samples give Sc ~26000-31000 / Cn ~15000-18000 reads matching the paper's ZYMO ~27000/~15000; C. neoformans markers 170 vs paper 150. C3 within-tol: false positives mean 30 reads / <8 markers, all flagged ambiguous (paper ~44/<8). C4 within-tol: CORRAL Sc:Cn ratio mean 1.43 (median 1.38) within the paper's 1-2 and never the >=10x EukDetect gives - reproducing the central CORRAL-vs-EukDetect contrast. NOT attempted: simulated-data benchmarks (need authors' simulation code), DIABIMMUNE, MicrobiomeDB deployment (all out of scope). Small quantitative gaps (Cn 170-vs-150 markers, FP 30-vs-44 reads) attributable to EukDetect-DB-snapshot / marker_alignments version drift, NOT fabrication - every value is derivable from the open data + published tool.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ 58c2a03c3e49
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether a Markov-clustering-based approach to processing eukaryotic marker gene alignments (CORRAL) can achieve more sensitive and accurate detection of microbial eukaryotes in shotgun metagenomic data than existing MAPQ-filter-based methods (EukDetect), including detection of species not represented in the reference database.

Core claims
  • MAPQ ≥30 filtering improves precision but substantially reduces recall, especially for unrepresented/divergent eukaryotic taxa finding
  • CORRAL uses Bowtie2 alignment followed by Markov clustering (MCL) of marker genes and taxa to detect eukaryotes, including species not present in the reference, without relying on MAPQ filtering method
  • CORRAL achieves sensitivity and specificity comparable to or better than EukDetect while retaining ability to detect unrepresented species finding
  • CORRAL can detect low-abundance taxa with as few as two reads aligning to two different markers, a lower limit of detection than EukDetect finding
  • CORRAL reports ambiguous/uncertain calls for taxa that do not perfectly match the reference, unlike EukDetect finding
  • In a ZymoBIOMICS mock community, CORRAL produced more accurate relative abundance estimates between S. cerevisiae and C. neoformans than EukDetect, which was skewed toward S. cerevisiae finding
  • CORRAL has been deployed on MicrobiomeDB.org to create an automated, running atlas of eukaryotes across human microbiome studies resource
  • The reference-alignment clustering approach may generalize to other non-exhaustive reference matching problems, such as bacterial virulence gene or viral read classification mechanism
Experimental setups
Assay System Perturbation Readout Platform
in silico read simulation and remapping EukDetect marker gene reference database (3977 taxa) MAPQ ≥30 filter vs no filter precision and recall of taxon assignment Bowtie2
holdout simulation (unrepresented species) EukDetect reference, 371 holdout taxa vs 3343 remaining taxa species removed from reference before mapping precision/recall, same-genus assignment rate Bowtie2
mutation-rate simulation EukDetect marker gene reference reads simulated read mutation rate 0-0.2 recall and precision vs mutation rate Bowtie2
low-abundance detection simulation 338 simulated single-taxon samples at minimal read abundance tool comparison (Bowtie2 only, EukDetect default/sensitive, CORRAL) sensitivity and specificity Bowtie2/EukDetect/CORRAL
unrepresented species detection simulation 338 simulated samples, single holdout taxon at 0.1x genome coverage tool comparison (EukDetect vs CORRAL) number of samples with results, single-species calls, same-genus calls, ambiguous calls CORRAL/EukDetect
shotgun metagenomic sequencing analysis ZymoBIOMICS mock community standard (S. cerevisiae, C. neoformans, 8 bacterial species) 6 DNA extraction methods (MP, MN, ZYMO, Q, PS, MetaHIT) copies per million (CPM), relative abundance ratio of S. cerevisiae to C. neoformans CORRAL/EukDetect
pairwise confusability analysis / PCA simulated reads across 4558 taxa pairs from EukDetect reference none read emission/acceptance rates between taxon pairs; PC1/PC2 variance explained
Key results
  • Exact-match simulation: baseline recall and precision both 95.1%; MAPQ≥30 filter raised precision to 99.7% but dropped recall to 91.7% 95.1% to 99.7% precision; 95.1% to 91.7% recall
  • Holdout (unrepresented species) simulation: without filter, same-genus precision 82%/recall 30%; with MAPQ filter, precision 83.6%/recall dropped to 7% recall 30% to 7%
  • Mutation-rate simulation: recall declined from 95.1% to <10% at mutation rate 0.2; at rate 0.1, MAPQ filter dropped recall from 68.3% to 5.0% 68.3% to 5.0%
  • Low-abundance detection: Bowtie2-only sensitivity 100%/specificity 93.4%; EukDetect default sensitivity 72.5%/specificity 100%; EukDetect sensitive and CORRAL both >95% sensitivity and specificity >95% sensitivity and specificity
  • Unrepresented species simulation: CORRAL returned results for 205/338 samples, single-species calls for 164/338, same-genus calls for 136/338, ambiguous calls for 139/205 205/338 samples with results
  • ZymoBIOMICS mock community: CORRAL detected both fungal species with abundance ratio between 1 and 2 (near theoretical 1:1); EukDetect overestimated S. cerevisiae abundance by at least an order of magnitude regardless of extraction method ratio 1-2 (CORRAL) vs ≥10-fold skew (EukDetect)
  • PCA of taxon-pair confusability: PC1 explained 61.6% of variance (overall confusability), PC2 explained 21.5% (bias toward identifying one member of a pair) 61.6% and 21.5% variance explained
Key statistics
  • other recall 95.1%, precision 95.1% (baseline exact-match simulation, no filter)
  • other precision 99.7%, recall 91.7% (exact-match simulation with MAPQ≥30 filter)
  • count 3977 taxa mapped; 1908 taxa at 100% precision pre-filter; 1105 additional taxa reach 100% precision post-filter; 146 taxa below 95.1% precision post-filter (species-level MAPQ filtering impact)
  • other same-genus precision 82%/recall 30% (no filter) vs precision 83.6%/recall 7% (MAPQ filter) (holdout unrepresented-species simulation)
  • count 205/338 samples with results; 164/338 single-species; 136/338 same-genus; 139/205 flagged ambiguous (CORRAL performance on unrepresented species simulation)
  • fold_change at least an order of magnitude higher S. cerevisiae abundance estimate (EukDetect abundance bias in ZymoBIOMICS mock community vs CORRAL ratio of 1-2)
  • mean ~27,000 reads / 182 markers (S. cerevisiae) and ~15,000 reads / 150 markers (C. neoformans) vs ~44 reads / <8 markers for false positives (ZYMO-extracted replicate samples, read/marker counts by taxon)
  • other PC1 = 61.6% variance, PC2 = 21.5% variance (PCA of pairwise taxon confusability features)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is primarily a bioinformatics software/benchmarking paper describing CORRAL, a tool for eukaryotic detection in metagenomic data. Performance was assessed using simulated read datasets, a mock community standard (6 replicates per DNA extraction method), and public metagenomic data, with results reported mainly as descriptive percentages (sensitivity, specificity, precision, recall) and visualized via scatter plots, box plots, and a heatmap. A principal component analysis (PCA) was used to characterize pairwise taxon 'confusability.' No formal inferential hypothesis tests (e.g., t-tests, ANOVA) or p-values were described in the presented text.

Replicationbiological Sample sizeDescribed per analysis: e.g., 338 simulated single-taxon samples for lower-limit-of-detection and unrepresented-species tests; 6 replicate extractions per method (six extraction protocols) for the ZymoBIOMICS mock community; 4558 taxon pairs for the PCA/confusability analysis GroupsCORRAL vs. EukDetect (default/sensitive settings) vs. unfiltered Bowtie2 mapping; comparisons across six DNA extraction methods (MP, MN, ZYMO, Q, PS, MetaHIT) Pairingunclear Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Principal component analysis (PCA) Analysis of pairwise taxon confusability (emit/accept read rates between taxon pairs), Fig. S1C 4558 pairs of taxa, 8 rate computations each not stated
Approaches that could also have been used
  • Sensitivity, specificity, precision, and recall for CORRAL and EukDetect were reported as single point-estimate percentages across sets of simulated samples (e.g., 338 samples).
    Could also: Bootstrap or exact binomial confidence intervals around these proportions — Since these are proportions estimated from a finite number of simulated samples, confidence intervals would convey the precision of the sensitivity/specificity estimates and make it easier to judge how much sampling variability might contribute to differences between tools.
  • Differences in performance between CORRAL, EukDetect (default), and EukDetect (sensitive) were described narratively based on the reported percentages, without a formal statistical comparison.
    Could also: A paired test for proportions across the same simulated samples, such as McNemar's test, or a mixed-effects logistic model with tool as a fixed effect — Because the same simulated samples were evaluated by each tool, a paired approach would account for the shared sample structure and could quantify how much of the observed difference in detection performance exceeds what might be expected by chance.
  • Read-count ratios between S. cerevisiae and C. neoformans across six DNA extraction methods were shown as box plots (Fig. 4D) without an accompanying statistical test.
    Could also: A Kruskal-Wallis test (or one-way ANOVA, if normality holds) across extraction methods, with post-hoc pairwise comparisons — With six replicates per method already collected, a formal test would let readers assess whether extraction-method differences in the estimated abundance ratio are larger than would be expected from replicate-to-replicate variation alone.
  • PCA was used to summarize an 8-feature representation of pairwise taxon confusability, reporting percent variance explained by PC1 and PC2.
    Could also: A permutation test (e.g., parallel analysis) to assess how many PCs explain more variance than expected by chance, alongside the reported variance percentages — This would provide an additional check on how many components meaningfully capture structure in the confusability data versus noise, complementing the raw variance-explained figures already reported.
  • Comparisons among CORRAL, EukDetect default, and EukDetect sensitive settings on the unrepresented-species simulation (Fig. 4B) were reported as counts/fractions of samples with results, single-species calls, and same-genus calls.
    Could also: A chi-square or Fisher's exact test comparing the count distributions across tools — This would offer a formal way to characterize whether the differences in categorical outcome counts (no result / single species / same genus) across tools are larger than expected under random variation.
Software: Bowtie2 · Markov Clustering (MCL) · Nextflow · Python (CORRAL module)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 37032329 (CORRAL)

Paper: Bazant et al. 2023, Microbiome 11:84. "Improved eukaryotic detection compatible with large-scale automated analysis of metagenomes." DOI 10.1186/s40168-023-01505-1 · PMCID PMC10084625.

Tool (P16 third-party-but-equally-valid): CORRAL — a Nextflow pipeline wrapping the marker_alignments Python module. Repo: https://github.com/wbazant/CORRAL pinned at commit 2c990f6030eedd9f6b0ce90362161e3f26ebf030 (main, 2023-10-14). CORRAL = Bowtie2 alignment of metagenomic reads to EukDetect universal marker genes, followed by marker_alignments filtering + correlation/MCL-style clustering to call taxa and flag ambiguity. License MIT.

Exact default pipeline command (from nextflow.config at pinned commit):

  • bowtie2Command = "bowtie2 --omit-sec-seq --no-discordant --no-unal -k10"
  • summarizeAlignmentsCommand = "marker_alignments --min-read-query-length 60 --min-taxon-num-markers 2 --min-taxon-num-reads 2 --min-taxon-better-marker-cluster-averages-ratio 1.01 --threshold-avg-match-identity-to-call-known-taxon 0.97 --threshold-num-taxa-to-call-unknown-taxon 1 --threshold-num-markers-to-call-unknown-taxon 4 --threshold-num-reads-to-call-unknown-taxon 8"
  • summaryColumn = "cpm" These defaults match the thresholds described in the paper Methods (≥60 nt read length, ≥97% identity to call a known taxon, ≥4 markers / ≥8 reads to call an unknown taxon).

In scope (pipeline-derived, reproducible with the paper's OWN data)

The paper's mock-community validation uses PRJEB38036 — "Assessment of fecal DNA extraction protocols for metagenomic study" — specifically the 36 ZymoBIOMICS D6300 mock-community (MMC) extractions (6 DNA-extraction methods × 6 replicates). The D6300 standard contains 2 fungal strains: Saccharomyces cerevisiae and Cryptococcus neoformans. This is the directly reproducible target: run CORRAL (default params) on the 36 D6300 runs and compare the eukaryotic calls.

id claim (reported) location pipeline reproducible here
C1 CORRAL detects both expected fungi S. cerevisiae (unambiguous) and C. neoformans (flagged ambiguous) in the MMC Fig 4C CORRAL default YES — run on 36 D6300 samples
C2 In ZYMO-method MMC samples, S. cerevisiae ≈27,000 reads / 182 markers; C. neoformans ≈15,000 reads / 150 markers Table S1 CORRAL default YES (per-method subset)
C3 False-positive hits average ≈44 reads / <8 markers, flagged ambiguous Table S1 CORRAL default YES
C4 CORRAL S. cerevisiae:C. neoformans relative-abundance ratio between 1 and 2 (vs EukDetect ≥10×) Fig 4D CORRAL default YES (CORRAL side); EukDetect side optional

Out of scope (not attempted / different data or wet-lab)

  • Simulated-data benchmarks (Fig 1, 2, 4A, 4B, 5): require the authors' read-simulation scripts and a 3977-taxon synthetic genome panel (lower-limit-of-detection, MAPQ filter, mutation-rate, unrepresented-species, closely-related-pairs). Not part of CORRAL's runnable pipeline; depends on custom simulation code not in the CORRAL repo. NOT attempted.
  • DIABIMMUNE real-data (1154 samples; 122/136 concordance, +97 detections): a different, much larger external accession; deprioritised behind the in-scope mock target.
  • MicrobiomeDB deployment (Release 30: 6337 samples, 1453 with eukaryotes, 190 taxa, Table 1 top-15): production-platform deployment over 8 external studies — out of scope.
  • HMP body-site analyses (Fig 7): separate HMP data; out of scope.

Heavy-compute plan («our HPC»/SLURM only)

  1. front1 (internet): clone CORRAL @ pinned commit on «infra»; conda env with nextflow, bowtie2, samtools, marker_alignments (pip); fetch EukDetect marker reference DB (bowtie2 index + marker-to-taxon map); download the 36 D6300 paired fastqs (~358 GB) into «infra» via ENA FTP.
  2. SLURM (compute only): `nex
Figures / tables: Fig 4CTableFig 4D
C1
Reported
CORRAL detects both mock fungi: S. cerevisiae (unambiguous) + C. neoformans (flagged ambiguous) in D6300 MMC (Fig 4C)
Reproduced
35/35 samples: S. cerevisiae called unambiguously (no '?'), C. neoformans flagged ambiguous ('?'), both detected
exact
C2
Reported
ZYMO-method MMC: S. cerevisiae ~27000 reads/182 markers; C. neoformans ~15000 reads/150 markers (Table S1)
Reproduced
Sc 182 markers EXACT (29/29 adequate-depth); deepest samples Sc 26082-31455 reads / Cn 14818-18043 reads (e.g. D6300_4_23 Sc26082/Cn15716, D6300_3_1 Sc29164/Cn14818); Cn markers mode 170
within tolerance
C3
Reported
False-positive hits ~44 reads/<8 markers, flagged ambiguous (Table S1)
Reproduced
False-positive ambiguous taxa: mean 30.2 reads, mean 4.6/max 10 markers, ALL '?'-flagged
within tolerance
C4
Reported
CORRAL S.cerevisiae:C.neoformans reads ratio 1-2 (EukDetect >=10x) (Fig 4D)
Reproduced
CORRAL ratio mean 1.43, median 1.38, max 3.77 (never >=10x), 22/35 in [1,2]
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 89/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

98.4 k
tokens (I/O) · 4.5 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.