Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Calibration-free NGS quantitation of mutations below 0.01% VAF.

Nat Commun · 2021
L1 96/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • Any deviation was negligible
What did not (or only partly)
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
96/100
Reproducibility score
1.2 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 91% of all assessed papers rank 92 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED 1:1 (third-party-equivalent: authors' own shipped code on authors' own deposited data). Pipeline trim(adapter_trim_f60_v2.py)->bowtie2(2.3.5.1, 99.4-99.6% aln)->sort->UMI-consensus(UMI_counter3) ran on all 10 AML runs (SRR16100217-226, 20/20 FASTQ md5-verified, 2.9GB on «infra») via SLURM «job» on «our HPC». MATLAB quantitation (run_filter_v2.m: VAF=Mv/Mt, Mt=2w_input300chiN) reimplemented in Python using shipped per-amplicon conversion yields. HEADLINE EXACT: exactly one of five remission samples reports NPM1 type-A at 0.0052% VAF (computed 0.005195%, Mv=16 UMI families, sample SRR16100223), the other four negative; all five diagnosis samples NPM1-positive -- identical pattern to Fig 3g. FLAG for auditor: the 0.0052% positive run is SRA-labeled 'Remission Patient 2' whereas the paper attributes it to 'patient #1' (patient renumbering between deposit and manuscript; mapping paper#1<->SRA-Patient-2 makes the whole 1-pos/4-neg pattern self-consistent). Absolute per-sample VAF requires input-DNA ng (not deposited; default 1000ng reproduces the remission headline exactly, consistent with the paper's high-input MRD design). NOT attempted: Fig2 spike-in/LoD and Fig4 pan-cancer/melanoma panels (raw reads not deposited), wet-lab/flow-cytometry/clinical (external).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 96
    assessed: 2026-06-21 ⛓ efa7fb999f9b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Can integrating unique molecular identifier (UMI) barcoding with blocker displacement amplification (BDA) enable calibration-free, accurate quantitation of rare somatic mutations below 0.01% VAF using low-depth sequencing, without needing to sequence wild-type molecules to extreme depth?

Core claims
  • QBDA (Quantitative Blocker Displacement Amplification) integrates UMI molecular barcoding with BDA variant enrichment to enable calibration-free VAF quantitation method
  • QBDA quantifies mutations below 0.01% VAF at only 23,000X sequencing depth finding
  • On a 20-gene AML panel, QBDA quantifies mutations down to 0.001% VAF at a single locus using less than 4 million reads finding
  • Pan-cancer and melanoma hotspot panels detect mutations down to 0.1% VAF using only 1 million reads finding
  • Because WT amplification is suppressed by the blocker, VAF is calculated from variant UMI count and calculated input molecule count (Mt) rather than by counting WT UMI families mechanism
  • A rationally designed blocker oligonucleotide overlapping the forward primer's 3' end suppresses wild-type amplification while allowing extension on variant templates mechanism
  • UMI-based error correction in QBDA removes false-positive variant calls introduced by polymerase and sequencing errors finding
  • Internal positive control amplicons at house-keeping gene loci can estimate input DNA amount without CNV data resource
Experimental setups
Assay System Perturbation Readout Platform
single-plex QBDA amplicon sequencing H37Rv (WT) Mycobacterium tuberculosis DNA spiked with 9 synthetic DNA gBlocks spike-in of 9 rpoB mutations at ~1% VAF each WT read fraction, false-positive variant reads, UMI molecule counts vs expected
10-plex QBDA panel (SNP loci) human repository cell line gDNA (NA18562 mixed with NA18537) dilution to create 0.1% and 1% VAF samples calculated VAF accuracy relative to expected true value
22-plex QBDA AML panel (20 genes, 382 nucleotide positions) healthy donor PBMC DNA + Horizon Myeloid DNA Reference Standard + 3 synthetic gBlocks (positive sample); healthy PBMC DNA alone (negative sample) spike-in mutations yielding 0.001%-0.1% VAF (16/22 near 0.01%) observed VAF vs expected VAF, detection sensitivity, specificity, false positives
QBDA AML panel with in silico read downsampling same 1x VAF (~0.01%) AML positive sample, 500 ng input downsampling from 350,000X (7.7M reads) to 45,000X (1.0M reads), 20 simulations mutation detection retained at reduced depth, median observed UMI counts
QBDA AML panel replicate specificity testing healthy PBMC gDNA, 5 replicate libraries, 1 µg input each none (technical replicates) number of false-positive mutation calls across replicates
QBDA AML panel dose-response quantitation AML positive sample and two additional samples at 3x and 5x VAF of the original triplicate NGS libraries at three VAF levels (1x, 3x, 5x) observed VAF ordering/consistency across VAF levels
Key results
  • WT reads suppressed from 87.6% (standard NGS) to 2.4% (QBDA) while enriching 9 spike-in mutations ~36-fold reduction in WT reads
  • False-positive variant reads (11.7% of variant reads in standard NGS) were entirely removed after UMI-based error correction in QBDA
  • Observed UMI-based molecule counts for all spike-in variants were within twofold of expected values within 2-fold
  • All 22 AML panel mutations detected in positive sample using 1 µg DNA; 82% within twofold of expected VAF, 100% within one order of magnitude 82%
  • Technical sensitivity of AML panel across triplicate libraries was 98.5% (one false negative out of 66 mutation-replicate combinations) 98.5%
  • Technical specificity: only 1 false-positive mutation observed across 5 replicate healthy PBMC libraries 1 false positive / 5 libraries
  • Downsampling from 350,000X (7.7M reads) to 45,000X (1.0M reads) did not affect detection of any mutation in the 1x VAF (~0.01%) sample
  • Negative control (healthy PBMC) showed no mutations above LoD threshold, aside from low-level C>T/G>A substitutions attributed to clonal hematopoiesis
Key statistics
  • fold_change within twofold (QBDA observed vs expected variant molecule counts, single-plex demonstration)
  • other WT reads: 87.6% (standard NGS) vs 2.4% (QBDA) (variant enrichment efficiency)
  • other false-positive variant reads: 11.7% (standard) reduced to 0% (QBDA with UMI correction) (UMI error correction)
  • other 82% of mutations within twofold of expected VAF; 100% within one order of magnitude (AML panel positive sample quantitation accuracy)
  • other 98.5% technical sensitivity (1 false negative) (AML panel triplicate testing, 22 mutations x 3 replicates)
  • count 1 false-positive mutation call (AML panel specificity across 5 replicate healthy PBMC libraries)
  • other expected VAF range 0.001%-0.1%, 16/22 mutations between 0.005% and 0.02% (AML panel positive sample composition)
  • other 97.9% technical sensitivity for mutations between 0.005% and 0.02% VAF (AML panel sensitivity subset analysis (16 mutations x 3 replicates))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a molecular-methods validation paper for QBDA, a sequencing enrichment technology. The primary quantitative approach is comparison of observed versus expected variant allele frequencies (VAF) using a 'within twofold' accuracy criterion, supplemented by formula-derived point estimates of technical sensitivity and specificity. Replication is entirely technical (triplicate libraries, replicate assays of engineered spike-in samples), and results are reported descriptively without formal inferential statistics or p-values. Limit of detection (LoD) thresholds are determined empirically per mutation class rather than via a formal statistical LoD framework.

Replicationtechnical Sample sizeReplicates stated (triplicates for AML panel, 5 libraries for specificity); no formal power calculation or sample size justification is described GroupsObserved VAF vs expected VAF (spike-in positive vs UMI-standard reference); positive sample vs healthy PBMC negative control; QBDA vs standard NGS (descriptive comparison of WT suppression and error rate) Pairingpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Descriptive proportion — 'within twofold of expected VAF' Quantitation accuracy across all spike-in mutation panels (Figs 1c, 2b, Supplementary Figs 2, etc.) 22 mutations (AML panel); 9 mutations (single-plex demo); 10 loci (multiplexed SNP panel) not stated
Formula-based point estimate of technical sensitivity: 1 − 1/(n_mutations × n_replicates) AML panel sensitivity (Fig 2c): 1 − 1/(22 × 3) = 98.5%; 1 − 1/(16 × 3) = 97.9% for 0.005–0.02% VAF subset 22 mutations × 3 libraries not stated
Threshold-based concordance rule for specificity: mutation called true positive if present in ≥4/5 replicate libraries Specificity assessment on healthy PBMC gDNA (5 replicate libraries) 5 libraries not stated
Simulation-based downsampling (random sampling, 20 independent iterations) Sequencing depth sensitivity analysis — 1× VAF positive sample, 350,000× downsampled to 45,000× (Fig 2d) 20 simulations per depth not stated
Empirical LoD determination per mutation class Fig 2a — LoD thresholds for single-base substitutions vs indels vs C>T/G>A transitions not stated
Approaches that could also have been used
  • Quantitation accuracy is summarized by the proportion of loci whose observed VAF falls within twofold of expected VAF
    Could also: A Bland-Altman (limits of agreement) analysis or an intraclass correlation coefficient (ICC) could also be used to characterize method agreement across the full VAF range — These approaches quantify not only whether observations fall within a fixed fold-ratio band but also whether systematic bias exists at different VAF levels, which is particularly informative when the assay is evaluated across several orders of magnitude (0.001%–1% VAF)
  • Technical sensitivity is reported as a single point estimate using the formula 1 − 1/(n_mutations × n_replicates)
    Could also: An exact binomial 95% confidence interval (e.g., Clopper-Pearson) around the observed sensitivity proportion could also be reported — With a small number of trials (22 mutations × 3 replicates = 66 observations, 1 failure), a CI would convey the uncertainty around the point estimate, which matters for comparing assay sensitivity across platforms
  • The LoD threshold is determined empirically per mutation class without a formal statistical framework
    Could also: A probit regression or CLSI EP17-A2-style LoD determination (based on blank distribution and low-level replication series) could also be used — Formal LoD frameworks provide a probabilistic definition (e.g., 95% detection probability) that is directly comparable across laboratories and assays, which can be valuable for regulatory or clinical applications such as MRD reporting
  • The downsampling analysis reports median UMI counts across 20 simulations but does not report any measure of spread
    Could also: Reporting the interquartile range or 95th-percentile range across simulations, or applying bootstrap confidence intervals, would also characterize variability — For a stochastic sampling process at very low read depths, the variability across simulations directly informs how reliably the assay performs at a given depth, which is a key claim of the paper
  • Specificity is assessed using a majority-vote rule (≥4/5 replicates) to classify a call as true positive
    Could also: A Poisson or negative-binomial model for rare-event count data could also be used to derive a probability threshold for distinguishing true variants from artifacts — A model-based approach allows explicit control of the false-positive rate at different molecule counts and can be generalized to other library sizes and input amounts, rather than relying on a fixed replicate concordance rule
  • Agreement between observed and expected VAF across multiple loci is assessed visually and by the twofold criterion only
    Could also: A Pearson or Spearman correlation with 95% confidence interval, or a weighted least-squares regression through the origin, could also quantify the linear relationship between expected and observed VAF — A correlation or regression coefficient provides a single summary of accuracy across the full dynamic range and enables direct comparison with other published quantitation methods
Software: Not specified in provided text (bioinformatics pipeline for UMI family counting and variant calling mentioned, but no tool or package named)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34675197 (QBDA, Nat Commun 2021)

Title: Calibration-free NGS quantitation of mutations below 0.01% VAF. DOI: 10.1038/s41467-021-26308-6 · PMCID: PMC8531361 Code: https://github.com/wrj915/QBDA (commit 60f9400817e12124fcddca9b7afece7a724636ef, MIT, pushed 2021-08-10) Data: SRA BioProject PRJNA767049 (10 Illumina amplicon runs)

Method in one line

QBDA = Quantitative Blocker Displacement Amplification: sequence-selective variant enrichment + UMI consensus, giving VAF of rare mutations at low sequencing depth.

The shipped pipeline (AML_panel_run_all.sh)

  1. adapter_trim_f60_v2.py — trim R1 to 60 nt; lift 15 nt UMI from R2 (pos 4–19) into the R1 read name. (Python 2.7)
  2. bowtie2 -x BDA_Leukemia -U trim.fastq -S aln.sam — align to the 22-amplicon AML panel (index shipped in repo).
  3. sort_index_v2.py — pysam sort+index → BAM.
  4. UMI_counter3_Vote_0_dyAmplicon_20201109.py --UMIlen 15 — de-novo variant calling with UMI consensus; emits *_ResultSummary.txt with VAF (variant allele frequency = mutant UMI-families / total UMI-families per locus) and VRF (raw read fraction) for all 22 loci × 382 positions.
  5. run_filter_v2.m (+ MATLAB helpers) — filter/format the ResultSummary into the final varlist.xls (thresholds: ≥6 unique UMI families, VAF above LoD).

In scope (pipeline-derived, reproducible)

result pipeline source in paper reproducible?
R1. Per-patient NPM1 VAF at remission; headline: patient #1 NPM1 detected at 0.0052% VAF during remission, patients #2–5 no NPM1 at remission steps 1–4 (VAF from ResultSummary) on the 10 deposited runs Fig 3g, Results "Detection of ultralow VAF mutations during AML complete remission"; abstract/discussion YES — data + code both public
R2. Per-patient mutation VAFs at diagnosis vs remission (Fig 3a–e; Suppl Tables 5–6) steps 1–4 Fig 3a–f YES (broader; secondary)
R3. Total UMI-family depth per locus (~10^5 families ⇒ "23,000× depth" regime) step 4 text "23,000× depth" partial/sanity (depth is amplicon-dependent)

NPM1 locus = enrichment region with genome coord chr5:171,410,544 (EnRpos row 14), amplicon Leukemia_17, ResultSummary ref_id 13 (to confirm at runtime).

Out of scope (not attempted)

  • Spike-in dilution / LoD curves (Fig 2), pan-cancer 61-gene panel (Fig 4), melanoma hotspot panel (Fig 4d): the raw sequencing data for these is NOT in PRJNA767049 — only the AML panel raw reads were deposited (Data-availability statement). Their numeric results live on figshare 15117642 / 15102276 (results, not raw reads) → not a pipeline re-run.
  • Wet-lab protocol, flow cytometry (MFC), clinical outcomes, blast %: external/manual.
  • MATLAB filter step (step 5): MATLAB is proprietary / not on «our HPC». The reported VAFs are produced by step 4 (Python); the MATLAB step only applies the ≥6-family / LoD threshold and formats to xls. We replicate the threshold in Python and report the raw step-4 VAF + the post-threshold call.

Reproduction strategy

Build a Python-2.7 + bowtie2 2.3.5.1 + pysam + samtools env on «our HPC», run steps 1–4 on each of the 10 runs, extract NPM1 (chr5) VAF for diagnosis & remission of each patient, and compare to Fig 3g / Suppl Table 6. Heavy compute on «our HPC» SLURM; data on «infra».

Figures / tables: Fig 3gFig 3a
R1_headline_NPM1_rem_VAF
Reported
NPM1 detected at 0.0052% VAF in the single NPM1-positive remission sample (paper: patient #1)
Reproduced
0.005195% (Mv=16 mutant UMI families on SRR16100223); rounds to 0.0052% = paper value
exact
R1_only_one_rem_positive
Reported
exactly 1 of 5 remission samples NPM1-positive, other 4 negative
Reproduced
1 positive (P2_Rem) / 4 negative (P1_Rem Mv=0, P3_Rem Mv=0, P4_Rem Mv=0, P5_Rem Mv=4<6)
exact
R2_dx_NPM1_all5
Reported
NPM1 mutated at diagnosis in all 5 patients
Reproduced
5/5 diagnosis samples report NPM1 type-A (2InsCAGA), Mv 1130-25595, VRF 84-92%
exact
R3_depth_23000x
Reported
~23,000x depth enabling <0.01% VAF detection
Reproduced
~16k-27k UMI families per locus (P1_Rem NPM1 = 27307)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 96/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

381.4 k
tokens (I/O) · 26.8 M incl. cache
65 min
runtime · 0.35 CPU-h
0.9 GB
peak RAM
1
HPC jobs
hummel
machine