Calibration-free NGS quantitation of mutations below 0.01% VAF.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓Any deviation was negligible
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED 1:1 (third-party-equivalent: authors' own shipped code on authors' own deposited data). Pipeline trim(adapter_trim_f60_v2.py)->bowtie2(2.3.5.1, 99.4-99.6% aln)->sort->UMI-consensus(UMI_counter3) ran on all 10 AML runs (SRR16100217-226, 20/20 FASTQ md5-verified, 2.9GB on «infra») via SLURM «job» on «our HPC». MATLAB quantitation (run_filter_v2.m: VAF=Mv/Mt, Mt=2w_input300chiN) reimplemented in Python using shipped per-amplicon conversion yields. HEADLINE EXACT: exactly one of five remission samples reports NPM1 type-A at 0.0052% VAF (computed 0.005195%, Mv=16 UMI families, sample SRR16100223), the other four negative; all five diagnosis samples NPM1-positive -- identical pattern to Fig 3g. FLAG for auditor: the 0.0052% positive run is SRA-labeled 'Remission Patient 2' whereas the paper attributes it to 'patient #1' (patient renumbering between deposit and manuscript; mapping paper#1<->SRA-Patient-2 makes the whole 1-pos/4-neg pattern self-consistent). Absolute per-sample VAF requires input-DNA ng (not deposited; default 1000ng reproduces the remission headline exactly, consistent with the paper's high-input MRD design). NOT attempted: Fig2 spike-in/LoD and Fig4 pan-cancer/melanoma panels (raw reads not deposited), wet-lab/flow-cytometry/clinical (external).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 96assessed: 2026-06-21 ⛓ efa7fb999f9b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-21
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetCan integrating unique molecular identifier (UMI) barcoding with blocker displacement amplification (BDA) enable calibration-free, accurate quantitation of rare somatic mutations below 0.01% VAF using low-depth sequencing, without needing to sequence wild-type molecules to extreme depth?
- ★ QBDA (Quantitative Blocker Displacement Amplification) integrates UMI molecular barcoding with BDA variant enrichment to enable calibration-free VAF quantitation method
- ★ QBDA quantifies mutations below 0.01% VAF at only 23,000X sequencing depth finding
- ★ On a 20-gene AML panel, QBDA quantifies mutations down to 0.001% VAF at a single locus using less than 4 million reads finding
- ★ Pan-cancer and melanoma hotspot panels detect mutations down to 0.1% VAF using only 1 million reads finding
- ★ Because WT amplification is suppressed by the blocker, VAF is calculated from variant UMI count and calculated input molecule count (Mt) rather than by counting WT UMI families mechanism
- ★ A rationally designed blocker oligonucleotide overlapping the forward primer's 3' end suppresses wild-type amplification while allowing extension on variant templates mechanism
- ★ UMI-based error correction in QBDA removes false-positive variant calls introduced by polymerase and sequencing errors finding
- Internal positive control amplicons at house-keeping gene loci can estimate input DNA amount without CNV data resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-plex QBDA amplicon sequencing | H37Rv (WT) Mycobacterium tuberculosis DNA spiked with 9 synthetic DNA gBlocks | spike-in of 9 rpoB mutations at ~1% VAF each | WT read fraction, false-positive variant reads, UMI molecule counts vs expected | — |
| 10-plex QBDA panel (SNP loci) | human repository cell line gDNA (NA18562 mixed with NA18537) | dilution to create 0.1% and 1% VAF samples | calculated VAF accuracy relative to expected true value | — |
| 22-plex QBDA AML panel (20 genes, 382 nucleotide positions) | healthy donor PBMC DNA + Horizon Myeloid DNA Reference Standard + 3 synthetic gBlocks (positive sample); healthy PBMC DNA alone (negative sample) | spike-in mutations yielding 0.001%-0.1% VAF (16/22 near 0.01%) | observed VAF vs expected VAF, detection sensitivity, specificity, false positives | — |
| QBDA AML panel with in silico read downsampling | same 1x VAF (~0.01%) AML positive sample, 500 ng input | downsampling from 350,000X (7.7M reads) to 45,000X (1.0M reads), 20 simulations | mutation detection retained at reduced depth, median observed UMI counts | — |
| QBDA AML panel replicate specificity testing | healthy PBMC gDNA, 5 replicate libraries, 1 µg input each | none (technical replicates) | number of false-positive mutation calls across replicates | — |
| QBDA AML panel dose-response quantitation | AML positive sample and two additional samples at 3x and 5x VAF of the original | triplicate NGS libraries at three VAF levels (1x, 3x, 5x) | observed VAF ordering/consistency across VAF levels | — |
- ▼ WT reads suppressed from 87.6% (standard NGS) to 2.4% (QBDA) while enriching 9 spike-in mutations ~36-fold reduction in WT reads
- ▼ False-positive variant reads (11.7% of variant reads in standard NGS) were entirely removed after UMI-based error correction in QBDA
- – Observed UMI-based molecule counts for all spike-in variants were within twofold of expected values within 2-fold
- – All 22 AML panel mutations detected in positive sample using 1 µg DNA; 82% within twofold of expected VAF, 100% within one order of magnitude 82%
- – Technical sensitivity of AML panel across triplicate libraries was 98.5% (one false negative out of 66 mutation-replicate combinations) 98.5%
- – Technical specificity: only 1 false-positive mutation observed across 5 replicate healthy PBMC libraries 1 false positive / 5 libraries
- – Downsampling from 350,000X (7.7M reads) to 45,000X (1.0M reads) did not affect detection of any mutation in the 1x VAF (~0.01%) sample
- – Negative control (healthy PBMC) showed no mutations above LoD threshold, aside from low-level C>T/G>A substitutions attributed to clonal hematopoiesis
- fold_change within twofold (QBDA observed vs expected variant molecule counts, single-plex demonstration)
- other WT reads: 87.6% (standard NGS) vs 2.4% (QBDA) (variant enrichment efficiency)
- other false-positive variant reads: 11.7% (standard) reduced to 0% (QBDA with UMI correction) (UMI error correction)
- other 82% of mutations within twofold of expected VAF; 100% within one order of magnitude (AML panel positive sample quantitation accuracy)
- other 98.5% technical sensitivity (1 false negative) (AML panel triplicate testing, 22 mutations x 3 replicates)
- count 1 false-positive mutation call (AML panel specificity across 5 replicate healthy PBMC libraries)
- other expected VAF range 0.001%-0.1%, 16/22 mutations between 0.005% and 0.02% (AML panel positive sample composition)
- other 97.9% technical sensitivity for mutations between 0.005% and 0.02% VAF (AML panel sensitivity subset analysis (16 mutations x 3 replicates))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a molecular-methods validation paper for QBDA, a sequencing enrichment technology. The primary quantitative approach is comparison of observed versus expected variant allele frequencies (VAF) using a 'within twofold' accuracy criterion, supplemented by formula-derived point estimates of technical sensitivity and specificity. Replication is entirely technical (triplicate libraries, replicate assays of engineered spike-in samples), and results are reported descriptively without formal inferential statistics or p-values. Limit of detection (LoD) thresholds are determined empirically per mutation class rather than via a formal statistical LoD framework.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Descriptive proportion — 'within twofold of expected VAF' | Quantitation accuracy across all spike-in mutation panels (Figs 1c, 2b, Supplementary Figs 2, etc.) | 22 mutations (AML panel); 9 mutations (single-plex demo); 10 loci (multiplexed SNP panel) | not stated |
| Formula-based point estimate of technical sensitivity: 1 − 1/(n_mutations × n_replicates) | AML panel sensitivity (Fig 2c): 1 − 1/(22 × 3) = 98.5%; 1 − 1/(16 × 3) = 97.9% for 0.005–0.02% VAF subset | 22 mutations × 3 libraries | not stated |
| Threshold-based concordance rule for specificity: mutation called true positive if present in ≥4/5 replicate libraries | Specificity assessment on healthy PBMC gDNA (5 replicate libraries) | 5 libraries | not stated |
| Simulation-based downsampling (random sampling, 20 independent iterations) | Sequencing depth sensitivity analysis — 1× VAF positive sample, 350,000× downsampled to 45,000× (Fig 2d) | 20 simulations per depth | not stated |
| Empirical LoD determination per mutation class | Fig 2a — LoD thresholds for single-base substitutions vs indels vs C>T/G>A transitions | — | not stated |
-
Quantitation accuracy is summarized by the proportion of loci whose observed VAF falls within twofold of expected VAF↳ Could also: A Bland-Altman (limits of agreement) analysis or an intraclass correlation coefficient (ICC) could also be used to characterize method agreement across the full VAF range — These approaches quantify not only whether observations fall within a fixed fold-ratio band but also whether systematic bias exists at different VAF levels, which is particularly informative when the assay is evaluated across several orders of magnitude (0.001%–1% VAF)
-
Technical sensitivity is reported as a single point estimate using the formula 1 − 1/(n_mutations × n_replicates)↳ Could also: An exact binomial 95% confidence interval (e.g., Clopper-Pearson) around the observed sensitivity proportion could also be reported — With a small number of trials (22 mutations × 3 replicates = 66 observations, 1 failure), a CI would convey the uncertainty around the point estimate, which matters for comparing assay sensitivity across platforms
-
The LoD threshold is determined empirically per mutation class without a formal statistical framework↳ Could also: A probit regression or CLSI EP17-A2-style LoD determination (based on blank distribution and low-level replication series) could also be used — Formal LoD frameworks provide a probabilistic definition (e.g., 95% detection probability) that is directly comparable across laboratories and assays, which can be valuable for regulatory or clinical applications such as MRD reporting
-
The downsampling analysis reports median UMI counts across 20 simulations but does not report any measure of spread↳ Could also: Reporting the interquartile range or 95th-percentile range across simulations, or applying bootstrap confidence intervals, would also characterize variability — For a stochastic sampling process at very low read depths, the variability across simulations directly informs how reliably the assay performs at a given depth, which is a key claim of the paper
-
Specificity is assessed using a majority-vote rule (≥4/5 replicates) to classify a call as true positive↳ Could also: A Poisson or negative-binomial model for rare-event count data could also be used to derive a probability threshold for distinguishing true variants from artifacts — A model-based approach allows explicit control of the false-positive rate at different molecule counts and can be generalized to other library sizes and input amounts, rather than relying on a fixed replicate concordance rule
-
Agreement between observed and expected VAF across multiple loci is assessed visually and by the twofold criterion only↳ Could also: A Pearson or Spearman correlation with 95% confidence interval, or a weighted least-squares regression through the origin, could also quantify the linear relationship between expected and observed VAF — A correlation or regression coefficient provides a single summary of accuracy across the full dynamic range and enables direct comparison with other published quantitation methods
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34675197 (QBDA, Nat Commun 2021)
Title: Calibration-free NGS quantitation of mutations below 0.01% VAF.
DOI: 10.1038/s41467-021-26308-6 · PMCID: PMC8531361
Code: https://github.com/wrj915/QBDA (commit 60f9400817e12124fcddca9b7afece7a724636ef, MIT, pushed 2021-08-10)
Data: SRA BioProject PRJNA767049 (10 Illumina amplicon runs)
Method in one line
QBDA = Quantitative Blocker Displacement Amplification: sequence-selective variant enrichment + UMI consensus, giving VAF of rare mutations at low sequencing depth.
The shipped pipeline (AML_panel_run_all.sh)
adapter_trim_f60_v2.py— trim R1 to 60 nt; lift 15 nt UMI from R2 (pos 4–19) into the R1 read name. (Python 2.7)bowtie2 -x BDA_Leukemia -U trim.fastq -S aln.sam— align to the 22-amplicon AML panel (index shipped in repo).sort_index_v2.py— pysam sort+index → BAM.UMI_counter3_Vote_0_dyAmplicon_20201109.py --UMIlen 15— de-novo variant calling with UMI consensus; emits*_ResultSummary.txtwith VAF (variant allele frequency = mutant UMI-families / total UMI-families per locus) and VRF (raw read fraction) for all 22 loci × 382 positions.run_filter_v2.m(+ MATLAB helpers) — filter/format the ResultSummary into the finalvarlist.xls(thresholds: ≥6 unique UMI families, VAF above LoD).
In scope (pipeline-derived, reproducible)
| result | pipeline | source in paper | reproducible? |
|---|---|---|---|
| R1. Per-patient NPM1 VAF at remission; headline: patient #1 NPM1 detected at 0.0052% VAF during remission, patients #2–5 no NPM1 at remission | steps 1–4 (VAF from ResultSummary) on the 10 deposited runs | Fig 3g, Results "Detection of ultralow VAF mutations during AML complete remission"; abstract/discussion | YES — data + code both public |
| R2. Per-patient mutation VAFs at diagnosis vs remission (Fig 3a–e; Suppl Tables 5–6) | steps 1–4 | Fig 3a–f | YES (broader; secondary) |
| R3. Total UMI-family depth per locus (~10^5 families ⇒ "23,000× depth" regime) | step 4 | text "23,000× depth" | partial/sanity (depth is amplicon-dependent) |
NPM1 locus = enrichment region with genome coord chr5:171,410,544 (EnRpos row 14),
amplicon Leukemia_17, ResultSummary ref_id 13 (to confirm at runtime).
Out of scope (not attempted)
- Spike-in dilution / LoD curves (Fig 2), pan-cancer 61-gene panel (Fig 4), melanoma hotspot panel (Fig 4d): the raw sequencing data for these is NOT in PRJNA767049 — only the AML panel raw reads were deposited (Data-availability statement). Their numeric results live on figshare 15117642 / 15102276 (results, not raw reads) → not a pipeline re-run.
- Wet-lab protocol, flow cytometry (MFC), clinical outcomes, blast %: external/manual.
- MATLAB filter step (step 5): MATLAB is proprietary / not on «our HPC». The reported VAFs are produced by step 4 (Python); the MATLAB step only applies the ≥6-family / LoD threshold and formats to xls. We replicate the threshold in Python and report the raw step-4 VAF + the post-threshold call.
Reproduction strategy
Build a Python-2.7 + bowtie2 2.3.5.1 + pysam + samtools env on «our HPC», run steps 1–4 on each of the 10 runs, extract NPM1 (chr5) VAF for diagnosis & remission of each patient, and compare to Fig 3g / Suppl Table 6. Heavy compute on «our HPC» SLURM; data on «infra».
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.