Single duplex DNA sequencing with CODEC detects mutations with high sensitivity.
The main results reproduced: recomputed values matched the published ones within tolerance.
- Nothing in this column.
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Tool-level reproduction SUCCEEDED; figure-level reproduction is data-restricted. CODECsuite (broadinstitute/CODECsuite @ 418f6aad, codec v1.1.5) was built from source on a «our HPC» compute node and run end-to-end on its bundled test/human_wgs fastqs: codec trim processed 10000 read pairs (HIGH_CONF 94.7%) and emitted a uBAM carrying the CODEC-specific single-molecule duplex tags (RX/QX UMI, s5/q5 + s3/q3 dual stems, sl/ql linker, bc barcode); BWA-MEM aligned the trimmed reads (37.2% primary-mapped to GRCh38 chr20, the test set being chr20-enriched). This validates the code artifact and the method's core CODEC adapter/duplex-tagging behaviour (P16). Building on a current gcc>=10 toolchain required 3 documented portability fixes (conda include path for bzlib.h; -fcommon shim for fermi-lite's rle_auxtab tentative definition; excluding conda htslib so -lhts resolves to SeqLib's bundled libhts that exports bam_hdr_destroy) — none affect results. NOT attempted: all 17 headline quantitative figures (Fig 2-5, residual error rates, fold-improvements, CH burden, mutational-signature cosine, HRD correlation, MSI sensitivity) — each is pipeline-derived from raw sequencing deposited at dbGaP phs003255.v1.p1 (controlled access, clinical human samples), and the cited Zenodo 'data' DOI is actually the source-code archive. No fabrication signal: open code that builds and emits exactly the claimed duplex structure, with figure data behind standard controlled-access governance. Grades provisional; a human reviewer signs off.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-19 ⛓ 17efd7739fbb
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe authors hypothesized that physically linking the Watson and Crick strand information of a DNA duplex before strand dissociation would allow standard next-generation sequencing to achieve single-duplex resolution with much higher efficiency than existing duplex sequencing methods.
- ★ CODEC concatenates both strands of an original DNA duplex into a single NGS read pair via an adapter quadruplex and strand-displacing extension, enabling single-duplex resolution method
- ★ CODEC affords 1,000-fold higher accuracy than standard NGS while using up to 100-fold fewer reads than duplex sequencing finding
- ★ CODEC revealed a sperm mutation frequency of 2.72 x 10^-8 in a 39-year-old individual finding
- ★ CODEC detects genome-wide clonal hematopoiesis and age-acquired somatic mutations in blood cells from single DNA duplexes finding
- ★ CODEC detects microsatellite instability with 10-fold greater sensitivity and reveals mutational signatures finding
- The CODEC adapter/index design suppresses index hopping compared to typical unique dual indices finding
- ★ CODEC WGS costs 87-fold less than duplex sequencing while retaining higher accuracy than standard WGS finding
- ★ CODEC detects specific tumor mutations from tumor genomes and liquid biopsies using up to 100-fold fewer reads finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Targeted NGS with pan-cancer hybridization capture panel (800 kb) | cfDNA from a cancer patient and a healthy donor | none | residual SNV/indel frequency, comparison to duplex sequencing/SSC/R1+R2 consensus | Illumina NGS |
| Whole-genome sequencing (WGS) | NA12878 pilot genome (Genome in a Bottle Consortium) | none | residual mutation frequency and sequencing cost, CODEC vs duplex sequencing vs standard WGS | Illumina NGS |
| Whole-exome sequencing (WES) | human genomic DNA | none | residual SNV frequency | Illumina NGS |
| WGS with varied end-repair/dA-tailing (ER/AT) methods | human sperm DNA (39-year-old donor) | commercial ER/AT vs Duplex-Repair vs ddBTP-blocked ER/AT | residual SNV frequency | Illumina NGS |
| Downsampled WGS germline variant calling | NA12878 | coverage downsampling (1x-40x) | false positive/false negative rates for germline SNPs | Illumina NGS |
| Low-coverage WGS (6x) with Duplex-Repair | buffy coat germline DNA from 15 breast cancer patients | none (age as variable) | somatic mutation count vs age, clonal hematopoiesis detection | Illumina NGS |
| Targeted deep duplex sequencing (validation) | buffy coat DNA from 8 breast cancer patients | none | cross-validation rate of CODEC-detected somatic mutations | Illumina NGS |
| CODEC WGS with duplex sequencing cross-validation | cfDNA and matched buffy coat DNA from 4 healthy donors and 4 breast cancer patients | none | fraction of cfDNA single-duplex mutations validated in buffy coat, VAF, trinucleotide context | Illumina NGS |
- – CODEC residual SNV frequency was similar to duplex sequencing on a targeted panel 2.9x10^-7 vs 4.3x10^-7
- ▼ CODEC residual mutation frequency was much lower than single-strand consensus (SSC) 234-fold
- ▼ CODEC recovered unique original duplexes using far fewer raw read pairs than duplex sequencing 220-fold fewer reads
- ▼ CODEC WGS was substantially cheaper than duplex sequencing while more accurate than standard WGS 87-fold lower cost
- ▼ ddBTP-blocked ER/AT paired with CODEC reduced sperm residual SNV frequency relative to commercial ER/AT kit 18.6-fold (to 2.72x10^-8)
- – CODEC showed fewer false positives but more false negatives than standard WGS for germline SNPs at low coverage 21-fold fewer FP, 2-fold more FN
- ▲ Only CODEC revealed a linear relationship between somatic mutation count and donor age R^2=0.80
- ▼ CODEC adapter/index structure reduced index hopping rate relative to standard unique dual indices 0.056% vs 0.16%
- correlation R^2=0.80 (linear relationship between somatic mutation count and age in buffy coat DNA via CODEC)
- fold_change 234-fold (CODEC residual SNV frequency vs single-strand consensus (SSC))
- fold_change 220-fold (fewer read pairs needed by CODEC vs duplex sequencing to recover original duplexes)
- fold_change 87-fold (cost of CODEC WGS vs duplex sequencing WGS on NA12878)
- other 2.72x10^-8 (residual/mutation frequency in sperm DNA with ddBTP-blocked ER/AT plus CODEC)
- other 21-fold fewer FP, 2-fold more FN (CODEC vs standard WGS germline SNP calling at 1x-5x coverage)
- count 20.2 vs 19.8 mutations per year (somatic mutations acquired per year in mature white blood cells, CODEC vs prior report)
- other 0.056% vs 0.16% (index hopping rate, CODEC structure vs typical unique dual indices)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a methods/technical paper introducing a new duplex sequencing approach (CODEC) that is benchmarked against standard NGS, conventional duplex sequencing, and other consensus strategies. Results are reported mainly as descriptive comparisons of mutation frequencies, fold-differences, and recovery rates across cfDNA, buffy coat, sperm, and reference-genome samples, with data points and error bars representing means and 95% binomial confidence intervals (Wilson method). A simple linear regression (with R²) was used to relate somatic mutation counts to donor age; no explicit hypothesis-testing framework (e.g., p-values, t-tests, ANOVA) or multiple-comparisons correction is described in the excerpted text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Linear regression (least-squares, R² reported) | Relationship between number of somatic (clonal hematopoiesis) mutations and donor age in 6x CODEC vs standard WGS of buffy coat DNA (Fig. 3c) | 15 breast cancer patients | not stated |
| 95% binomial confidence intervals (Wilson method) around mean residual mutation/recovery frequencies | Residual SNV/indel frequencies and duplex recovery comparisons across CODEC, duplex sequencing, and other consensus methods (Fig. 2a,b,d,e) | varies by panel/sample as described per experiment (e.g., 2 individuals for Fig. 2a; not a single pooled n) | not stated |
-
Differences in residual mutation frequencies between methods (e.g., CODEC vs duplex sequencing vs SSC in Fig. 2a) are described using point estimates with 95% Wilson confidence intervals rather than a formal comparative test.↳ Could also: A rate-ratio test such as a two-sample Poisson test, negative-binomial regression, or Fisher's exact test on mutation/base counts — Would yield a formal p-value or effect-size estimate for the between-method difference in mutation frequency, complementing the descriptive confidence-interval comparison already shown.
-
The relationship between somatic mutation counts and donor age (Fig. 3c) is summarized with a linear regression R² value.↳ Could also: Reporting the regression slope with its confidence interval and p-value, or fitting a Poisson/negative-binomial regression given that mutation counts are count data — Count-based outcomes often show overdispersion relative to a linear model's assumptions, so a count-data regression could describe the variance structure directly, and a slope estimate with CI/p-value would quantify the age association's strength and uncertainty beyond R² alone.
-
False positive and false negative rates for germline variant calling between CODEC and standard WGS at matched coverages (Fig. 3b) are compared as point values on the same samples.↳ Could also: A paired-proportion approach such as McNemar's test — Because both pipelines were applied to the same reference sample (NA12878), a paired test could formally assess whether the observed difference in error rates exceeds what would be expected from sampling variability alone.
-
Multiple comparisons are drawn across coverage levels, cohorts, and methods (targeted panels, WGS, sperm, cfDNA, buffy coat) within the same set of analyses.↳ Could also: A multiplicity-control procedure such as Benjamini-Hochberg false discovery rate adjustment — Would help bound the overall false-positive rate when many related comparisons are summarized together in one study.
-
Cross-validation of single-duplex mutations is reported as simple percentages (e.g., 8.2%, 4.6%, 9.5% of mutations reobserved) without accompanying interval estimates.↳ Could also: Exact (Clopper-Pearson) or Wilson confidence intervals around these validation proportions — Would communicate the precision of the reobservation rate given the modest number of patients/donors contributing to each percentage (e.g., n=8 for one validation cohort).
-
Cohort sizes for each experiment (e.g., 15 or 8 patients, 4 donors and 4 patients) were set by the specific dataset used rather than derived from a stated power calculation.↳ Could also: An a priori power/sample-size calculation based on an assumed effect size — Could clarify, in advance, how likely the chosen sample sizes were to detect the mutation-frequency or recovery differences of interest.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37106072 (CODEC, Bae et al., Nat Genet 2023)
Title: Single duplex DNA sequencing with CODEC detects mutations with high sensitivity.
PMID 37106072 · PMCID PMC10181940 · DOI 10.1038/s41588-023-01376-0
Code: github.com/broadinstitute/CODECsuite @ commit 418f6aad09ebd38e10da8d727db66ddf45cc759e (master, 2026-04-22) · Zenodo 10.5281/zenodo.7705860
Raw data: dbGaP phs003255.v1.p1 (controlled access)
Reference data: NA12878 PacBio CCS from GIAB (open)
What CODEC is / the pipeline
CODEC (Concatenating Original Duplex for Error Correction) is a library-prep + analysis method giving single-molecule duplex information from ~standard WGS depth. The analysis tool CODECsuite (C++/CMake, with a Snakemake end-to-end workflow) performs 5 steps:
codec demux— demultiplex lane fastqs by sample sheetcodec trim— adapter/UMI trimming → uBAM with RX/QX/duplex tags- align (BWA-MEM) to HG38 (SMaHT duplex ref, no decoy)
- duplicate-collapse / consensus (fgbio GroupReadsByUmi + CallMolecularConsensusReads)
codec call— Single Fragment Caller (SFC) somatic mutation calling Downstream: codec2maf/maf2vcf, Mutect2 MAFs, CODEC-MSI (msisensor), mutational-signature (cosine), HRD scoring.
Data-availability verdict (drives scope)
- Code: fully available (GitHub + Zenodo). Reproducible as a tool (P16 third-party/own-tool run is valid).
- Raw sequencing data: dbGaP phs003255.v1.p1 = CONTROLLED ACCESS. Requires an approved
dbGaP Data Access Request (human-subjects). NOT publicly obtainable →
data_restricted. - Every headline quantitative figure (Fig. 2–5, Ext. Data Figs.) is pipeline-derived from this controlled data. None can be numerically reproduced from public sources.
- The repo bundles a tiny smoke-test input only:
test/human_wgs/CODEC_test.{1,2}.fastq.gz(~1.2 + 1.4 MB raw read pairs). No sample sheet / germline BAM / region BED / expected output ships with it → it validates that the software runs, not any paper number.
IN SCOPE (attempted)
| id | what | pipeline | feasible? |
|---|---|---|---|
| S1 | Build CODECsuite from source (CMake/C++14) at pinned commit | build | yes |
| S2 | Run codec demux+codec trim+BWA align on bundled test/human_wgs fastqs; confirm CODEC-specific uBAM (RX/QX duplex tags) + trim metrics are produced end-to-end |
demux/trim/align | yes (tool-level smoke test, no paper-number comparison) |
OUT OF SCOPE (not attempted — reason)
| id | reported result | location | reason not attempted |
|---|---|---|---|
| C1 | CODEC residual SNV freq 2.9×10⁻⁷ (vs duplex 4.3×10⁻⁷) | Fig 2a | data_restricted (dbGaP) |
| C2 | 234-fold lower residual mut freq vs single-strand consensus | Fig 2b | data_restricted |
| C3 | 220-fold fewer reads to recover duplexes; ~2.7 vs 600 rp/target | Fig 2c | data_restricted |
| C4 | standard NGS residual ~1e-4–1e-5; CODEC 87-fold cheaper than duplex | Fig 2d | data_restricted |
| C5 | sperm ddBTP-blocked end-repair residual 2.72×10⁻⁸ | Fig 2e | data_restricted |
| C6 | 72.8% reads retain both-strand (correct product ratio) | Ext Fig 2a | data_restricted (test fastqs give a value but not comparable to paper sample) |
| C7 | 21-fold fewer germline-SNV false positives at 1–5× | Fig 3b | data_restricted |
| C8 | clonal-hematopoiesis mutations vs age R²=0.80, 20.2 mut/yr (WBC) | Fig 3c | data_restricted |
| C9 | 8.2% cfDNA single-duplex mutations confirmed at 2,311× duplex depth | Ext Fig 7a | data_restricted |
| C10 | 2× CODEC 83-fold higher validated SNV fraction vs 2× WGS | Fig 4a | data_restricted |
| C11 | VAF≥0.042% mutations CODEC-exclusive, >25% validated | Fig 4c | data_restricted |
| C12 | mutation-signature cosine 0.98 (CODEC) vs 0.61 (NGS); >0.9 to 0.05× (140-fold) | Fig 4d–e | data_restricted |
| C13 | HRD Pearson 0.91 (CODEC vs Mutect2) | Fig 4h | data_restricted |
| C14 | up to 100-fold fewer read pairs for tracked-mutation d |
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.