Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Single duplex DNA sequencing with CODEC detects mutations with high sensitivity.

Nat Genet · 2023
L1 78/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 51% of all assessed papers rank 533 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Tool-level reproduction SUCCEEDED; figure-level reproduction is data-restricted. CODECsuite (broadinstitute/CODECsuite @ 418f6aad, codec v1.1.5) was built from source on a «our HPC» compute node and run end-to-end on its bundled test/human_wgs fastqs: codec trim processed 10000 read pairs (HIGH_CONF 94.7%) and emitted a uBAM carrying the CODEC-specific single-molecule duplex tags (RX/QX UMI, s5/q5 + s3/q3 dual stems, sl/ql linker, bc barcode); BWA-MEM aligned the trimmed reads (37.2% primary-mapped to GRCh38 chr20, the test set being chr20-enriched). This validates the code artifact and the method's core CODEC adapter/duplex-tagging behaviour (P16). Building on a current gcc>=10 toolchain required 3 documented portability fixes (conda include path for bzlib.h; -fcommon shim for fermi-lite's rle_auxtab tentative definition; excluding conda htslib so -lhts resolves to SeqLib's bundled libhts that exports bam_hdr_destroy) — none affect results. NOT attempted: all 17 headline quantitative figures (Fig 2-5, residual error rates, fold-improvements, CH burden, mutational-signature cosine, HRD correlation, MSI sensitivity) — each is pipeline-derived from raw sequencing deposited at dbGaP phs003255.v1.p1 (controlled access, clinical human samples), and the cited Zenodo 'data' DOI is actually the source-code archive. No fabrication signal: open code that builds and emits exactly the claimed duplex structure, with figure data behind standard controlled-access governance. Grades provisional; a human reviewer signs off.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.7705860

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ 17efd7739fbb
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The authors hypothesized that physically linking the Watson and Crick strand information of a DNA duplex before strand dissociation would allow standard next-generation sequencing to achieve single-duplex resolution with much higher efficiency than existing duplex sequencing methods.

Core claims
  • CODEC concatenates both strands of an original DNA duplex into a single NGS read pair via an adapter quadruplex and strand-displacing extension, enabling single-duplex resolution method
  • CODEC affords 1,000-fold higher accuracy than standard NGS while using up to 100-fold fewer reads than duplex sequencing finding
  • CODEC revealed a sperm mutation frequency of 2.72 x 10^-8 in a 39-year-old individual finding
  • CODEC detects genome-wide clonal hematopoiesis and age-acquired somatic mutations in blood cells from single DNA duplexes finding
  • CODEC detects microsatellite instability with 10-fold greater sensitivity and reveals mutational signatures finding
  • The CODEC adapter/index design suppresses index hopping compared to typical unique dual indices finding
  • CODEC WGS costs 87-fold less than duplex sequencing while retaining higher accuracy than standard WGS finding
  • CODEC detects specific tumor mutations from tumor genomes and liquid biopsies using up to 100-fold fewer reads finding
Experimental setups
Assay System Perturbation Readout Platform
Targeted NGS with pan-cancer hybridization capture panel (800 kb) cfDNA from a cancer patient and a healthy donor none residual SNV/indel frequency, comparison to duplex sequencing/SSC/R1+R2 consensus Illumina NGS
Whole-genome sequencing (WGS) NA12878 pilot genome (Genome in a Bottle Consortium) none residual mutation frequency and sequencing cost, CODEC vs duplex sequencing vs standard WGS Illumina NGS
Whole-exome sequencing (WES) human genomic DNA none residual SNV frequency Illumina NGS
WGS with varied end-repair/dA-tailing (ER/AT) methods human sperm DNA (39-year-old donor) commercial ER/AT vs Duplex-Repair vs ddBTP-blocked ER/AT residual SNV frequency Illumina NGS
Downsampled WGS germline variant calling NA12878 coverage downsampling (1x-40x) false positive/false negative rates for germline SNPs Illumina NGS
Low-coverage WGS (6x) with Duplex-Repair buffy coat germline DNA from 15 breast cancer patients none (age as variable) somatic mutation count vs age, clonal hematopoiesis detection Illumina NGS
Targeted deep duplex sequencing (validation) buffy coat DNA from 8 breast cancer patients none cross-validation rate of CODEC-detected somatic mutations Illumina NGS
CODEC WGS with duplex sequencing cross-validation cfDNA and matched buffy coat DNA from 4 healthy donors and 4 breast cancer patients none fraction of cfDNA single-duplex mutations validated in buffy coat, VAF, trinucleotide context Illumina NGS
Key results
  • CODEC residual SNV frequency was similar to duplex sequencing on a targeted panel 2.9x10^-7 vs 4.3x10^-7
  • CODEC residual mutation frequency was much lower than single-strand consensus (SSC) 234-fold
  • CODEC recovered unique original duplexes using far fewer raw read pairs than duplex sequencing 220-fold fewer reads
  • CODEC WGS was substantially cheaper than duplex sequencing while more accurate than standard WGS 87-fold lower cost
  • ddBTP-blocked ER/AT paired with CODEC reduced sperm residual SNV frequency relative to commercial ER/AT kit 18.6-fold (to 2.72x10^-8)
  • CODEC showed fewer false positives but more false negatives than standard WGS for germline SNPs at low coverage 21-fold fewer FP, 2-fold more FN
  • Only CODEC revealed a linear relationship between somatic mutation count and donor age R^2=0.80
  • CODEC adapter/index structure reduced index hopping rate relative to standard unique dual indices 0.056% vs 0.16%
Key statistics
  • correlation R^2=0.80 (linear relationship between somatic mutation count and age in buffy coat DNA via CODEC)
  • fold_change 234-fold (CODEC residual SNV frequency vs single-strand consensus (SSC))
  • fold_change 220-fold (fewer read pairs needed by CODEC vs duplex sequencing to recover original duplexes)
  • fold_change 87-fold (cost of CODEC WGS vs duplex sequencing WGS on NA12878)
  • other 2.72x10^-8 (residual/mutation frequency in sperm DNA with ddBTP-blocked ER/AT plus CODEC)
  • other 21-fold fewer FP, 2-fold more FN (CODEC vs standard WGS germline SNP calling at 1x-5x coverage)
  • count 20.2 vs 19.8 mutations per year (somatic mutations acquired per year in mature white blood cells, CODEC vs prior report)
  • other 0.056% vs 0.16% (index hopping rate, CODEC structure vs typical unique dual indices)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a methods/technical paper introducing a new duplex sequencing approach (CODEC) that is benchmarked against standard NGS, conventional duplex sequencing, and other consensus strategies. Results are reported mainly as descriptive comparisons of mutation frequencies, fold-differences, and recovery rates across cfDNA, buffy coat, sperm, and reference-genome samples, with data points and error bars representing means and 95% binomial confidence intervals (Wilson method). A simple linear regression (with R²) was used to relate somatic mutation counts to donor age; no explicit hypothesis-testing framework (e.g., p-values, t-tests, ANOVA) or multiple-comparisons correction is described in the excerpted text.

Replicationbiological Sample sizeSample sizes are given per experiment/cohort (e.g., 2 individuals for pan-cancer panel comparison, 15 breast cancer patients for CH/age analysis, 8 breast cancer patients for cross-validation, 4 healthy donors and 4 breast cancer patients for cfDNA CH validation, a single reference genome NA12878, one 39-year-old sperm donor), but no a priori power or sample-size calculation is described. GroupsCODEC vs standard NGS/WGS vs duplex sequencing and other consensus methods (SSC, R1+R2) across cfDNA, buffy coat, sperm, and reference-genome samples Pairingmixed Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesyes Confidence intervalsyes
Statistical tests used
Test Applied to n Assumptions
Linear regression (least-squares, R² reported) Relationship between number of somatic (clonal hematopoiesis) mutations and donor age in 6x CODEC vs standard WGS of buffy coat DNA (Fig. 3c) 15 breast cancer patients not stated
95% binomial confidence intervals (Wilson method) around mean residual mutation/recovery frequencies Residual SNV/indel frequencies and duplex recovery comparisons across CODEC, duplex sequencing, and other consensus methods (Fig. 2a,b,d,e) varies by panel/sample as described per experiment (e.g., 2 individuals for Fig. 2a; not a single pooled n) not stated
Approaches that could also have been used
  • Differences in residual mutation frequencies between methods (e.g., CODEC vs duplex sequencing vs SSC in Fig. 2a) are described using point estimates with 95% Wilson confidence intervals rather than a formal comparative test.
    Could also: A rate-ratio test such as a two-sample Poisson test, negative-binomial regression, or Fisher's exact test on mutation/base counts — Would yield a formal p-value or effect-size estimate for the between-method difference in mutation frequency, complementing the descriptive confidence-interval comparison already shown.
  • The relationship between somatic mutation counts and donor age (Fig. 3c) is summarized with a linear regression R² value.
    Could also: Reporting the regression slope with its confidence interval and p-value, or fitting a Poisson/negative-binomial regression given that mutation counts are count data — Count-based outcomes often show overdispersion relative to a linear model's assumptions, so a count-data regression could describe the variance structure directly, and a slope estimate with CI/p-value would quantify the age association's strength and uncertainty beyond R² alone.
  • False positive and false negative rates for germline variant calling between CODEC and standard WGS at matched coverages (Fig. 3b) are compared as point values on the same samples.
    Could also: A paired-proportion approach such as McNemar's test — Because both pipelines were applied to the same reference sample (NA12878), a paired test could formally assess whether the observed difference in error rates exceeds what would be expected from sampling variability alone.
  • Multiple comparisons are drawn across coverage levels, cohorts, and methods (targeted panels, WGS, sperm, cfDNA, buffy coat) within the same set of analyses.
    Could also: A multiplicity-control procedure such as Benjamini-Hochberg false discovery rate adjustment — Would help bound the overall false-positive rate when many related comparisons are summarized together in one study.
  • Cross-validation of single-duplex mutations is reported as simple percentages (e.g., 8.2%, 4.6%, 9.5% of mutations reobserved) without accompanying interval estimates.
    Could also: Exact (Clopper-Pearson) or Wilson confidence intervals around these validation proportions — Would communicate the precision of the reobservation rate given the modest number of patients/donors contributing to each percentage (e.g., n=8 for one validation cohort).
  • Cohort sizes for each experiment (e.g., 15 or 8 patients, 4 donors and 4 patients) were set by the specific dataset used rather than derived from a stated power calculation.
    Could also: An a priori power/sample-size calculation based on an assumed effect size — Could clarify, in advance, how likely the chosen sample sizes were to detect the mutation-frequency or recovery differences of interest.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37106072 (CODEC, Bae et al., Nat Genet 2023)

Title: Single duplex DNA sequencing with CODEC detects mutations with high sensitivity. PMID 37106072 · PMCID PMC10181940 · DOI 10.1038/s41588-023-01376-0 Code: github.com/broadinstitute/CODECsuite @ commit 418f6aad09ebd38e10da8d727db66ddf45cc759e (master, 2026-04-22) · Zenodo 10.5281/zenodo.7705860 Raw data: dbGaP phs003255.v1.p1 (controlled access) Reference data: NA12878 PacBio CCS from GIAB (open)

What CODEC is / the pipeline

CODEC (Concatenating Original Duplex for Error Correction) is a library-prep + analysis method giving single-molecule duplex information from ~standard WGS depth. The analysis tool CODECsuite (C++/CMake, with a Snakemake end-to-end workflow) performs 5 steps:

  1. codec demux — demultiplex lane fastqs by sample sheet
  2. codec trim — adapter/UMI trimming → uBAM with RX/QX/duplex tags
  3. align (BWA-MEM) to HG38 (SMaHT duplex ref, no decoy)
  4. duplicate-collapse / consensus (fgbio GroupReadsByUmi + CallMolecularConsensusReads)
  5. codec call — Single Fragment Caller (SFC) somatic mutation calling Downstream: codec2maf/maf2vcf, Mutect2 MAFs, CODEC-MSI (msisensor), mutational-signature (cosine), HRD scoring.

Data-availability verdict (drives scope)

  • Code: fully available (GitHub + Zenodo). Reproducible as a tool (P16 third-party/own-tool run is valid).
  • Raw sequencing data: dbGaP phs003255.v1.p1 = CONTROLLED ACCESS. Requires an approved dbGaP Data Access Request (human-subjects). NOT publicly obtainable → data_restricted.
  • Every headline quantitative figure (Fig. 2–5, Ext. Data Figs.) is pipeline-derived from this controlled data. None can be numerically reproduced from public sources.
  • The repo bundles a tiny smoke-test input only: test/human_wgs/CODEC_test.{1,2}.fastq.gz (~1.2 + 1.4 MB raw read pairs). No sample sheet / germline BAM / region BED / expected output ships with it → it validates that the software runs, not any paper number.

IN SCOPE (attempted)

id what pipeline feasible?
S1 Build CODECsuite from source (CMake/C++14) at pinned commit build yes
S2 Run codec demux+codec trim+BWA align on bundled test/human_wgs fastqs; confirm CODEC-specific uBAM (RX/QX duplex tags) + trim metrics are produced end-to-end demux/trim/align yes (tool-level smoke test, no paper-number comparison)

OUT OF SCOPE (not attempted — reason)

id reported result location reason not attempted
C1 CODEC residual SNV freq 2.9×10⁻⁷ (vs duplex 4.3×10⁻⁷) Fig 2a data_restricted (dbGaP)
C2 234-fold lower residual mut freq vs single-strand consensus Fig 2b data_restricted
C3 220-fold fewer reads to recover duplexes; ~2.7 vs 600 rp/target Fig 2c data_restricted
C4 standard NGS residual ~1e-4–1e-5; CODEC 87-fold cheaper than duplex Fig 2d data_restricted
C5 sperm ddBTP-blocked end-repair residual 2.72×10⁻⁸ Fig 2e data_restricted
C6 72.8% reads retain both-strand (correct product ratio) Ext Fig 2a data_restricted (test fastqs give a value but not comparable to paper sample)
C7 21-fold fewer germline-SNV false positives at 1–5× Fig 3b data_restricted
C8 clonal-hematopoiesis mutations vs age R²=0.80, 20.2 mut/yr (WBC) Fig 3c data_restricted
C9 8.2% cfDNA single-duplex mutations confirmed at 2,311× duplex depth Ext Fig 7a data_restricted
C10 2× CODEC 83-fold higher validated SNV fraction vs 2× WGS Fig 4a data_restricted
C11 VAF≥0.042% mutations CODEC-exclusive, >25% validated Fig 4c data_restricted
C12 mutation-signature cosine 0.98 (CODEC) vs 0.61 (NGS); >0.9 to 0.05× (140-fold) Fig 4d–e data_restricted
C13 HRD Pearson 0.91 (CODEC vs Mutect2) Fig 4h data_restricted
C14 up to 100-fold fewer read pairs for tracked-mutation d
Figures / tables: Fig 2aFig 2bFig 2cFig 2dFig 2eFig 3b
S1
Reported
CODECsuite builds from source (CMake/C++14) at pinned commit 418f6aad
Reproduced
BUILD OK: codec v1.1.5; cmake rc=0, make rc=0 («our HPC» «job»)
exact
S2
Reported
codec demux+trim+align runs end-to-end on bundled test/human_wgs fastqs -> CODEC uBAM with duplex tags
Reproduced
TRIM OK: 10000 read pairs, HIGH_CONF=9467 (94.7%); uBAM (19846 primary recs) carries CODEC duplex tags RX/QX/s5/q5/s3/q3/sl/ql/bc; BWA-MEM aligned (37.2% primary mapped to chr20). demux N/A (single-sample input)
within tolerance
C1-C17
Reported
all headline figures: residual error 2.9e-7 (Fig2a), 234x/220x/87x advantages, CH burden R^2=0.80, signature cosine 0.98, HRD r=0.91, MSI 0.01% (Fig2-5)
Reproduced
not attempted
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

114.7 k
tokens (I/O) · 4.4 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.