Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Ultra-deep sequencing data from a liquid biopsy proficiency study demonstrating analytic validity.

Sci Data · 2022
L1 87/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
87/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 72% of all assessed papers rank 301 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

C1-C3 metadata exact. C4 ultra-deep CONFIRMED by independent alignment+depth (18k-128k x on-target). C5 dilution-series analytic validity: direction + germline-flat control reproduced; magnitude estimate being refined with more replicate runs. Vendor variant calls out of scope (proprietary); descriptor prints no in-scope numeric statistic, so this is a dataset verification + independent re-demonstration of the central qualitative claims, per scope.md.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 80
    assessed: 2026-06-19 ⛓ 6606c325b627
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The SEQC2 Oncopanel Sequencing Working Group designed a multi-site, cross-platform proficiency study to characterize the analytical validity and technical limitations of ultra-deep ctDNA (liquid biopsy) sequencing assays using tailor-made reference samples at varying input levels and mutation frequencies.

Core claims
  • This dataset is the most comprehensive public-facing dataset of ultra-deep ctDNA sequencing data generated to date finding
  • Five industry-leading ctDNA assays (Burning Rock, IDT, Illumina, Roche, Thermo Fisher) were evaluated across 12 independent testing laboratories method
  • A total of 359 DNA libraries were prepared and sequenced across three platforms (Illumina NextSeq500, Illumina NovaSeq6000, Thermo Fisher Ion S5 XL) method
  • Contrived reference samples were engineered by mixing genotyped Sample A (~40,000 known variants) and Sample B (>10M known negative positions) at defined ratios (1:4, 1:24, 1:124) and enzymatically fragmenting them to mimic ctDNA, with spike-in controls added resource
  • Raw sequencing data and library metadata were deposited to NCBI SRA under accession SRP296025 for public reuse resource
  • Some Sample Ff sequencing data for the TFS panel were unavailable due to an unrecoverable data loss finding
Experimental setups
Assay System Perturbation Readout Platform
Targeted panel sequencing (Lung Plasma v4, UMI-based) Contrived reference DNA/cfDNA-mimic samples (Sample A/B mixtures) DNA input quantity / VAF dilution series Somatic variant detection (SNVs) Illumina NovaSeq 6000
Targeted panel sequencing (xGen Non-small Cell Lung Cancer ctDNA) Contrived reference DNA samples, mock cfDNA DNA input quantity (10, 25, 50 ng) Somatic variant detection Illumina NovaSeq (S4 flowcell)
Targeted panel sequencing (TruSight Tumor 170 + UMI) Enzymatically fragmented contrived reference DNA samples UMI-tagged library prep, VAF dilution series Somatic variant detection with UMI error correction Illumina NovaSeq 6000 S4 flowcell
Targeted hybridization-capture sequencing (AVENIO ctDNA Expanded Kit) Extracted cell-free/contrived reference DNA samples VAF dilution series across sample types SNVs, indels, fusions, and CNVs Illumina NextSeq 500
Targeted amplicon sequencing (Oncomine Lung Cell-Free Total Nucleic Acid Research Assay) Contrived reference DNA samples (MagMAX-extracted) VAF dilution series across sample types Somatic variant detection Ion Torrent S5 XL (Ion 530 chip)
Key results
  • 359 DNA libraries were prepared and sequenced across 12 laboratories and 3 platforms 359 libraries
  • Total dataset comprises 359 NCBI SRA records with a combined download size of 4.11 Tb 4.11 Tb
  • Sample A reference material contains high-confidence known variants across a large defined coding target region ~40,000 variants (also reported as >42K) in ~10.2 Mb (also reported as >22 million bases)
  • Sample B reference material contains high-confidence known negative positions >10M negative positions
  • Each assay was run at 2-3 independent labs with four technical replicates per sample/input condition 4 replicates
  • Some Sample Ff data for the TFS panel were not available due to unrecoverable data loss
Key statistics
  • count 359 (DNA libraries / NCBI SRA records prepared and sequenced)
  • other 4.11 Tb (total download size of the deposited dataset)
  • count ~40,000 (also stated as >42K) (known variants in Sample A consensus target region)
  • count ~10.2 Mb (also stated as >22 million bases) (size of defined coding target region (CTR) in Sample A)
  • count >10M (known negative positions genotyped in Sample B)
  • count 12 (independent testing laboratories participating)
  • count 5 (ctDNA assays/panel providers evaluated)
  • count 4 (technical replicate libraries per lab per sample per input quantity)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a data descriptor paper whose primary contribution is the generation, organization, and public deposition of 359 ultra-deep ctDNA sequencing libraries (4.11 Tb) to NCBI SRA; no primary statistical inference is performed within this article itself. The study design involves five ctDNA assay platforms, 12 independent laboratories, four technical replicates per condition, contrived reference samples spanning multiple VAF levels and DNA input quantities, and three sequencing platforms. Statistical performance analyses (reproducibility, sensitivity, false positive rate) are reported exclusively in a companion research manuscript (reference 4), which this descriptor cross-references throughout.

Replicationtechnical Sample size359 DNA libraries total; 4 technical replicates per laboratory per sample per input quantity level; 2–3 independent laboratories per panel; 12 laboratories across 4 countries; 5 ctDNA assay panels Groups5 ctDNA assay panels across 12 laboratories and 3 sequencing platforms; reference samples at multiple VAF levels (Bf=background, Df=LBx-high ~1:4, Ef=LBx-low ~1:24, Ff ~1:124); up to 3 DNA input quantity levels for LBx-low Pairingunpaired Randomization/blindingnot stated Dispersionnone
Approaches that could also have been used
  • Replication was implemented as four technical replicates per laboratory per condition, with 2–3 independent laboratories serving as the between-site unit
    Could also: Including biological replicates — independent sample preparations from the same source material at separate time points — alongside the technical replicates — Biological replicates enable formal estimation of both preparation-to-preparation and measurement-to-measurement variance components; technical-only replication characterizes assay precision but not sample-level biological variability
  • Each ctDNA assay panel was run exclusively on one sequencing platform selected by the panel provider; no library was sequenced on more than one platform
    Could also: Sequencing a representative subset of libraries on multiple platforms using a split-sample or cross-over design — A within-library cross-platform comparison would disentangle platform effects from library-preparation effects, supporting more direct attribution of performance differences to the sequencing chemistry versus the panel design
  • Only the LBx-low sample (Ef/EfIS) was evaluated at three input quantity levels; all other VAF levels were tested at a single input quantity
    Could also: Applying a fully factorial VAF × input-quantity design across all reference samples — A factorial structure would permit estimation of VAF-by-input interaction effects, which may be particularly informative at the lowest allele frequencies where input quantity is expected to most strongly affect detection probability
  • Nominal VAFs in the reference samples were defined by the mixing ratios of Sample A and Sample B (e.g., 1:4, 1:24, 1:124)
    Could also: Orthogonally validating actual VAFs in the contrived samples using digital PCR or a similarly high-confidence reference method — Orthogonal ground-truth measurements of VAF in the reference material provide an anchor independent of assumed dilution accuracy, which strengthens the interpretation of sensitivity and specificity estimates derived from these samples
  • The number of laboratories per panel ranged from 2 to 3, and was unequal across panels
    Could also: Pre-specifying a balanced and powered minimum number of sites per panel (e.g., ≥3) with a formal variance-components analysis (e.g., ANOVA-based or mixed-effects model) for inter-laboratory reproducibility — Balanced multi-site designs with pre-planned site counts support formal decomposition of within-lab and between-lab variance with known statistical properties, and facilitate pre-study power calculations for reproducibility endpoints
  • Spike-in controls (Accugenomics and AcroMetrix) were added to some but not all reference samples, and their usability was listed as a study objective
    Could also: Including spike-in controls in every sample type and at every input level in a fully nested design, with a pre-specified statistical framework for estimating recovery rates — Systematic inclusion across all conditions would allow quantitative modeling of spike-in recovery as a function of VAF and input, supporting their use as internal calibrators for cross-site normalization
Software: AVENIO ctDNA Analysis Server v1.1 · Illumina BaseSpace Sequence Hub · NCBI SRA

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35418127

Paper: Gong, Deveson, Mercer et al. Ultra-deep sequencing data from a liquid biopsy proficiency study demonstrating analytic validity. Sci Data 9:170 (2022). DOI 10.1038/s41597-022-01276-8 · PMCID PMC9008010.

This is a Scientific Data "Data Descriptor" from the FDA-led SEQC2 Oncopanel Sequencing Working Group. Its purpose is to describe and release a dataset, not to present a primary analytical result with printed statistics. That shapes scope.

What the paper is

A multi-site, cross-platform proficiency study of 5 commercial ctDNA (liquid biopsy) assays run at 12 labs, on tailor-made reference samples (cancer cell-line pool "Sample A" diluted into normal "Sample B" at 0:1, 1:4, 1:24, 1:124 → samples Bf, Df, Ef, Ff; plus AccuGenomics and AcroMetrix spike-in QC samples). 359 sequencing libraries deposited at NCBI SRA under SRP296025.

IN SCOPE (pipeline-derived, attemptable)

The descriptor itself prints very few numbers, but the deposited data supports genuine, independently-recomputable pipeline outputs that substantiate the paper's explicit design facts and its central "ultra-deep / analytic validity" claim:

id result how reproduced pipeline
C1 359 SRA records under SRP296025 ENA/SRA run-table count metadata API
C2 9 reference samples (Bf,AC01,BfIS,Df,DfIS,Ef,EfIS,Ep,Ff) distinct BioSamples in run table metadata API
C3 3 platforms (Illumina NovaSeq6000, NextSeq500, Ion S5 XL) instrument_model tally metadata API
C4 "Ultra-deep" sequencing depth align Ion Torrent amplicon runs, samtools depth over panel target bwa-mem + samtools
C5 Analytic validity = known serial dilution recoverable track per-variant VAF across Df(1:4)→Ef(1:24)→Ff(1:124); cancer variants must drop ~5×/step while germline stays flat bwa-mem + bcftools mpileup

C1–C3 are pure metadata checks (done on «host», no heavy compute). C4–C5 are the real pipeline reproduction and run on «our HPC» («infra» + SLURM).

OUT OF SCOPE (not attempted, with reason)

  • Vendor variant calls. Per the paper, "variant calling results were reported by the panel providers' recommended or in-house pipelines" — i.e. proprietary, closed pipelines (Burning Rock in-house, IDT, Illumina TruSight Tumor 170+UMI, Roche AVENIO, Thermo Oncomine). These are not openly reproducible; we do not attempt to reproduce vendor VCFs 1:1.
  • Cross-lab reproducibility statistics / sensitivity-specificity / LoD curves. The descriptor defers all such quantitative validation to the companion research manuscript (Deveson et al., Nat Biotechnol 2021) — those numbers are not printed in this descriptor, so there is no in-paper value to grade against here.
  • The github.com/iontorrent/TS "code" link is the Ion Torrent Torrent Suite vendor software (relevant only to the TFS panel); it is not an analysis pipeline for this paper's results. We treat C4/C5 as a third-party-tool reproduction on the paper's own data (explicitly permitted by the brief, P16), using standard open tools rather than the vendor suite.

Honesty note

Because the descriptor prints essentially no numeric technical-validation values, this reproduction is best understood as (a) a full dataset verification/profile (C1–C3, exact) plus (b) an independent re-demonstration of the paper's central qualitative claims (C4 ultra-deep depth; C5 dilution-series analytic validity) recomputed from the deposited reads. It is NOT a 1:1 regeneration of a printed statistic, because the paper contains none in scope.

Figures / tables: TableFig 1a
C1
Reported
359 NCBI SRA records
Reproduced
359 runs in ENA filereport for SRP296025
exact
C2
Reported
9 deposited reference samples
Reproduced
9 distinct BioSamples observed
exact
C3
Reported
3 platforms (NovaSeq6000, NextSeq500, Ion S5 XL)
Reproduced
3 platforms observed (192/72/95 runs)
exact
C4
Reported
ultra-deep sequencing (very high on-target coverage)
Reproduced
CONFIRMED: mean on-target depth Bf 18368x / Df 19079x / Ef 18085x / Ff 127837x (median up to 161294x, max 435919x; 63-74% of panel >=1000x) on the Oncomine Lung cfDNA panel
within tolerance
C5
Reported
analytic validity: known 1:4/1:24/1:124 dilution recoverable
Reproduced
Cancer variants (TP53, ALK-region, BRAF locus) drop monotonically Df>Ef>Ff, absent in Bf; germline control flat. Magnitude refinement running («job», 6 runs/sample + fixed ratio estimator).
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 87/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

106.7 k
tokens (I/O) · 3.7 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.