Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Evaluating genome sequencing strategies: trio, singleton, and standard testing in rare disease diagnosis

Genome Medicine · 2025
L1 88/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
88/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 74% of all assessed papers rank 276 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction, described well enough at the summary layer but NOT at the pipeline layer. The paper is a blinded prospective CLINICAL diagnostic-yield study; its headline result is the product of (a) the proprietary licensed Illumina DRAGEN variant-calling appliance and (b) manual ACMG/AMP/ACGS classification by blinded expert teams in proprietary software (TruSight, Alissa) — neither of which is reproducible here. The raw sequencing data (the true pipeline input) are controlled/on-request with no public accession (data_restricted), and DRAGEN/TruSight/Alissa are unobtainable proprietary software (env_unresolvable). HOWEVER, the six headline diagnostic-yield figures — the central quantitative claims (tGS/sGS/SoC, prospective + retrospective) — were recomputed EXACTLY from the deposited supplementary per-case data (Additional file 1, Table S3): all six numerators (incl. the fractional 96.5/117.5) and rounded percentages reconcile to the exact reported values, over the 335-family trio cohort with denominator 335. No fabrication detected — the reported numbers are fully traceable to the shipped data. NOT attempted: independent re-derivation of the variant calls and ACMG classifications (restricted data + proprietary/manual). Verdicts provisional; a human reviewer signs off (AUDIT.md).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 88
    assessed: 2026-06-18 ⛓ 41baae11348e
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether short-read genome sequencing (as singleton GS or trio GS) can serve as a superior, unifying first-tier ('one-test-for-all') diagnostic strategy compared to standard-of-care exome-based testing for patients with rare diseases.

Core claims
  • Trio genome sequencing (tGS) achieves higher prospective diagnostic yield than standard-of-care (SoC) and singleton genome sequencing (sGS) even when performed by a newly trained team. finding
  • Retrospective analysis under uniform conditions shows GS (sGS and tGS) technically outperforms SoC in diagnostic yield. finding
  • GS detects deep intronic, non-coding, and small copy-number variants that are missed by SoC methods. finding
  • Trio approach is particularly valuable for less experienced analysis teams because inheritance data aids variant interpretation. finding
  • Experienced teams can achieve comparable diagnostic yield using singleton analysis alone, supporting sGS as a more cost-effective alternative. finding
  • tGS identified three de novo variants classified as likely pathogenic based on recent GeneMatcher collaborations and newly published gene-disease association studies. finding
  • GS should be considered a first-tier genetic test to reduce the diagnostic odyssey for rare disease patients. resource
Experimental setups
Assay System Perturbation Readout Platform
short-read genome sequencing (GS) 1011 individuals (patients and relatives) from 416 index rare disease cases, blood-derived DNA none SNVs, CNVs, mitochondrial variants, repeat expansions, structural variants; diagnostic P/LP classification Illumina NovaSeq 6000, DRAGEN Germline Pipeline v3.7.5, Manta, ExpansionHunter, TruSight Software Suite v2.6
exome sequencing (ES) index/mother/father blood samples, rare disease patients none coding-region variants for diagnostic P/LP classification Illumina NovaSeq 6000, Nextera Flex for Enrichment kit, xGen Exome Research Panel v2, IKMB DRAGEN Pipeline v3.10.4, Alissa Interpret v5.4.2
array comparative genomic hybridisation (array-CGH) patient blood samples none copy-number variants (CNVs) Agilent SurePrint G3 Human CGH Microarray Kit (8x60K), Genomic Workbench v7.0
karyotyping patient blood metaphase samples none chromosomal abnormalities (aneuploidies, translocations, deletions, duplications) G-banded trypsin-Giemsa staining, NEON interface v1.3 (MetaSystems)
Key results
  • Prospective diagnostic yield for P/LP variants: tGS 36.1%, SoC 35.1%, sGS 28.8% tGS 36.1% vs SoC 35.1% vs sGS 28.8%
  • Retrospective diagnostic yield: SoC 36.7%, sGS 39.1%, tGS 40.0% tGS 40.0% vs sGS 39.1% vs SoC 36.7%
  • GS superior yield attributed to detection of deep intronic, non-coding, and small copy-number variants missed by SoC
  • tGS identified three de novo likely pathogenic variants via GeneMatcher collaborations and new gene-disease associations 3 variants
  • Sequencing repeated for 68 of 1148 samples due to technical issues or low quality 6.55%
Key statistics
  • count 416 index cases / 1011 genomes included; 1148 individuals recruited and sequenced; 32 cases excluded (study cohort composition)
  • fold_change 36.1% (tGS) vs 35.1% (SoC) vs 28.8% (sGS) (prospective diagnostic yield)
  • fold_change 36.7% (SoC) vs 39.1% (sGS) vs 40.0% (tGS) (retrospective diagnostic yield)
  • count 81 singletons, 51 duos, 258 trios, 11 quartets, 1 quintet, 1 sextet (family structure composition of cohort)
  • mean GS mean coverage 38x (anticipated minimum 29x) (genome sequencing depth)
  • mean ES mean sequencing depth index/mother/father: 202.8x/159.1x/155.2x (exome sequencing depth)
  • other 67.1% male; 84.8% of patients below 20 years of age; 279 male and 137 female patients (cohort demographics)
  • count 68 samples (6.55%) required repeat sequencing (GS sample quality control)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This prospective, blinded cohort study compared diagnostic yields of three genetic testing strategies (singleton GS, trio GS, and exome-based SoC) in 416 rare-disease index cases. Diagnostic yield was defined as a binary outcome (resolved/unresolved) and reported as percentages separately for a prospective stage (each team's real-time performance) and a retrospective stage (uniform re-evaluation to assess technical detectability independent of team experience). Results were descriptively presented as proportions; the statistical analysis section is truncated in the provided text, so inferential test details are not recoverable from the supplied excerpt.

Replicationbiological Sample size416 index cases meeting inclusion criteria described narratively; no formal sample size or power calculation mentioned in the provided text GroupssGS vs tGS vs exome-based SoC, across 416 rare-disease index patients; 81 singletons and 335 multi-member families Pairingmixed Randomization/blindingstated Dispersionnone Effect sizesno
Statistical tests used
Test Applied to n Assumptions
Not recoverable — statistical analysis section truncated in supplied text Comparison of diagnostic yields (P/LP rate) across sGS, tGS, and SoC, prospective and retrospective stages 416 index cases (81 singletons, 335 trio-family cases) not stated
Approaches that could also have been used
  • Diagnostic yields were compared across three groups (sGS, tGS, SoC) as proportions, but the inferential test(s) are not visible in the supplied excerpt
    Could also: A chi-square test or Fisher's exact test could compare pairwise proportions of resolved cases; given that all 416 patients contributed data to the SoC and sGS comparisons simultaneously, a correlated-proportions approach (e.g., McNemar's test for paired binary outcomes within the same patient, or a generalised estimating equations model) would also account for the within-patient dependency — Because each index patient was evaluated by both the SoC and sGS strategies, the two diagnostic outcomes are correlated within the same individual; methods that acknowledge this pairing may produce narrower confidence intervals and more precise inference than independent-samples tests
  • Diagnostic yield was summarised as point-estimate percentages (e.g., 36.1%, 28.8%) with no confidence intervals reported
    Could also: 95% confidence intervals for each proportion (e.g., Wilson or Clopper-Pearson) and for the difference or ratio between proportions could also be reported alongside the point estimates — Confidence intervals convey the precision of each yield estimate given the sample size, allowing readers to judge whether observed differences (e.g., 36.1% vs 28.8%) are compatible with a range of true effect magnitudes, which is especially informative in comparisons with moderate n
  • Three pairwise comparisons of diagnostic yield (sGS vs SoC, tGS vs SoC, tGS vs sGS) are implicit in the results, but no multiplicity correction is mentioned
    Could also: A Bonferroni correction, Holm step-down procedure, or a single omnibus test (e.g., Cochran Q for three correlated proportions) could also be applied to control the family-wise error rate across the three comparisons — When multiple pairwise tests are conducted on the same outcome, some form of multiplicity adjustment is a standard option; applying one would make the reported significance thresholds directly interpretable as experiment-wise error rates
  • No formal sample size or power calculation is described; the enrolled n appears to reflect routine clinical intake during the recruitment window
    Could also: A prospective power calculation targeting a minimum detectable difference in diagnostic yield between strategies (e.g., 10 percentage-point difference at 80% power) could also have been used to pre-specify the recruitment target — A priori power calculations help readers assess whether the study was sized to detect clinically meaningful differences; they are also increasingly requested by journals and funding bodies for prospective diagnostic accuracy studies
  • Two analytical stages (prospective real-time and retrospective uniform re-evaluation) were used to separate team-experience effects from technical detectability, but no statistical model explicitly adjusts for team experience as a covariate
    Could also: A logistic regression model with team (or experience level) and sequencing strategy as co-predictors could also estimate the independent contribution of strategy while statistically controlling for inter-team differences in a single model — Modelling both factors simultaneously would quantify how much of the yield difference attributable to strategy persists after accounting for experience, rather than relying on the two-stage descriptive decomposition alone
  • Disease category (e.g., developmental disorder, neurological syndrome) was used to describe the cohort but does not appear to be included as a factor in the primary yield comparisons
    Could also: A stratified analysis or an interaction term in a regression model could also examine whether the relative advantage of tGS or sGS over SoC varies by disease category — Diagnostic yield is known to vary substantially by phenotypic group in rare-disease genomics; a stratified or interaction analysis would indicate whether one strategy is consistently superior across categories or whether the advantage is concentrated in specific subgroups
Software: DRAGEN Bio IT Platform (variant calling) v3.7.5 · Manta Structural Variant Caller · ExpansionHunter · TruSight Software Suite (GS interpretation) v2.6 · Alissa Interpret (ES interpretation) v5.4.2 · Genomic Workbench (array-CGH analysis) v7.0 · IGV (variant visualization) 2.13.2–2.17.4 · Statistical analysis software: not stated in available text

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 40963120

Title: Evaluating genome sequencing strategies: trio, singleton, and standard testing in rare disease diagnosis Journal: Genome Medicine (2025) · DOI: 10.1186/s13073-025-01516-7 · PMCID: PMC12445032 Design: Blinded, prospective clinical diagnostic study. 416 index cases (335 trio families + 81 singleton patients; 1,011 individuals total) with rare disease. Three strategies compared: trio genome sequencing (tGS), singleton GS (sGS), exome/standard-of-care (SoC). Primary outcome = diagnostic yield (% cases with a P/LP molecular diagnosis), prospective and retrospective.

Reported headline results (Table S3 / abstract)

Strategy Prospective P/LP Retrospective P/LP
tGS 121/335 = 36.1% 134/335 = 40.0%
sGS 96.5/335 = 28.8% 131/335 = 39.1%
SoC 117.5/335 = 35.1% 123/335 = 36.7%
(Fractional numerators arise from cases counted as partial/0.5 in the source table.)

Pipeline vs manual — what is in scope

OUT OF SCOPE — manual / wet-lab / proprietary (NOT a reproducible pipeline)

The headline diagnostic yields are the product of manual ACMG/AMP/ACGS variant classification by blinded expert clinical teams using proprietary interpretation software (TruSight Software Suite v2.6 for GS; Alissa Interpret v5.4.2 for ES), GeneMatcher, OMIM/ClinVar/Decipher lookups, and literature curation. These are manual clinical-interpretation outputs, not deterministic pipeline outputs → non_pipeline. They cannot be regenerated by re-running code.

Upstream pipeline (named, but NOT reproducible here)

  • Alignment + variant calling: Illumina DRAGEN Bio-IT Platform v3.7.5 (GS) and IKMB DRAGEN pipeline v3.10.4 (ES). Repo: https://github.com/ikmb/dragen-variant-calling (Apache-2.0, Nextflow wrapper, master @ f6fd797016248eff8395f61a92753dade944b884). → Requires the proprietary, licensed Illumina DRAGEN appliance (hardware/ software; see --dragen_unit_cost, DRAGEN reference hashtables, Illumina support-docs links in docs/development.md). Not available on «our HPC», not obtainable open-source → env_unresolvable.
  • SV/CNV: Manta; repeat expansions: ExpansionHunter (both open-source) — but they consume DRAGEN/BAM outputs derived from the restricted raw reads.

Input data — NOT obtainable

Raw sequencing reads for the 1,011 individuals are controlled / on-request only:

"Raw sequencing data are not publicly available due to ethical and legal restrictions but may be available upon request with appropriate approvals." No public accession (no EGA/dbGaP/SRA). Clinically relevant variants were submitted to ClinVar (no accession list given). → data_restricted.

Decision (updated after the audit below)

Outcome = PARTIAL. The upstream pipeline-derived result (DRAGEN variant calls) is NOT independently reproducible, but the paper's six headline diagnostic-yield numbers were reproduced EXACTLY from the deposited supplement (see "What we CAN audit" — all six match to the exact numerator and rounded %, no fabrication).

No upstream pipeline-derived result is independently reproducible:

  1. data_restricted — the genomes (pipeline input) are not publicly obtainable.
  2. env_unresolvable — the variant-calling pipeline needs the proprietary licensed Illumina DRAGEN appliance; interpretation needs proprietary TruSight/Alissa.
  3. non_pipeline — the headline yields are manual expert ACMG classification.

Primary drop_reason: data_restricted (first hard blocker; even a fully open pipeline could not run without the genomes). Secondary: env_unresolvable + non_pipeline.

What we CAN audit (and do, below)

The supplementary Table S3 reports the yield numerators/denominators. We download the supplement («infra») and verify the internal arithmetic consistency of the reported percentages (anti-fabrication cross-check). This audits the reported numbers; it is NOT a reproduction of the pipeline.

Figures / tables: Table
yield_prosp_tGS
Reported
121/335 = 36.1%
Reproduced
121.0/335 = 36.12%
exact
yield_prosp_sGS
Reported
96.5/335 = 28.8%
Reproduced
96.5/335 = 28.81%
exact
yield_prosp_SoC
Reported
117.5/335 = 35.1%
Reproduced
117.5/335 = 35.07%
exact
yield_retro_tGS
Reported
134/335 = 40.0%
Reproduced
134.0/335 = 40.00%
exact
yield_retro_sGS
Reported
131/335 = 39.1%
Reproduced
131.0/335 = 39.10%
exact
yield_retro_SoC
Reported
123/335 = 36.7%
Reproduced
123.0/335 = 36.72%
exact
cohort_total
Reported
416 index cases
Reproduced
416 included (448 ID rows, 32 Excluded)
exact
cohort_trio
Reported
335 trio families
Reproduced
336 trio-applicable rows; denom 335 exact
within tolerance
variant_calling_pipeline
Reported
Illumina DRAGEN v3.7.5 / v3.10.4
Reproduced
not attempted (proprietary appliance + restricted raw data)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 88/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

218.7 k
tokens (I/O) · 19.9 M incl. cache
35 min
runtime · 0 CPU-h
0 GB
peak RAM
2
HPC jobs
hummel
machine