Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

TOSCA: an automated Tumor Only Somatic CAlling workflow for somatic mutation detection without matched normal samples.

Bioinform Adv · 2022
L1 50/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

IN PROGRESS. TOSCA (Snakemake tumor-only somatic caller) repo present, pinned commit 14a849e (2023-03-04), pipeline + I/O format fully understood; reproduction targets defined (T-LBL panel variant count 655 + germline/somatic split + pure/hybrid sens-spec; HCC1395 WES sens-spec). PRIMARY data PRJEB36436 is the Muenster T-LBL cohort; paper used a 16-patient subset (10 tumor + 6 normal) of a larger targeted-capture deposit (~200+ runs incl. TG/TP/TR/WES aliases) -> N concordance needs exact count from ENA filereport. BLOCKER right now: «our HPC» login node («host»:22) unreachable -> all heavy compute + reliable ENA download paused, retrying centrally-managed VPN per brief. No result fabricated; counts left null until observed.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ d3de269ea8cf
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can an automated, open-source tumor-only somatic calling workflow reliably distinguish somatic from germline variants in whole-exome and targeted panel sequencing data without a matched normal sample, yielding estimates consistent with matched-normal paired analyses?

Core claims
  • TOSCA is the first automated, open-source, end-to-end tumor-only somatic calling workflow for WES and targeted panel sequencing data. resource
  • TOSCA's tumor-only somatic/germline classifications are consistent with results from matched-normal paired analyses. finding
  • Adding unmatched normal samples ('hybrid' mode) with tumor purity/ploidy and copy-number estimation via PureCN substantially improves classification accuracy over 'pure' tumor-only mode. finding
  • Somatic status is inferred via two approaches: a decision-tree variant filtration strategy (extended from Sukhai et al. 2019) and PureCN-based tumor purity/ploidy estimation. method
  • For clinical tumor-only WES, inclusion of at least a small panel of normal samples should be mandatory because pure tumor-only mode yields very low precision for germline discrimination. finding
  • TOSCA is implemented as a Snakemake workflow with conda environments enabling scalable, reproducible end-to-end analysis usable by non-coders. method
Experimental setups
Assay System Perturbation Readout Platform
Targeted panel sequencing (78-gene capture panel) T-cell lymphoblastic lymphoma (T-LBL) patient cohort, tumor and matched normal (16 patients) none somatic vs germline variant classification (sensitivity/specificity) Agilent Technologies capture-based technique; data from SRA PRJEB36436
Whole-exome sequencing (WES) HCC1395 triple-negative breast cancer cell line and HCC1395BL normal B-lymphocyte cell line (same donor) none somatic SNV detection sensitivity and precision/specificity vs truth set SEQC-II Consortium data; SRA SRP162370
Key results
  • In pure tumor-only mode on T-LBL targeted panel data, somatic and germline variants classified correctly with sensitivity 91% and specificity 88%. sensitivity 91%, specificity 88%
  • In hybrid mode on T-LBL data, TOSCA reached 96% sensitivity and 96% specificity, increasing accuracy by 4% and 7% respectively over pure mode. 96% / 96%; +4% / +7%
  • TOSCA produced somatic vs germline calls for 100% of variants in T-LBL samples. 100%
  • In pure tumor-only WES mode (HCC1395), SNV detection showed lowest sensitivity and specificity of 69% and 75%. sensitivity 69%, specificity 75%
  • In hybrid WES mode incorporating HCC1395BL pool of normals and copy-number profile, somatic/germline classified with 95.7% and 94.9% accuracy. 95.7% / 94.9%
  • External validation: Sukhai et al. filtration algorithm achieved ~100% sensitivity and ~95.5% specificity on a 48-gene panel, and 99.6%/86.9% on a 555-gene panel. 48-gene: 100%/95.5%; 555-gene: 99.6%/86.9%
  • External validation: PureCN (Oh et al. 2020) achieved sensitivity/specificity of 96.1%/88.1% and 97.2%/96.6% on two WES datasets. 96.1%/88.1%; 97.2%/96.6%
Key statistics
  • count 655 unique variants detected; 57% (N=376) germline and 43% (N=279) somatic (T-LBL targeted panel gold standard variants)
  • count 1089 SNVs and 57 Indels (SEQC-II high confidence truth set for WES validation)
  • other sensitivity 91%, specificity 88% (pure tumor-only mode, T-LBL targeted panel)
  • other sensitivity 96%, specificity 96% (hybrid mode, T-LBL targeted panel)
  • other sensitivity 69%, specificity 75% (pure tumor-only mode, WES HCC1395)
  • other 95.7% and 94.9% (hybrid mode WES accuracy for somatic and germline)
  • count 16 patients; 10 tumor for calling, 6 normals for normal database (T-LBL cohort design)
  • count 12 high coverage sequencing replicates each for HCC1395BL and HCC1395 (SEQC-II WES dataset replicates)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a bioinformatics tool/method paper describing a computational workflow (TOSCA) for tumor-only somatic variant calling, validated by comparing its variant classifications against a 'gold standard' derived from matched tumor-normal sequencing in two independent datasets (a targeted panel cohort and a WES reference cohort). Performance was reported descriptively using sensitivity, specificity, and precision point estimates for each analysis mode ('pure' tumor-only vs 'hybrid' with unmatched normals), without inferential hypothesis testing, p-values, or confidence intervals.

Replicationunclear Sample sizeTargeted-sequencing validation used 16 T-LBL patients (10 tumor samples for tumor-only calling, 6 normals for a reference database), yielding 655 unique variants (376 germline, 279 somatic per gold standard). WES validation used the SEQC-II HCC1395/HCC1395BL reference pair with 12 high-coverage sequencing replicates each, evaluated against a high-confidence truth set of 1089 SNVs and 57 indels. GroupsTOSCA tumor-only calls ('pure' mode) vs TOSCA hybrid-mode calls (using unmatched normals) vs a gold-standard/truth set derived from matched tumor-normal sequencing Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Approaches that could also have been used
  • Sensitivity, specificity, and precision are reported as single point-estimate percentages for each dataset and mode.
    Could also: Binomial or Wilson/Clopper-Pearson confidence intervals around these proportions — Since these metrics are based on counts of variants classified correctly or incorrectly, a confidence interval would communicate the precision of each estimate, particularly useful given the moderate sample sizes (e.g. 655 variants, 57 indels).
  • Performance of 'pure' tumor-only mode is compared descriptively to 'hybrid' mode by reporting the percentage-point increase in sensitivity and specificity.
    Could also: A paired statistical comparison such as McNemar's test on the paired classification outcomes (since both modes are applied to the same variants/samples) — A paired test would formally quantify whether the difference in classification accuracy between the two modes is unlikely to be due to chance, complementing the raw percentage-point comparison already reported.
  • Variant classification is evaluated using fixed thresholds (e.g. variant allele fraction, database allele-frequency cutoffs) with resulting single sensitivity/specificity pairs.
    Could also: ROC or precision-recall curve analysis across a range of thresholds, summarized with an AUC — This would show how the sensitivity/specificity trade-off varies with threshold choice, which can be informative for users tuning the workflow to their own accuracy priorities.
  • The gold standard for the WES dataset is itself derived from consensus of multiple callers/replicates rather than an independently validated ground truth.
    Could also: Bootstrap resampling across the 12 sequencing replicates to estimate variability in the sensitivity/specificity estimates — Bootstrapping over replicates could give an empirical sense of estimate stability without requiring additional data collection.
  • Comparisons between the panel-based (Sukhai et al.) and copy-number-based (PureCN) filtration strategies are described narratively by citing each method's previously published sensitivity/specificity figures.
    Could also: A formal meta-analytic or statistical synthesis (e.g. combining sensitivity/specificity across studies with corresponding intervals) — This could provide a pooled, quantitative comparison across the cited external validation studies rather than a narrative side-by-side listing of percentages.
Software: Snakemake · PureCN (R package) · MultiQC

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36699358 (TOSCA)

Paper: Del Corvo M, Mazzara S, Pileri SA. TOSCA: an automated Tumor Only Somatic CAlling workflow for somatic mutation detection without matched normal samples. Bioinform Adv 2022; 2(1):vbac070. PMID 36699358 / PMC9710689.

Repo (third-party-own = authors' own): https://github.com/mdelcorvo/TOSCA default branch main, pinned commit 14a849e39a0b3ea052c8fe11e3602faaf856efe0 (2023-03-04). Snakemake + conda workflow.

What TOSCA is

A Snakemake pipeline for somatic SNV/indel calling from tumor samples without a matched normal. Two modes:

  • Pure tumor-only mode (mandatory rules): align → call → 3-phase filtration (quality → germline-DB filter at 1% MAF + COSMIC marking → ClinVar benign removal) following Sukhai et al. 2019 decision tree.
  • Hybrid mode (optional, enabled by a normal-panel metafile meta_N, ≥5 normals): adds PureCN (R) for purity/ploidy + unmatched-normal PoN.

Final targets (Snakefile rule all):

  • results/Somatic_Prediction.txt ← the variant-call result we compare
  • results/Report.html (MultiQC)
  • PureCN/{sample}.rds (hybrid only)

Input = metafile TSV sample platform fq1 fq2; config = config/config.yaml (genome GRCh38 release 104, min depth 250x, min VAF 5%, target BED, snpEff db, germline/somatic DB URLs). Repo ships ExAC + 1000G PoN under resources/.

In scope (pipeline-derived, attempted)

id reported result paper loc pipeline tractability
TLBL-nvar 655 unique variants total Results (T-LBL TS) TOSCA pure runs the pipeline; count variants
TLBL-split 57% germline (N=376) / 43% somatic (N=279) Results TOSCA pure classification output
TLBL-pure pure tumor-only: sensitivity 91%, specificity 88% Results TOSCA pure needs labeled truth
TLBL-hybrid hybrid (+unmatched normals): sensitivity 96%, specificity 96% Results TOSCA hybrid needs ≥5 normals + truth
HCC-truth truth set 1089 SNV + 57 indel (SEQC-II) Results (WES) external truth downloadable
HCC-pure pure: SNV sensitivity 69%, specificity 75% Results TOSCA pure needs SEQC-II FASTQ + truth
HCC-hybrid hybrid: SNV sensitivity 95.7%, specificity 94.9% Results TOSCA hybrid heavy WES compute

Primary target (quick minimum): get TOSCA pure mode to RUN on the T-LBL targeted-panel data (PRJEB36436) and produce Somatic_Prediction.txt; compare the variant count / germline-somatic split to the reported 655 / 376 / 279. Panel data is small → most tractable. Then push toward sens/spec and the hybrid mode and the HCC1395 WES result.

Datasets

  • PRJEB36436 (ENA/SRA) — T-LBL targeted-panel; paper: 16 patients (10 tumor analyzed tumor-only + 6 normals for coverage-normalization DB). PRIMARY data.
  • SRP162370 (SEQC-II) — HCC1395 / HCC1395BL WES, 12 high-cov replicates each. Large WES; secondary.
  • SEQC-II somatic truth set (NIST/GIAB-adjacent) — 1089 SNV + 57 indel high-confidence somatic truth for HCC1395; needed for HCC sens/spec.

Out of scope (not pipeline-derived)

  • External-literature comparison numbers quoted from Sukhai et al. 2019 (100% / 95.5%; 99.6% / 86.9%) and Oh/PureCN 2020 (96.1% / 88.1%; 97.2% / 96.6%) — these are other papers' results cited for context, not produced by TOSCA here.
  • Wet-lab / clinical interpretation.

Known blockers / risks

  • COSMIC download requires registration/login (email+password) — the DB annotation/marking step may need a COSMIC VCF. Mitigation: COSMIC only marks variants (it does not filter them out in phase-2); pipeline may still produce calls without it, or we supply COSMIC via authenticated download.
  • Truth labels for T-LBL (which of the 655 are truly germline vs somatic) come from the original cohort study's matched analysis — needed for sens/spec; the variant COUNT and germline/somatic split are reproducible from pipeline output
    • DB filter
TLBL-nvar
Reported
655 unique variants (T-LBL targeted panel)
Reproduced
partial
TLBL-split
Reported
57% germline (376) / 43% somatic (279)
Reproduced
partial
TLBL-pure
Reported
pure tumor-only sensitivity 91% / specificity 88%
Reproduced
partial
TLBL-hybrid
Reported
hybrid sensitivity 96% / specificity 96%
Reproduced
partial
HCC-pure
Reported
HCC1395 WES pure SNV sens 69% / spec 75%
Reproduced
partial
HCC-hybrid
Reported
HCC1395 WES hybrid SNV sens 95.7% / spec 94.9%
Reproduced
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

97.9 k
tokens (I/O) · 5.1 M incl. cache
16 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.