Genetic demultiplexing of pooled single-cell RNA-sequencing samples in cancer facilitates effective experimental design.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓No authors-side cause for any deviation
- ✓Any deviation was negligible
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well. Benchmark of genetic demultiplexing tools (cellSNP-lite/Vireo, demuxlet) on pooled cancer scRNA-seq with fully reproducible Snakemake code (github lmweber/snp-dmx-cancer @ec583e0). The ONE pipeline-derived result deposited under open access -- per-sample cell counts (Table 1, HGSOC X2/X3/X4 = 7123/1533/6546) -- reproduces EXACTLY from the deposited GEO count matrix (GSE158937; barcode counts and mtx column counts both match to the cell). The paper's CORE results (Fig 2 demultiplexing precision/recall; Fig 4 runtimes) cannot be reproduced from open data: genetic demultiplexing requires read-level SNP pileups from BAM files, and the raw HGSOC reads are deposited only under dbGaP controlled access (phs002262), the lung data only under EGA controlled access (EGAD00001005054). The deposited GEO matrix contains gene-expression counts with no genotype/read information, so it cannot drive the demultiplexing pipeline. Cost-savings (Fig S4) are hardcoded external 'Cost Per Cell' calculator values (not pipeline-derived; arithmetic ~60% confirmed). An open third-party path exists (healthy iPSC souporcell data, ENA ERS2630502-06) but is a supplementary non-cancer panel requiring hundreds of GB of fastq + multi-day Cell Ranger; not attempted. Verdict: PARTIAL -- open-deposited pipeline output (cell counts) reproduces 1:1 exactly; the main benchmark is honestly blocked by legitimately controlled-access patient genomic data.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 87assessed: 2026-06-18 ⛓ 3423d8f05415
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan genetic variation–based computational demultiplexing tools effectively recover sample identities of pooled single-cell RNA-seq samples in cancer, where somatic mutations (high SNV or high CNV burden) might obscure the natural genetic variation signal used to distinguish individuals?
- ★ Genetic variation–based demultiplexing tools can be effectively deployed on cancer scRNA-seq tissue using a pooled experimental design, achieving high recall at acceptable precision-recall tradeoffs in both high-CNV (HGSOC) and high-SNV (lung adenocarcinoma) cancers, even with extremely high doublet proportions. finding
- ★ cellSNP/Vireo with a genotype reference generated by bcftools from matched bulk RNA-seq is the best-performing tool combination for demultiplexing cancer scRNA-seq samples. finding
- ★ High proportions of ambient RNA from simulated cell debris reduce demultiplexing performance (mainly recall), though the effect is minimized when using the top-performing bulkBcftools_cellSNPVireo combination; demuxlet is more sensitive to ambient RNA than cellSNP/Vireo. finding
- ★ cellSNP/Vireo more reliably identifies true doublets than demuxlet in cancer samples (99.2% vs 31.9% of called doublets being true identifiable doublets), supporting cost-saving super-loading designs. finding
- Pooled library preparation with genetic demultiplexing provides significant cost savings compared to individually prepared libraries. finding
- ★ A reproducible, modular Snakemake workflow built around the best-performing tools is provided for experimental design, available on GitHub. resource
- In silico benchmark simulations constructed by combining raw sequencing reads from multiple single-cell samples with known sample identity can evaluate demultiplexing under varying doublet and ambient-RNA proportions. method
- The 1000 Genomes Project SNP reference (filtered to 3' UTR) with cellSNP/Vireo performs nearly as well as the unfiltered reference while keeping runtimes lower, enabling demultiplexing without matched bulk RNA-seq. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA-seq (3' tag, 10x/Cell Ranger) | high-grade serous ovarian cancer (HGSOC) tissue, 3 samples (X2, X3, X4) | none (in silico pooling of samples with simulated doublets and debris) | precision and recall of recovering sample identity of true singlet cells | Cell Ranger; cellSNP/Vireo, demuxlet, bcftools, samtools |
| single-cell RNA-seq (3' tag) | lung adenocarcinoma tissue, 6 samples (T08, T09, T20, T25, T28, T31) | none (in silico pooling with simulated doublets and debris) | precision and recall of sample identity recovery for singlets | Cell Ranger; cellSNP/Vireo with 1000 Genomes filtered reference |
| matched bulk RNA-seq | HGSOC samples | none | genotype reference list of SNPs for demultiplexing | bcftools |
| in silico doublet simulation | HGSOC and lung adenocarcinoma scRNA-seq | combining raw reads from multiple cell barcodes (no doublets, 20%, 30% doublets) | demultiplexing precision/recall and doublet identification accuracy | — |
| in silico ambient RNA / debris simulation | HGSOC and lung adenocarcinoma scRNA-seq | reassigning all reads from 10%, 20%, or 40% of cell barcodes randomly to other barcodes | effect on demultiplexing recall and precision | — |
| downstream doublet detection | HGSOC and lung scRNA-seq (cellSNP/Vireo, bulk reference) | 20% and 30% doublet scenarios | false-positive/false-negative doublet calls | scDblFinder |
| single-cell RNA-seq (baseline) | healthy non-cancer cell lines | none | baseline demultiplexing performance for comparison | cellSNP/Vireo |
- ▲ cellSNP/Vireo with matched bulk RNA-seq reference (bulkBcftools_cellSNPVireo) achieved highest recall in HGSOC across no/20%/30% doublet scenarios 99.0%, 99.9%, 99.9% recall
- ▼ Precision of bulkBcftools_cellSNPVireo decreased as doublet proportion increased in HGSOC 100%, 85.9%, 77.4%
- ▲ Of cell barcodes called doublets by Vireo (HGSOC, 30% doublets, bulk reference), nearly all were true identifiable doublets, vs demuxlet 99.2% vs 31.9% for demuxlet
- – demuxlet with bulk reference gave higher precision but large reduction in recall (20%/30% doublets) in HGSOC precision 91.3%/84.3%, recall 53.0%/52.1%
- ▲ 1000 Genomes unfiltered reference with cellSNP/Vireo performed well in HGSOC without matched bulk RNA-seq recall 97.9%/99.0%/99.0%; precision 100%/84.0%/74.5%
- ▲ 1000 Genomes filtered (3' UTR) reference showed only minor loss vs unfiltered in HGSOC recall 94.3%/95.0%/95.3%; precision 100%/83.7%/73.8%
- ▼ 10% debris simulation reduced HGSOC recall with bulk reference; demuxlet far more affected bulk ref recall 85.1%/86.2%/86.7%; demuxlet recall 7.8%/6.7%/6.4%
- ▲ Calling SNPs directly from scRNA-seq with cellSNP/Vireo gave comparable recall in HGSOC recall 91.5%/91.9%/92.1%; precision 99.7%/82.3%/72.7%
- other recall 99.0%, 99.9%, 99.9% (no/20%/30% doublets) (bulkBcftools_cellSNPVireo recall, HGSOC, averaged across 3 samples)
- other precision 100%, 85.9%, 77.4% (bulkBcftools_cellSNPVireo precision, HGSOC across doublet scenarios)
- count 99.2% (proportion of Vireo-called doublets that were true identifiable doublets (HGSOC, 30% doublets, bulk reference))
- count 31.9% (proportion of demuxlet-called doublets that were true identifiable doublets (HGSOC, 30% doublets, bulk reference))
- other recall 7.8%, 6.7%, 6.4% (demuxlet recall with 10% debris, bulk reference, HGSOC)
- other >10 or >20 SNVs per megabase (definition of high-TMB cancers)
- other ~1,000 SNPs per megabase (MAF >1%) (population SNP frequency, ~2 orders of magnitude higher than cancer SNVs)
- count HGSOC cells per sample: X2=7,123; X3=1,533; X4=6,546 (number of cells per HGSOC scRNA-seq sample from Cell Ranger)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper performs in silico benchmark evaluations of genetic variation–based demultiplexing algorithms for pooled scRNA-seq cancer data, constructing simulated pooled samples from real single-cell datasets with known per-cell sample identity (3 HGSOC samples; 6 lung adenocarcinoma samples). Performance is characterized descriptively using precision and recall metrics across combinations of demultiplexing tools (Vireo, demuxlet) and SNP reference strategies, under varying simulated doublet proportions (0%, 20%, 30%) and ambient RNA debris proportions (10%, 20%, 40%). No formal statistical hypothesis tests are applied; results are reported as mean precision and recall values averaged across samples within each scenario.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Precision and recall (descriptive performance metrics; no formal statistical hypothesis test applied) | All benchmark comparisons of demultiplexing tool/reference-strategy combinations across HGSOC and lung adenocarcinoma datasets and all simulation scenarios | 3 HGSOC patient samples; 6 lung adenocarcinoma patient samples | na |
-
Precision and recall per scenario are reported as means averaged across 3–6 biological samples, with no accompanying dispersion measure↳ Could also: Report mean ± SD or range of precision/recall across samples, or display individual sample values as points overlaid on the precision-recall plots — With only 3–6 biological replicates, between-patient variability is an important aspect of generalizability; showing spread would let readers gauge how consistently each strategy performs across individuals
-
Comparisons between tool/reference combinations are made by visual inspection of precision-recall plots, without formal statistical testing↳ Could also: Apply a McNemar's test or bootstrap/permutation test to assess whether observed differences in correctly classified cells between strategies exceed chance given the sample sizes — Formal testing would help distinguish differences likely to replicate from those within the noise of a small number of biological replicates, complementing the descriptive summaries
-
Performance is summarized using two separate metrics (precision and recall), requiring readers to judge trade-offs scenario by scenario↳ Could also: Also report the F1 score (harmonic mean of precision and recall) or the area under the precision-recall curve (AUC-PR) as a scalar summary per scenario — A single composite metric simplifies direct ranking of scenarios and is especially useful when precision and recall move in opposite directions across conditions
-
Doublet proportions are evaluated at three fixed levels (0%, 20%, 30%) and debris at three fixed levels (10%, 20%, 40%)↳ Could also: Fit a regression or spline model of precision/recall as a continuous function of doublet or debris proportion — A continuous model would characterize the performance-degradation curve more fully and allow interpolation to proportions not explicitly simulated, potentially informing experimental design thresholds
-
Benchmark datasets consist of 3 HGSOC and 6 lung adenocarcinoma patient samples pooled in silico↳ Could also: Use bootstrap resampling of the available patient samples or leave-one-out cross-validation to quantify uncertainty in the observed performance rankings — Resampling would provide confidence intervals on mean precision/recall without requiring additional experimental data, making the stability of performance rankings across tools more explicit
-
Ambient RNA debris is simulated by reassigning all reads from a fixed percentage of barcodes to other barcodes uniformly at random↳ Could also: Also evaluate performance using computationally estimated ambient RNA profiles (e.g., via SoupX or CellBender) rather than random reassignment, or vary the distribution of reassignment — Real ambient RNA has a non-uniform composition reflecting lysed cells; alternative simulation models would test robustness of the conclusions to different debris-contamination patterns
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34553212
Paper: Weber et al. 2021, GigaScience. "Genetic demultiplexing of pooled single-cell RNA-sequencing samples in cancer facilitates effective experimental design." DOI 10.1093/gigascience/giab062.
Code: https://github.com/lmweber/snp-dmx-cancer @ ec583e038b05ec8c2bb4dcb6c85f476b77dff8c0
(fully reproducible Snakemake workflow + benchmark + evaluation R scripts — authors' own).
What the paper does (computational pipeline)
A benchmark of genetic (SNP-based) demultiplexing tools for pooled scRNA-seq of cancer samples. Pipeline per scenario:
- Cell Ranger align each individual sample's 10x scRNA-seq -> filtered count matrix + position-sorted BAM with cell barcodes.
- In-silico pooling: merge per-sample BAMs/barcodes (suffix encodes the true sample-of-origin = ground truth) and simulate doublets at 0/20/30%.
- Genotype reference built one of several ways: 1000 Genomes (filtered to
3'UTR / unfiltered), matched bulk RNA-seq via
bcftoolsorcellSNP, or directly from scRNA-seq viacellSNP. - Demultiplex:
cellSNP-litepileup + Vireo, or demuxlet (popscle). - Evaluate: precision & recall of recovering true sample-of-origin
(
evaluations/*.R), runtimes (Fig 4), cost savings (Fig S4).
Datasets
| dataset | role | repository | access |
|---|---|---|---|
| HGSOC (3 patients, scRNA + bulk RNA-seq) | main | GEO GSE158937 (processed matrices) + dbGaP phs002262.v1.p1 (raw reads) | matrices OPEN; raw CONTROLLED |
| Lung adenocarcinoma (Kim 2020) | main | EGA EGAD00001005054 | CONTROLLED |
| Healthy iPSC (Heaton/souporcell 2020) | supplementary | ENA ERS2630502–ERS2630506 (fastq) | OPEN |
In scope (pipeline-derived)
- C1 — per-sample cell counts, HGSOC (Table 1). Cell Ranger output; deposited as the GEO count matrix. Reproducible from open data. → DONE, exact.
- C3 — demultiplexing precision/recall (Fig 2). Pipeline-derived (core result).
- C4 — runtimes (Fig 4). Pipeline-derived.
- C5 — per-sample cell counts, lung (Table 1). Pipeline-derived.
Out of scope
- C2 — cost savings (Fig S4). NOT pipeline-derived: the numbers are hardcoded
in
estimated_cost_savings.R, taken from the external Satija Lab "Cost Per Cell" web calculator. Deterministic plot of manual inputs; we verify the ~60% arithmetic but it is not a computation on the paper's data (flagged as not-derivable-from-data). - Wet-lab (10x library prep, sequencing), manual SNP-array supplementary steps.
Feasibility / blockers
- C3, C4, C5 are BLOCKED. Genetic demultiplexing operates on read-level SNP pileups from BAM files. The deposited GEO matrix (GSE158937) holds only gene-expression counts — no genotype/read information — and therefore cannot drive the demultiplexing pipeline. The raw reads needed exist only under controlled access (dbGaP phs002262 for HGSOC; EGA EGAD00001005054 for lung), which require an approved Data Access Request — out of reach here. This is a legitimate, expected restriction for human patient genomic data (the authors DID deposit raw data to dbGaP/EGA and ship fully reproducible code).
- Open third-party path (not attempted): the supplementary healthy-iPSC analysis uses open ENA souporcell data and could in principle drive cellSNP/Vireo end-to-end. It is a supplementary non-cancer panel requiring hundreds of GB of fastq across 5 pooled 10x runs (~32 fastq pairs each) plus multi-day Cell Ranger; the cost/benefit (a supplementary figure on non-cancer data, while the main cancer benchmark stays blocked) did not justify the cluster burn. Documented as available, not done.
Verdict
PARTIAL. The single pipeline-derived result deposited under open access (Table 1 cell counts) reproduces EXACTLY. The paper's core benchmark (Fig 2 / Fig 4) is honestly unreproducible from open data due to controlled-access raw reads — not a code/doc deficiency.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.