Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Estimating and Correcting for Off-Target Cellular Contamination in Brain Cell Type Specific RNA-Seq Data.

Front Mol Neurosci · 2021
L1 50/100 PQI 83
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Paper describes the method well, but the shipped analysis code (github.com/jsicherman/sct-contamination @ 68c21dc; NOT STAR as the brief's code_url stated) is not runnable end-to-end. Scope = the only self-contained result, the SYNTHETIC benchmark = Table 1 (FPF/FNF/AUROC, corrected vs uncorrected), reproduced on the tasic2016data reference (paper used the un-shipped Allen-2018 transcrip.tome). PARTIAL: I reproduced the full environment (R 4.0.5 + Seurat 3.2.3 [v3 required] + MAST + DESeq2 + compcodeR + tasic2016data on «our HPC»; spatstat pinned to 1.64) and the upstream pipeline (as_Seurat -> simplify_reference/MAST markers -> generate_prior -> run_synthetic generated the synthetic contaminated data). But NO Table-1 number was obtained: every condition fails at the DESeq step with 'NA values are not allowed in the count matrix', because run_synthetic appends marker-gene rows keyed by gene name that are DUPLICATED across clusters, leaving NA rows that propagate into the reconstructed counts. The driver also calls get_DEGs(), a function NEVER DEFINED in the repo (aliased to do.DEG). No random seeds are set anywhere, so even when patched it is not exactly reproducible. NOT a fabrication signal -- a code-completeness/robustness problem; the paper's numbers are plausible but could not be confirmed or refuted from the shipped code. NOT ATTEMPTED (the hard 20%): the real LCM-seq/TRAP-seq/manual-sort results (contamination coefficients 0.65/1.35/2.21; per-cell-type DE-gene counts; AIC 64-67%) require external data the repo does not ship and loads via hard-coded local paths (Allen transcrip.tome, counts_exon.rds, qc.tsv for GSE119183/GSE145521/GSE141337/GSE141464/GSE74985), also with no seeds. A one-line dedupe patch on the marker rows would likely let the synthetic benchmark complete; not applied because finalization was requested.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-15 ⛓ a2677f0c9353
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Single-cell type RNA-seq (sctRNA-seq) datasets from brain tissue are susceptible to off-target cellular contamination from surrounding cell types, and a computational reference-based approach can estimate this contamination and, when used as covariates, improve downstream differential expression analyses.

Core claims
  • A computational method using high-quality scRNA-seq reference data can estimate per-sample, per-cell-type off-target contamination coefficients in sctRNA-seq datasets. method
  • Most brain LCM-seq sctRNA-seq datasets show measurable off-target mRNA contamination from surrounding cell types, whereas scRNA-seq reference/control data show little to none. finding
  • Including contamination coefficients as covariates in differential expression analysis improves model quality (lower gene-wise AIC) and typically yields more differentially expressed genes. finding
  • Interneuron subtypes (PV, SST, VIP) are contaminated by pyramidal cells, while pyramidal samples show less contamination, consistent with larger/more numerous pyramidal neurons contributing more off-target mRNA. mechanism
  • Contamination coefficients are scaled against expected biological cell-type similarity so that values near 1.0 indicate no contamination and values >1.0 indicate contamination. method
  • The method provides a flexible resource for detecting and controlling off-target cell type contamination in sctRNA-seq datasets. resource
Experimental setups
Assay System Perturbation Readout Platform
LCM-seq (single-cell type RNA-seq) mouse brain neurons (pyramidal cells; PV, SST, VIP interneurons), Aging dataset (GSE119183) aging (case/control condition) cell type-specific gene expression / marker expression and contamination coefficients STAR v2.7.5 alignment to mm10_ensembl98
LCM-seq (single-cell type RNA-seq) mouse brain neurons (pyramidal cells; PV, SST, VIP interneurons), Stress dataset (GSE145521) stress (case/control condition) cell type-specific gene expression / marker expression and contamination coefficients STAR v2.7.5; FISH used for cell labeling prior to LCM
scRNA-seq (reference) mouse whole cortex and hippocampus, Allen Institute SMART-seq (2019), Tasic et al. 2018 none per-cell-type expression centroids and marker genes SMART-seq
scRNA-seq (control validation) mouse brain, alternative dataset (Tasic et al. 2016) none scaled contamination coefficients (expected near 1.0)
Marker gene determination / differential expression on scRNA-seq Allen Brain single cell data none cell type specific marker genes Seurat v3.1.5 FindAllMarkers with MAST v1.12.0
Differential expression analysis LCM-seq sctRNA-seq case/control samples condition (aging/stress) +/- contamination covariates differentially expressed genes, model AIC, RRHO overlap DESeq2 v1.26.0; RRHO v1.26.0
Synthetic data simulation in silico synthetic cell type clusters simulated contamination and confounding bias by condition DEG recovery / DE analysis performance
Key results
  • Pyramidal marker Slc17a7 showed considerable expression in PV, SST, and VIP interneuron LCM-seq samples, indicating off-target pyramidal contamination
  • Control scRNA-seq dataset (Tasic 2016) contamination coefficients centered at or near 1.0, indicating little/no contamination ~1.0
  • Each LCM-seq sctRNA-seq sample showed contamination from multiple off-target cell types (pyramidal, oligodendrocytes, astrocytes)
  • SST interneurons in pyramidal/other samples contributed little off-target contamination (low Sst expression in other types)
  • Pyramidal cell samples showed less off-target contamination than profiled interneuronal cell types
  • Contamination level was similar across the two LCM-seq datasets (e.g., SST samples most contaminated by astrocytes, pyramidal, oligodendrocytes regardless of dataset)
  • Including contamination covariates increased model quality and typically yielded more differentially expressed genes
Key statistics
  • fold_change at least two-fold difference in expression (FDR 0.01) for retained marker genes (marker gene selection threshold in MAST)
  • other at least 40% difference in fraction of detection between populations (stringent marker gene on/off threshold)
  • other 60% of marker genes subsampled, 10,000 iterations (per-sample correlation confidence estimation)
  • other FDR threshold of 0.1 (DESeq2 differentially expressed gene selection)
  • count 20,000 synthetic genes with 3,000 differentially expressed (synthetic data generation)
  • other r^2 = 1 (synthetic pure clusters correlate perfectly with reference cell type centroids)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper presents a computational pipeline to estimate off-target cellular contamination in single-cell type RNA-seq (sctRNA-seq) datasets by computing Pearson correlations between sctRNA-seq sample marker-gene profiles and collapsed single-cell reference centroids, with empirical confidence intervals derived from 10,000 bootstrap iterations over 60% subsamples of marker genes. Marker genes were identified using MAST hurdle-model differential expression testing within Seurat, filtered at FDR 0.01 and a ≥2-fold / ≥40% detection-fraction threshold. Differential expression of case/control contrasts was performed with DESeq2 (FDR < 0.1) in models with and without contamination covariates; gene-wise model fit was compared by AIC, and DEG-list overlap between the two models was assessed via hypergeometric p-values using the RRHO package.

Replicationbiological Sample sizeSample sizes for the two LCM-seq datasets (GSE119183, GSE145521) are not stated in the excerpt; synthetic data comprised 20,000 genes (3,000 DEGs); reference scRNA-seq n not stated GroupsCase/control conditions (aging, stress) across brain cell types (pyramidal cells, SST, PV, VIP interneurons) in mouse LCM-seq; scRNA-seq reference dataset versus LCM-seq test datasets as a secondary comparison Pairingunclear Randomization/blindingnot stated Dispersionunclear Effect sizesyes Confidence intervalsyes Multiplicity correctionBenjamini-Hochberg FDR (DESeq2 default; not explicitly named in text)
Statistical tests used
Test Applied to n Assumptions
MAST hurdle model (via Seurat v3.1.5 FindAllMarkers) Identification of cell-type-specific marker genes from single-cell Allen Brain Institute reference data (mouse whole cortex and hippocampus SMART-seq 2019) not stated
Pearson correlation (bootstrapped; 10,000 iterations, 60% marker gene subsample per iteration) Per-sample estimation of off-target contamination coefficients: each sctRNA-seq sample correlated against reference cell-type centroids not stated
DESeq2 Wald test (negative binomial GLM) Differential expression analysis in LCM-seq Aging (GSE119183) and Stress (GSE145521) datasets, run separately with and without contamination covariates not stated
Hypergeometric test (RRHO v1.26.0, R) Assessment of DEG-list overlap between the contamination-covariate model and the baseline model not stated
Akaike Information Criterion (AIC) Gene-wise comparison of DESeq2 model fit with versus without contamination covariates na
Approaches that could also have been used
  • Pearson correlation was used to score transcriptional similarity between each sctRNA-seq sample and reference cell-type centroids
    Could also: Spearman rank correlation or cosine similarity could also score this similarity — Spearman correlation is robust to non-normal expression distributions and outlier marker genes; cosine similarity is common in single-cell deconvolution pipelines and does not assume linearity — either would produce a comparable contamination-scoring framework without the normality assumption implicit in Pearson r
  • MAST hurdle model (via Seurat FindAllMarkers) was used to identify cell-type-specific marker genes from scRNA-seq reference data
    Could also: Wilcoxon rank-sum test, edgeR likelihood ratio test, or logistic regression could also identify cell-type markers from scRNA-seq data — Benchmarking studies show that Wilcoxon-based approaches are particularly robust to zero-inflation common in scRNA-seq; edgeR captures overdispersion in count data; the choice of marker-detection method can influence which genes enter the contamination-scoring step and thus affect downstream coefficient estimates
  • AIC was computed gene-wise to compare DESeq2 models with and without contamination covariates
    Could also: DESeq2's built-in likelihood ratio test (LRT) or the Bayesian Information Criterion (BIC) could also compare nested models — DESeq2's LRT directly tests whether adding covariates significantly improves fit and yields a p-value per gene that can itself be FDR-corrected; BIC applies a stronger complexity penalty than AIC and is often preferred when sample sizes are moderate, making it a complementary perspective on the same model-selection question
  • An FDR threshold of 0.1 was used to define differentially expressed genes in DESeq2
    Could also: An FDR threshold of 0.05 is also a standard cutoff widely used in transcriptomics — The choice of threshold reflects a sensitivity–specificity tradeoff; FDR 0.05 is the predominant convention in many genomics fields and would yield a more conservative DEG set, which could be useful when comparing the two models, since differences in list size at 0.1 versus 0.05 can themselves reveal differential model sensitivity
  • Hypergeometric p-values (RRHO) were used to assess overlap between DEG lists from models with and without contamination covariates
    Could also: Fisher's exact test or the Jaccard similarity index could also quantify two-list overlap — Fisher's exact test is the most widely used, easily interpretable method for two-list overlap and does not require the RRHO rank-rank framework; the Jaccard index provides a symmetric, threshold-free similarity metric and is robust to variation in list size between the two models
  • Bootstrap subsampling of 60% of marker genes over 10,000 iterations was used to estimate variability in contamination coefficients
    Could also: Jackknife resampling over samples, or leave-one-out cross-validation, could also assess the stability of contamination estimates — Sample-level resampling (rather than gene-level subsampling) would capture inter-sample variability in contamination scores directly, which may be more relevant when these scores are subsequently used as per-sample covariates in differential expression models
Software: STAR 2.7.5 · Seurat 3.1.5 · MAST 1.12.0 · DESeq2 1.26.0 · R/RRHO 1.26.0 · R

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
9
Impact: low
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 75/100
stands on reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE115746 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE119183 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE141337 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE141464 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE145521 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE71585 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

Downstream reach in the literature

57 downstream papers · 6 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

GSE145521 GEO reused by 2 papers in the literature
Most-cited downstream papers:
GSE119183 GEO reused by 1 papers in the literature
GSE141337 GEO reused by 1 papers in the literature
GSE141464 GEO reused by 1 papers in the literature

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33746712

Paper: Sicherman et al. 2021, Estimating and Correcting for Off-Target Cellular Contamination in Brain Cell Type Specific RNA-Seq Data. Front Mol Neurosci. DOI 10.3389/fnmol.2021.637143.

Method (what the paper actually does): marker-gene correlation scaling. For each cell-type-specific RNA-seq sample they compute Pearson correlation of its marker-gene expression against collapsed single-cell reference profiles (bootstrap subsampling of markers, n.iter=1e4), scale the resulting coefficients to an empirical prior, and add the per-off-target-cell-type coefficients as covariates in the DESeq2 differential-expression model. This removes false positives caused by off-target cellular contamination.

Code: the brief's code_url (github.com/alexdobin/STAR) is WRONG — STAR is only the read aligner mentioned in passing. The authors' actual analysis code is https://github.com/jsicherman/sct-contamination (2 commits, R-only: src/pipeline.R = function library, src/sample.R = driver). Per P16 this is the code to run.

In scope (attempted) — the SYNTHETIC benchmark = Table 1

The synthetic-data benchmark is the ONLY part of the repo that is self-contained: run_synthetic() in pipeline.R generates synthetic contaminated count data with a KNOWN differential-expression ground truth, using only:

  • the tasic2016data R package (Allen Inst. Tasic-2016 scRNA-seq) as reference;
  • the compcodeR Pickrell/Cheung mu/phi estimates (shipped in the package). No external download is needed. The driver then runs DESeq2 with vs. without the contamination covariates and reports the false-positive fraction (FPF) and AUROC. This maps 1:1 onto Table 1 (FPF/FNF/AUROC, uncorrected vs corrected, across 4 contamination×confound conditions). Headline: at High contamination + High confound, FPF 79%→5%, AUROC 0.618→0.677.

Target conditions (sample.R grid: contamination∈{0,.05,.1,.2}, confound∈{.5,.6,.7,1}):

  • None/None = contam 0, confound 0.5 (control)
  • High/High = contam 0.2, confound 1.0 (headline)
  • (High/Low = contam 0.2, confound 0.6; Low/Low = contam 0.05, confound 0.6 — if time)

Out of scope (NOT attempted) — the real-data results (the hard 20%)

The real LCM-seq / TRAP-seq / manual-sorting results (contamination coefficients 0.65 / 1.35 / 2.21 for oligodendrocytes-in-pyramidal across methods; per-cell-type DE-gene counts e.g. SST 8→58; AIC model-fit 64–67%) require external data the repo does NOT ship and the loader functions hard-code local paths to:

  • data/AllenInstitute/Mouse/transcrip.tome (Allen 2018 SMART-seq, large; needs the scrattch.io package; loaded via loadReference),
  • counts_exon.rds / counts_intron.rds / qc.tsv for GSE119183 (Aging), GSE145521 (Stress), GSE141337/GSE141464 (TRAP), GSE74985 (Hipposeq) — loaded via loadAging/loadStress with hard-coded relative paths. These also set no random seeds, so they are not byte-reproducible by design. Not attempted (80/20).

Known repo defects (documented patches)

  1. get_DEGs() is called by the driver but never defined in either file. Its signature and use match do.DEG(data, model, mapping=NULL) exactly, so the reproduction aliases get_DEGs <- do.DEG (a necessary, minimal patch).
  2. trace(glmer, edit=T) (sample.R l.107) is interactive (opens $EDITOR) and only affects an unused mixed-model variant — skipped.
  3. Two generate_prior definitions; the 2nd (ref, ref.collapsed, …) shadows the 1st. The synthetic path passes prior= explicitly, so this is harmless.

Deviation from paper (honest note)

The paper built the reference + collapsed profiles from the Allen-2018 transcrip.tome. We use the available tasic2016data package as the reference instead (same lab, earlier dataset, the only self-contained option). The synthetic benchmark's qualitative claim (covariates sharply reduce FPF under high contamination+confound) is robust to reference choice, but exact Table-1 c

Figures / tables: Table
C3
Reported
FPF 79% -> 5% (High contamination + High confound, Table 1 row 4 / Results)
Reproduced
not obtained: DESeq DE step errors 'NA values are not allowed in the count matrix'
partial
C4
Reported
AUROC 0.618 -> 0.677 (High/High, Table 1 row 4)
Reproduced
not obtained (same blocker)
partial
C1
Reported
FPF 12% -> 2% (None/None, Table 1 row 1)
Reproduced
not obtained (same blocker)
partial
I1
Reported
(intermediate) reference cell counts per cluster
Reproduced
Astro 43, Endo 14, L[1-6] 812, Oligo 38, Pvalb 275, Sst 202, Vip 186 (tasic2016data)
partial
I2
Reported
(intermediate) marker genes per cluster (MAST)
Reproduced
Astro 1388, Endo 3636, L[1-6] 203, Oligo 1676, Pvalb 209, Sst 80, Vip 163
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

205.5 k
tokens (I/O) · 14.3 M incl. cache
32 min
runtime · 8.53 CPU-h
93.1 GB
peak RAM
3 (2 failed)
HPC jobs
hummel
machine