Single-Cell Transcriptome Analysis Revealed Heterogeneity and Identified Novel Therapeutic Targets for Breast Cancer Subtypes.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Secondary computational re-analysis of the public Wu et al 2021 BC scRNA-seq atlas (GSE176078); listed code is the third-party tool AltAnalyze (P16). «our HPC» reachable throughout; all compute ran as SLURM jobs on compute node n093 (4 jobs, all COMPLETED exit 0). REPRODUCED EXACTLY: 26 patients (R1b); the GEO 3-way subtype labeling 11 ER+/5 HER2+/10 TNBC (R1c - which the paper itself states is how GEO labels them); and supplementary DE-table sizes 381 (S1, ER+ vs HER2+) and 321 (S4, TNBC vs ER+) plus S2/S3/S5/S6 = 220/386/229/290. NOT REPRODUCIBLE AS DESCRIBED: the headline 49,899-cell count (R1a) - the deposit holds 100,064 cells (full atlas); 49,899 is a ~49.9% downstream subset from the paper's unspecified AltAnalyze ICGS2 cell QC, not derivable from the shipped metadata; and the 13/44/29 DepMap therapeutic-target counts (R3a-c) - a faithful best-effort DE-up x DepMap(<=-0.3 over breast lines) intersection overshoots ~10x (144/110/76) with the wrong rank order, indicating the target-selection rule is under-specified. NOT ATTEMPTED: R4 KMplot survival (external cohort), OPLS-DA (SIMCA, commercial), STRING (web), wet-lab CFU knockdown. No evidence of outright fabrication, but two material methodology-opacity / non-reproducibility findings (49,899 subset; 13/44/29 targets) that a human reviewer should weigh. GSE176078 dataset itself is grade A: complete, internally concordant, delivers exactly what the source atlas promised. All grades PROVISIONAL and human-checkable.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-19 ⛓ 6b88477fe2ee
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan integrating single-cell transcriptomic data of EPCAM+ tumor epithelial cells with CRISPR-Cas9 functional screen data delineate the cellular heterogeneity of breast cancer molecular subtypes (ER+, HER2+, ER+HER2+, TNBC) and identify novel subtype-specific therapeutic targets?
- ★ Single-cell transcriptomic analysis of EPCAM+Lin- tumor epithelial cells identifies unique gene signatures that classify ER+, HER2+, ER+HER2+, and TNBC subtypes with high specificity and sensitivity finding
- ★ Integrating single-cell transcriptomics with CRISPR-Cas9 gene-effect data identified 13 therapeutic targets for ER+, 44 for HER2+, and 29 for TNBC resource
- ★ ENO1, FDPS, CCT6A, TUBB2A, and PGK1 predict worse relapse-free survival in basal breast cancer and are elevated in aggressive BLIS TNBC finding
- ★ Targeted depletion of ENO1 and FDPS reduces TNBC cell proliferation, colony formation, migration, and organoid growth while increasing cell death mechanism
- ★ Several identified targets (e.g., RPS4X, RPL34, VMP1 for ER+; RPS29 for HER2+) outperform current standard-of-care targets (ESR1/ERBB2) in CRISPR gene-effect potency finding
- FDPS-high TNBC is enriched in cell cycle and mitosis categories, whereas ENO1-high is associated with cell cycle, glycolysis, and ATP metabolic processes finding
- ★ Integration of single-cell transcriptomics with CRISPR-Cas9 screens provides the first comprehensive dependency map for each BC molecular subtype method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA-seq (re-analysis via ICGS2/UMAP/MarkerFinder) | 26 breast cancer patients (12 ER+, 3 HER2+, 2 ER+HER2+, 9 TNBC), 49,899 single cells, EPCAM+Lin- epithelial cells | none | cellular composition and subtype-specific differentially expressed gene markers | GSE176078 dataset; AltAnalyze v2.1.3 |
| CRISPR-Cas9 functional screen (data integration) | cancer cell lines from Achilles project | genome-wide CRISPR-Cas9 knockout | gene effect scores (≤-0.3 threshold) for essential genes | Achilles project (DepMap) |
| discriminant analysis (OPLS-DA) / ROC | BC subtype gene marker profiles | none | classification AUC/sensitivity/specificity across subtypes | SIMCA v16 (Umetrics) |
| survival analysis (Kaplan-Meier RFS) | 442 basal breast cancer patients | none | relapse-free survival stratified by median gene expression | KMplot database |
| bulk RNA-seq analysis | 360 TNBC patient cohort (BLIS, IM, LAR, MES subtypes) | none | differential expression and GO enrichment by ENO1/FDPS high vs low | KALLISTO 0.4.2.1, GENCODE v33, iDEP.951 |
| Colony Forming Unit (CFU) assay | MDA-MB-231 and BT-549 TNBC cell lines | siRNA knockdown of ENO1 and FDPS (30 nM) vs scrambled control | colony formation (crystal violet absorbance at 590 nm) | Lipofectamine 2000; crystal violet |
| AO/EtBr fluorescence cell death staining | MDA-MB-231 and BT-549 TNBC cell lines | siRNA knockdown of ENO1 and FDPS | number of dead (red/EtBr+) cells | Olympus IX73 fluorescence microscope; ImageJ |
| scratch/wound migration assay and 3D organoid dome culture | MDA-MB-231 and BT-549 TNBC cell lines | siRNA knockdown of ENO1 and FDPS | wound area closure at 24h; number of organoids | Matrigel GFR; ImageJ |
- – Identified 13 therapeutic targets for ER+, 44 for HER2+, and 29 for TNBC by integrating scRNA-seq with CRISPR-Cas9 data 13/44/29 targets
- – Gene classifiers discriminated the four BC subtypes with excellent ROC performance AUC: TNBC=0.98, ER+=0.94, HER2+=0.99, ER+HER2+=0.99
- ▲ ENO1, FDPS, CCT6A, TUBB2A, and PGK1 high expression predicted worse RFS in basal BC n=442
- ▲ Highest ENO1 expression in aggressive BLIS TNBC subtype; FDPS, PGK1, CCT6A highest in BLIS and LAR subtypes
- ▼ ENO1 and FDPS siRNA depletion reduced colony formation in MDA-MB-231 and BT-549 TNBC cells
- – RPS4X, RPL34, and VMP1 showed more profound CRISPR gene effects than ESR1 for ER+ BC
- – Differential expression identified subtype-comparison DEG counts (e.g., 381 ER+ vs HER2+; 321 TNBC vs ER+) 381/220/386/321/229/290 DEGs
- count 49,899 single cells (single cells analyzed from 26 BC patients)
- other AUC TNBC=0.98, ER+=0.94, HER2+=0.99, ER+HER2+=0.99 (ROC performance of gene classifiers)
- count 13 ER+, 44 HER2+, 29 TNBC targets (therapeutic targets identified per subtype)
- pvalue PPI enrichment p ≤ 1.0 × 10^-16 (HER2+ and TNBC gene target PPI network enrichment)
- pvalue PPI enrichment p = 0.000313 (ER+ gene target PPI network)
- pvalue GO response to interferon-alpha FDR p = 0.0076 (highest GO enrichment among ER+ targets)
- count 381 DEGs (differentially expressed genes ER+ vs HER2+)
- count 442 (basal BC patients in KMplot RFS analysis)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study combines computational analysis of public single-cell RNA-seq data (differential expression with fold-change and FDR-adjusted p-value cutoffs, OPLS-DA classification with ROC/AUC) with CRISPR-Cas9 dependency screen integration, PPI network enrichment, Kaplan-Meier/log-rank survival analysis in a clinical cohort, bulk RNA-seq differential expression/GO enrichment in a TNBC cohort, and in vitro functional assays (colony formation, apoptosis staining, scratch migration, 3D organoid growth) analyzed with pairwise statistics in GraphPad Prism. Results are reported primarily as fold-change/FDR thresholds, enrichment p-values, log-rank p-values, and mean ± SD for functional assays.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential expression analysis with 1.5-fold-change and FDR-adjusted p-value < 0.05 cutoff (AltAnalyze/iDEP.951) | Pairwise comparisons among BC molecular subtypes (Figure 2a) and ENO1/FDPS-high vs -low groups in the 360-patient TNBC cohort | Varies by comparison (e.g., 381, 220, 386, 321, 229, 290 differentially expressed genes reported per comparison); TNBC cohort n=360 | not stated |
| OPLS-DA discriminant analysis with ROC/AUC evaluation | Discrimination of ER+, HER2+, ER+HER2+, and TNBC subtypes based on identified gene markers (Figure 3e,f) | 49,899 single cells from 26 BC patients | not stated |
| Kaplan-Meier relapse-free survival analysis with log-rank test | Prognostic value of identified TNBC therapeutic targets (Figure 5a–e) | n = 442 basal BC patients (KMplot database) | not stated |
| PPI network enrichment analysis (STRING enrichment p-value) | Protein-protein interaction networks for ER+, HER2+, and TNBC target gene sets (Figure 4b,d,f) | not explicitly stated (based on identified target gene lists: 13, 44, and 29 genes) | not stated |
| Pairwise statistical analyses (specific test not named) in GraphPad Prism v9 | In vitro functional assays: colony formation (CFU), apoptosis (AO/EtBr), migration (scratch assay), and organoid growth following ENO1/FDPS knockdown (Figure 6 and related) | Experiments repeated at least twice; CFU data from four replicas; migration/organoid counts from three fields | not stated |
-
Differential expression was defined using a fixed 1.5-fold-change plus FDR-adjusted p-value < 0.05 cutoff via AltAnalyze/iDEP.951↳ Could also: A model-based differential expression tool with per-gene variance shrinkage (e.g., DESeq2, edgeR, or limma-voom) — These approaches explicitly model count/variance structure across replicates and can provide shrinkage-adjusted fold-change estimates, which some readers find useful alongside fixed-cutoff filtering
-
Functional assay comparisons (colony formation, apoptosis, migration, organoid growth) were analyzed with unspecified 'pairwise statistical analyses' in GraphPad Prism↳ Could also: Explicitly naming and justifying the test (e.g., Student's t-test for two groups meeting normality assumptions, or Mann-Whitney U as a non-parametric alternative), and using one-way/two-way ANOVA with a post-hoc correction when more than two groups or cell lines are compared — Naming the specific test and its assumptions supports reproducibility, and an ANOVA-based approach can jointly control error rate when multiple related comparisons (e.g., two cell lines × two targets) are made
-
Survival analysis dichotomized gene expression at the median to form high/low groups for Kaplan-Meier and log-rank testing↳ Could also: Cox proportional hazards regression treating gene expression as a continuous variable, or data-driven optimal cutpoint methods — Continuous modeling avoids information loss from dichotomization and allows adjustment for other clinical covariates, which can complement the median-split KM approach
-
No multiplicity-correction method is described for the multiple pairwise comparisons across in vitro functional assays (two targets × two cell lines × several readouts)↳ Could also: A formal multiple-comparison correction such as Holm-Bonferroni or FDR across that family of tests — This can help control the family-wise error rate when several related comparisons are drawn from the same experimental system
-
The OPLS-DA classifier's discriminative performance (ROC/AUC) was assessed on the same single-cell dataset used to derive the marker genes↳ Could also: Cross-validation (e.g., k-fold) or evaluation on an independent held-out cohort — Independent or cross-validated testing can give a more conservative estimate of how well the classifier generalizes beyond the discovery dataset
-
Dispersion was reported as SD for the colony-formation assay, with other assays not specifying a dispersion measure↳ Could also: Consistently reporting SD, SEM, or a 95% confidence interval across all quantified assays — Uniform reporting of a chosen dispersion measure (and CIs in particular) helps convey the precision of small-n functional assay estimates
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37190091
Paper: Vishnubalaji R, Alajez NM. Single-Cell Transcriptome Analysis Revealed Heterogeneity and Identified Novel Therapeutic Targets for Breast Cancer Subtypes. Cells 2023;12(8):1182. PMID 37190091 · PMCID PMC10137100 · DOI 10.3390/cells12081182.
Nature of the study: This is a secondary computational re-analysis of an existing public scRNA-seq atlas (GSE176078, Wu et al. Nat Genet 2021). The authors did NOT generate new sequencing data for the single-cell part; they downloaded the processed atlas and re-ran a clustering + marker + target-prioritisation pipeline, then added wet-lab validation. The listed "code" is the third-party tool AltAnalyze (https://github.com/nsalomonis/altanalyze) — per BRIEF P16, applying a third-party tool to the paper's data is an equally valid reproduction.
Pipeline(s) named in the paper
- AltAnalyze v2.1.3 — CPTT normalisation → ICGS2 (iterative clustering & guide-gene selection 2) → UMAP → MarkerFinder (marker genes per population).
- Differential expression between subtypes: 1.5 fold-change, FDR-adj p<0.05 (Tables S1–S6). Tool not fully specified (AltAnalyze / iDEP.951 mentioned).
- CRISPR essentiality intersection: DepMap/Achilles gene-effect score ≤ −0.3, crossed with up-regulated DE genes → "therapeutic targets" (13 ER+, 44 HER2+, 29 TNBC).
- OPLS-DA classifier in SIMCA v16 (commercial).
- STRING v11.5 PPI (web tool).
- Survival: 5 TNBC genes (ENO1, FDPS, CCT6A, TUBB2A, PGK1) vs relapse-free survival, n=442 basal BC (external meta-cohort, KM-plotter-style).
- Functional validation: ENO1 / FDPS siRNA knockdown, CFU assay (wet lab).
IN SCOPE (pipeline-derived, attempted)
| # | Result | Pipeline | Tractability | Notes |
|---|---|---|---|---|
| R1 | Dataset structure: 26 patients, subtype composition (12 ER+/3 HER2+/2 ER+HER2+/9 TNBC), 49,899 cells with per-subtype counts (16,350 / 7,824 / 11,487 / 14,238) | data download + metadata parse + subsetting | HIGH (clean 1:1) | Note: GEO labels the 26 as 11 ER+/5 HER2+/10 TNBC — paper RE-CLASSIFIES. Verifiable from the shipped metadata.csv. |
| R2 | ICGS2/UMAP cell populations (immune, fibroblast, pericyte, endothelial, epithelial); MarkerFinder markers | AltAnalyze ICGS2 | MEDIUM-LOW | Paper gives NO numeric cluster counts → only qualitative match possible. Stochastic; not bit-reproducible. |
| R3 | Therapeutic-target counts per subtype (13 / 44 / 29) | DE genes ∩ DepMap Achilles ≤−0.3 | MEDIUM | Needs supplementary DE gene lists (Tables S1–S6) + DepMap CRISPR (public). Intermediate DE counts not reported → reconstruct. |
| R4 | Survival: 5 TNBC genes predict worse RFS (n=442 basal) | KM-plotter-style external cohort | MEDIUM | Reproducible via public KM Plotter; external cohort, semi-in-scope. |
OUT OF SCOPE (not attempted; reason)
| Result | Reason |
|---|---|
| OPLS-DA classifier (SIMCA v16) | Commercial software, not obtainable/scriptable. |
| STRING PPI network figures | Manual web tool; figure-level, no quantitative claim. |
| ENO1/FDPS knockdown CFU reduction (56–61% / 76–81%) | Wet-lab experiment — not computational. |
| Other immunohistochemistry / protein validation | Wet lab. |
Reproduction strategy (floor → stretch)
- Floor (~80%, quick): R1 — download GSE176078 on «infra», parse metadata, verify patient count, subtype composition, total + per-subtype cell counts. Cleanest, most auditable 1:1.
- Stretch: R3 (target counts) using public supplementary gene lists + DepMap; R4 (survival) via KM Plotter; R2 a qualitative Seurat/AltAnalyze clustering sanity check.
Hard dependency
All data (GSE176078 processed tar.gz, 532.9 MB; DepMap; supplementary tables) lives on «infra», downloaded on «host». No data on «host». Requires the «our HPC» tunnel up.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.