Gene Dosage Analysis on the Single-Cell Transcriptomes Linking Cotranslational Protein Targeting to Metastatic Triple-Negative Breast Cancer.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to VERIFY THE ASSIGNED DATASET, but NOT to reproduce the paper's core pipeline. GSE75688 (the paper's validation set) downloaded + profiled on «infra»: it is a clean, complete, open Smart-seq2 TPM matrix (57,915 genes x 563 expression cols, 11 patients BC01-BC11). Two reported numeric facts about it reproduce EXACTLY by direct inspection/arithmetic: 57,915 genes, and 7,528,950 Z-scores = 57,915 genes x 130 cells (internally consistent; no fabrication evident). The '130 cells'/'5 TNBC patients' are post-filter/external-subtype facts (raw TNBC tumour single cells = 89; subtype labels are not in the deposit). The ONE shipped 'code' artifact is the generic vegan R package, cited only for the Fig 5 diversity indices; running vegan 2.7.1 diversity() on GSE75688 reproduces the qualitative Fig 5B finding (Shannon & Simpson strongly correlate: Pearson 0.86 / Spearman 0.95) -- though Fig 5 in the paper is computed on the PRIMARY dataset, so this is a method cross-check, not a like-for-like number. NOT ATTEMPTED / not reproducible: the central CNV-gene-dosage concordance pipeline and everything derived from it (20,651 CNV events, 86/94 CNG-UP genes, 33 SRP genes, 5 MCODE modules, GO p-values, cBioPortal survival) -- because (a) the authors shipped NO custom code (only a prose Methods + a generic package link), and (b) the CNV modality is entirely absent from GSE75688 and the room was given no matched DNA-seq. Reproducing those would be re-creation from prose, not reproduction, so they are honestly recorded as out-of-scope rather than fabricated. Overall: dataset solid + count claims exact; pipeline non-reproducible from shipped artifacts.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 53assessed: 2026-06-18 ⛓ a8ffe567570a
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a computational framework integrating concordant copy number variation (CNV) and single-cell gene expression reveal common gene dosage effects across cell types in metastatic triple-negative breast cancer (TNBC), and does amplification-induced upregulation of ribosome/protein-targeting genes drive its metastatic potential?
- ★ A computational framework integrating independently measured CNV (DNA sequencing) and single-cell RNA-seq from the same patients identifies recurrent concordant copy number gain and gene upregulation (CNG-UP) events and functional modules at the single-cell level. method
- ★ Consistent copy number gains and upregulation in metastatic TNBC are enriched for ribosome proteins involved in SRP-dependent cotranslational protein targeting to membranes. finding
- ★ SRP-dependent cotranslational protein targeting is the top functional module, validated as the prioritized module in an independent metastatic TNBC dataset. finding
- ★ Increased ribosome gene copies in TNBC associate with enhanced stemness, differentiation, and EMT/metastatic potential. mechanism
- ★ The 33 ribosome/protein-targeting genes are frequently mutated in metastatic breast cancers and their mutation associates with worse patient survival. finding
- Mutations in the 33 genes correlate with adverse clinical features including higher histological grade, later tumor stage, higher aneuploidy and hypoxia scores, and older diagnosis age. finding
- Ribosome protein modules represent a potential target for TNBC therapy. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA-seq combined with matched whole-exome/DNA sequencing CNV analysis (computational dosage framework) | metastatic TNBC patient tumors, GSE118389/GEO118390 (6 patients, 1533 cells) | none (observational CNV) | concordant CNG-UP events, Z-scored relative gene expression mapped to CNV | — |
| single-cell RNA-seq with matched CNV analysis (independent validation) | metastatic TNBC patients, GSE75688 (5 patients, 130 cells) | none | CNG-UP events and functional modules (SRP-dependent cotranslational protein targeting) | — |
| functional enrichment / GO & pathway analysis with MCODE module detection on human interactome | 94 top CNG-UP genes from TNBC single cells | none | enriched GO terms, pathways, and functional modules | MCODE algorithm |
| mutational frequency and survival analysis on bulk cancer genomics data | 6688 breast cancer samples across 12-15 studies; survival on 4821 patients | none | mutation frequency per gene, overall/relapse-free/disease-specific/disease-free/progression-free survival | public cancer genomic resource (cBioPortal-type) |
| intratumor heterogeneity and cell-state analysis (diversity indices + tSNE) | 6 TNBC patients, GSE118389 single cells | none | Shannon-Wiener index, Simpson index, marker-based cell states (stemness, pluripotency, differentiation, proliferation, EMT/metastasis) vs SRP module | tSNE |
- ▲ 47,514 CNG-UP events associated with 94 genes across 1145 cells identified in primary TNBC dataset 94 genes; 47,514 events (abstract: 47,198)
- ▲ Cotranslational protein targeting to membranes was the most significantly enriched functional term (20 genes) corrected p = 10×10^-25.133
- ▲ Independent dataset yielded 86 genes with CNG-UP in 6+ cells; confirmed SRP-dependent cotranslational protein targeting as top module 33 overlapping genes (7 in common); 29 ribosome protein genes
- ▲ 31 of 33 genes related to protein translation; 7 related to VEGFA-VEGFR2 signaling p = 10×10^-56.43 (translation); p = 10×10^-3.76 (VEGFA-VEGFR2)
- ▲ Three metastatic breast cancer cohorts highly mutated (>50% patients) in the 33 genes, unlike non-metastatic cohorts >50% of patients
- ▼ Patients with mutations in the 33 genes had shorter median overall survival (145.43 vs 175.30 months) 145.43 vs 175.30 months
- ▲ Ribosome (SRP) module positively associated with cell differentiation, stemness, and EMT/metastasis states
- ▲ Top mutated genes included MRPL13 (17%), SRP9 (15%), PABPC1 (15%), RPL8 (15%) 11-17% mutation frequency
- count 47,514 CNG-UP events / 94 genes / 1145 cells (primary TNBC dataset filtered concordant events)
- pvalue corrected p = 10×10^-25.133 (cotranslational protein targeting to membranes enrichment)
- pvalue 10×10^-56.43 (33 genes protein translation enrichment)
- pvalue logrank p = 8.94×10^-7, Q = 4.24×10^-6 (overall survival altered vs unaltered group)
- mean 145.43 vs 175.30 survival months (median OS mutated (1625 patients) vs unmutated)
- count 6688 breast cancer samples, 12 studies (mutational analysis cohort)
- count 89,524 meaningful CNV-expression events (|Z|>1.96) of 1,440,802 total in 1245 cells (primary dataset filtering)
- other mutation frequencies 11-17% (top ten mutated ribosome/SRP genes)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper presents a computational framework that maps CNV data from bulk DNA sequencing to Z-score-transformed single-cell RNA expression profiles from the same patients to identify concordant copy number gain and gene upregulation (CNG-UP) events in metastatic TNBC. Primary analysis was conducted on 1145 cells from 6 patients (GSE118389), with independent validation in a second cohort of 5 patients (GSE75688); enriched genes were clustered into functional modules using the MCODE network algorithm. Survival analyses using log-rank tests with Q-value correction were performed on pooled multi-cohort breast cancer data (up to 6688 samples), and intratumor heterogeneity was characterised using Shannon-Wiener and Simpson ecological diversity indices.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Z-score threshold (|Z| > 1.96) applied as a significance criterion for single-cell gene expression relative to the cross-cell mean | Identifying significantly up- or downregulated genes in each of the 1533 cells across the primary TNBC dataset; analogous threshold applied in validation dataset | 1533 cells from 6 patients (primary); 130 cells from 5 patients (validation) | not stated |
| Gene Ontology and biological pathway functional enrichment analysis with corrected p-values (method not named) | Enrichment of 94 CNG-UP genes (primary) and 33 overlapping SRP-module genes across GO terms and pathways including VEGFA-VEGFR2, Rho GTPase, spliceosome, glycolysis | 94 genes (primary); 86 genes (validation); 33 genes (overlap) | not stated |
| Log-rank test with Q-value correction | Five survival endpoints (overall survival, relapse-free, disease-specific, disease-free, progression-free survival) comparing patients with vs. without mutations in the 33-gene module | 4821 breast cancer patients (survival analysis); 6688 samples across 12 studies (mutation frequency) | not stated |
| Shannon-Wiener diversity index and Simpson diversity index | Quantifying intratumor expression heterogeneity per patient; correlation between the two indices across patients | 6 patients (GSE118389) | na |
| MCODE (Molecular Complex Detection) network clustering algorithm | Clustering GO-annotated CNG-UP genes from the human interactome into functional modules (5 modules in primary, 3 in validation) | 94 genes (primary); 86 genes (validation) | na |
| tSNE (t-distributed stochastic neighbour embedding) | Dimensionality reduction for visualising cell-state variation across all single cells and defining five cell states (stemness, pluripotency, differentiation, proliferation, EMT/metastasis) | 1145 cells (primary dataset) | na |
-
Survival associations were assessed with log-rank tests comparing a binary altered/unaltered grouping, and the paper separately documents that grade, stage, age, and chemotherapy receipt differ substantially between these groups↳ Could also: Multivariable Cox proportional hazards regression could also have been used, adjusting for the clinical covariates shown to differ between groups — A multivariable Cox model would allow assessment of whether the 33-gene module is independently prognostic beyond stage, grade, and treatment, and would yield hazard ratios with confidence intervals as interpretable, reportable effect sizes
-
Single-cell gene expression significance was determined by a fixed Z-score threshold of |Z| > 1.96, applied uniformly across all genes and cells↳ Could also: Statistical frameworks designed for single-cell count data — such as MAST (a hurdle model), DESeq2, or edgeR — could also identify differentially expressed genes while explicitly modelling dropout and the distributional properties of UMI or read counts — Dedicated scRNAseq differential expression tools account for zero-inflation (dropouts) and overdispersion characteristic of single-cell experiments, which the Z-score approach applied here does not model, potentially affecting the set of genes prioritised
-
The recurrence threshold for retaining CNG-UP events was set at ≥100 cells (primary) or ≥6 cells (validation) as absolute counts without a stated statistical rationale for the cut-off values↳ Could also: A permutation test or binomial/hypergeometric test could also be used to assess whether the observed co-occurrence of CNG and upregulation in a given number of cells exceeds chance expectation given the marginal frequencies of each event — A significance-based recurrence threshold would provide a statistical basis for the cutoff, yield a p-value or FDR for each candidate gene's recurrence, and allow estimation of false discovery rates among retained CNG-UP events
-
GO and pathway enrichment used a corrected p-value on a threshold-defined binary gene list (94 genes), with the correction method unnamed↳ Could also: Gene Set Enrichment Analysis (GSEA) on a continuously ranked gene list (e.g., ranked by recurrence frequency or mean Z-score) could also test for pathway enrichment without requiring a binary threshold — Rank-based enrichment methods use the full score distribution rather than a binary gene list, reducing sensitivity to the choice of cutoff and providing a normalized enrichment score as an effect-size analogue alongside the p-value
-
Clinical associations (race, grade, stage, chemotherapy, diagnosis age, aneuploidy score, hypoxia score) between altered and unaltered patient groups were described narratively with figures but without formal statistical tests or reported p-values↳ Could also: Chi-squared or Fisher exact tests (for categorical variables: race, grade, stage, chemotherapy) and Mann-Whitney U tests or t-tests (for continuous variables: diagnosis age, aneuploidy score, hypoxia score) could also formally quantify these associations, with multiplicity correction across the seven features — Formal tests with reported statistics, p-values, and effect sizes (e.g., odds ratios for categorical associations) would allow readers to assess the strength and precision of each observed association and distinguish signal from description
-
Intratumor expression heterogeneity was quantified per patient using Shannon-Wiener and Simpson ecological diversity indices treating expressed genes as analogues of species↳ Could also: Purpose-built single-cell heterogeneity metrics such as scEntropy, or variance decomposition via mixed-effects models partitioning expression variance into between-patient and within-patient components, could also quantify transcriptomic diversity — Ecological diversity indices were developed for species-abundance distributions; tools designed for transcriptomic data can account for the specific sparsity, gene-gene correlations, and dropout structure of scRNAseq, potentially yielding more biologically interpretable heterogeneity estimates
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34577617
Paper: Liu Y, Zhao M. Gene Dosage Analysis on the Single-Cell Transcriptomes Linking Cotranslational Protein Targeting to Metastatic Triple-Negative Breast Cancer. Pharmaceuticals (Basel) 2021. DOI 10.3390/ph14090918. PMCID PMC8472593.
Assigned data: geo:GSE75688 (Chung et al. 2017 breast-cancer scRNA-seq) —
this is the paper's validation dataset.
Assigned "code": github.com/vegandevs/vegan — the generic community-ecology
R package. The authors shipped NO custom code/scripts. vegan is cited only
for the diversity (heterogeneity) indices and an ordination correlation in Fig 5.
Reported results and their pipeline origin
| Result (paper) | Pipeline | In scope? | Why |
|---|---|---|---|
| GSE75688 has 57,915 genes; 7,528,950 Z-scores over 130 cells (5 TNBC patients) | direct count / Z-score transform | YES | derivable by direct inspection of the deposited matrix |
| Z-score transform on primary set: 1,533 cells, 21,785 genes, 33,396,405 Z-scores | custom prose pipeline on GSE118389/90 | NO (wrong dataset) | not GSE75688; primary dataset not assigned to this room |
| CNV–expression concordance → 89,524 consistent events → 94 CNG-UP genes (primary); 86 CNG-UP genes (GSE75688) | custom CNV/Z concordance | NO | no shipped code; requires matched patient CNV/DNA-seq which GSE75688 (expression-only TPM) does not contain |
| 33 SRP-dependent genes shared; 7 overlapping genes (Fig 2E/2F) | MCODE on gene network | NO | downstream of above; no code; needs the CNG-UP gene lists |
| 5 MCODE functional modules; GO enrichment p-values | MCODE / Metascape (external GUI) | NO | external interactive tool; not a scriptable shipped pipeline |
| Mutation frequencies (cBioPortal, 6,688 samples); survival logrank p=8.94e-7 | cBioPortal queries on external TCGA cohorts | NO | external cohorts (not GSE75688); GUI/portal, no code |
| Fig 5A Shannon-Wiener & Fig 5B Simpson~Shannon correlation; Fig 5C t-SNE envfit | vegan diversity + ordination |
PARTIAL | the ONE shipped tool; computed on the primary set in the paper (no GSE75688 numbers reported) → we reproduce the method on GSE75688 as a qualitative cross-check |
What this room attempts
- Dataset profiling of GSE75688 (required) — N genes/cells/patients reported vs observed, QC.
- Exact count claims about GSE75688: 57,915 genes; 130 cells (via 57,915×130 = 7,528,950 Z-scores).
- Methodological repro of the vegan diversity result (Fig 5A/5B) on GSE75688 tumour single cells: Shannon & Simpson per cell and their correlation — qualitative (paper reports no GSE75688 diversity numbers).
Explicitly NOT attempted (and why)
- The core CNV-gene-dosage concordance pipeline and all results derived from it (86/94 CNG-UP genes, 33 SRP genes, MCODE modules): no code shipped and the CNV modality is absent from GSE75688. Reproducing it would require re-implementing an under-specified prose pipeline and sourcing matched CNV data the room was not given — that is re-creation, not reproduction.
- Survival / mutation-frequency analyses: external TCGA/cBioPortal cohorts, GUI-driven.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The clean, checkable facts about GSE75688 reproduce exactly and are internally consistent (57,915 genes; 57,915×130 = 7,528,950 Z-scores), with no fabrication evident. However, the paper's central CNV-gene-dosage pipeline (20,651 CNV events, 86 CNG-UP genes, and all SRP/MCODE/GO/survival downstream) is non-reproducible from the shipped artifacts: no author code was shipped (the code link is the generic vegan package) and the assigned deposit is expression-only with no CNV modality. The deviation therefore sits on the authors'/availability side (incomplete code + a headline number attributed to a dataset lacking that modality), not in a demonstrated computational disagreement. Overall this is partial: solid and exact where checkable, core conclusion untested and explainably non-reproducible — yellow rather than red because nothing is proven false or fabricated.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.