Decoding the tumor microenvironment and molecular mechanism: unraveling cervical cancer subpopulations and prognostic signatures through scRNA-Seq and bulk RNA-
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the in-scope scRNA-seq cell counts close to 1:1. Standard Seurat v4.3.0.1 workflow on GEO GSE171894 (4 cervical-cancer samples) run on «our HPC» (final «job»). 6/7 pinned cell-count claims reproduce within tolerance; the same 6 cell-type identities appear in the same rank order with close magnitudes. Post-QC total 14286 vs reported 13770 (+3.7%) is mechanistically explained by our omission of DoubletFinder (paper ran it; ~3-4% doublet removal closes the gap). Only Myeloid (456 -> 550, +20.6%) is partial (small population, sensitive to clustering resolution + automated-vs-manual annotation boundary vs NK_T). Every in-scope reported value IS derivable from the shipped GEO data via the described pipeline => no fabrication concern. Honest data note: the deposited GSE171894 matrices have near-constant per-cell totals (appear library-size-normalized, not raw UMIs), so nCount-based QC is approximate, but the authors used the same matrices. NOT attempted (beyond the 80%): inferCNV malignant-EPC call + 5 tumor-EPC subclusters; LASSO 9-gene prognostic signature; survival ROC AUC (0.762/0.733/0.781) and Cox p=0.001; CIBERSORT/pRRophetic/CellChat -- these need TCGA-CESC bulk + many unstated parameters.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 80assessed: 2026-06-15 ⛓ 84e66f34c2a9
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study investigates the cellular heterogeneity of cervical carcinoma by combining scRNA-seq and bulk RNA-seq to identify tumor epithelial cell subpopulations driving CC progression and to build a prognostic signature, hypothesizing that a specific tumor epithelial progenitor subpopulation critically influences differentiation, progression, and prognosis of CC.
- ★ C3 PLP2+ Tumor Epithelial Progenitor Cells are a key subpopulation that drives differentiation and progression of cervical carcinoma finding
- ★ A PLP2+ Tumor EPCs score serves as an independent prognostic indicator, with high-score patients showing worse survival than low-score patients resource
- ★ Risk-score groups correlate with immune infiltration, enriched pathways, SNPs/somatic mutation, and drug sensitivity in CC finding
- ★ ATF6 significantly affects proliferation and migration of CC cell lines, validated by cellular experiments mechanism
- ★ Integrated scRNA-seq and bulk RNA-seq pipeline maps the CC tumor microenvironment and builds a LASSO/Cox prognostic signature method
- ★ inferCNV-based separation identified five tumor epithelial cell subgroups (C0 TMPRSS2+, C1 ANKRD36C+, C2 HK2+, C3 PLP2+, C4 MKI67+) finding
- scRNA-seq of CC resolves six major cell types within the tumor microenvironment finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA sequencing (scRNA-seq) | tumor tissue from 4 cervical carcinoma patients (GSE171894) | none | cell type clusters, gene expression, tumor epithelial subpopulations, CNV inference, trajectory, cell-cell interactions | — |
| bulk RNA-seq analysis | TCGA CESC cohort | none | DEGs, survival/prognosis, immune infiltration, prognostic signature | — |
| RT-qPCR | SiHa and Hela CC cell lines | siRNA knockdown of ATF6 (vs si-NC) | ATF6 mRNA expression (GAPDH reference) | SYBR Green Kit (TaKaRa); PrimeScript RT kit |
| CCK-8 proliferation assay | SiHa and Hela CC cell lines | ATF6 siRNA knockdown | OD value (cell proliferation over 1-5 days) | CCK-8 (Vazyme A311-01) |
| Wound healing assay | SiHa and Hela CC cell lines | ATF6 siRNA knockdown | cell migration at 0 and 48 hours | — |
| Transwell assay (with/without Matrigel) | SiHa and Hela CC cell lines | ATF6 siRNA knockdown | cell migration/invasion at 36 hours | BD Biosciences Matrigel; Corning chambers |
| somatic mutation / TMB analysis | TCGA CESC cohort | none | mutation distribution, TMB, CNV of modeled genes, survival | maftools |
| drug sensitivity prediction | TCGA CESC cohort (high vs low risk groups) | none | predicted IC50 of chemotherapeutic drugs | pRRophetic |
- ▼ High PLP2+ Tumor EPCs score group exhibited diminished survival compared to the low score group
- ▼ ATF6 knockdown reduced proliferation and migration of CC cell lines
- – C3 PLP2+ Tumor EPCs identified as progenitor subpopulation influencing CC differentiation and progression
- – 13,770 high-quality cells clustered into 23 states and annotated to six major cell types
- – Five tumor epithelial cell subgroups identified via inferCNV sub-clustering
- – No statistically significant difference in proportion of the six cell types between HPV+ and HPV- groups
- count 13770 cells retained after QC (total high-quality cells after quality control and batch correction)
- count 4 patients (CC samples used for scRNA-seq (GSM5236544-5236547))
- count 23 clusters (unique tissue states from dimensionality reduction)
- count NK_T 6975, Epithelial 5434, B_Plasma 707, Myeloid 456, pDCs 139, Fibroblasts 59 (cell numbers per major cell type)
- count C0 TMPRSS2+ 1266, C1 ANKRD36C+ 919, C2 HK2+ 489, C3 PLP2+ 440, C4 MKI67+ 39 (tumor epithelial subgroup cell numbers)
- other median survival 16.8 months (advanced-stage CC patients (background))
- other 5-year survival rate 72% (average CC survival (background))
- other chemotherapy response rate 29% to 63% (current chemotherapeutic agent efficacy (background))
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a computational bioinformatics study combining single-cell RNA-seq (4 cervical cancer patients, GSE171894) with TCGA bulk RNA-seq, supplemented by in vitro validation. The single-cell workflow used Seurat for QC, normalization, clustering (UMAP), Harmony batch correction, inferCNV for tumor-cell identification, and Wilcoxon rank-sum tests for marker-gene detection; trajectory (CytoTRACE/Monocle/Slingshot) and cell-cell communication (CellChat) analyses were applied. A prognostic risk score was built via univariate Cox plus LASSO regression, evaluated with Kaplan-Meier survival, time-dependent ROC, and a nomogram (C-index), with downstream immune-infiltration (CIBERSORT/ESTIMATE/xCell), DESeq2 differential expression, mutation (maftools), and drug-sensitivity (pRRophetic) analyses. Wet-lab validation used CCK-8, wound-healing, and Transwell assays following ATF6 siRNA knockdown.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Wilcoxon rank-sum test (Seurat FindAllMarkers) | identification of differentially expressed marker genes per cell type and per tumor epithelial subpopulation | 13,770 cells total; subpopulation cell counts stated (e.g., C3 PLP2+ = 440) | not stated |
| Paired Wilcoxon test | comparison of the proportion of the 6 cell types between HPV+ and HPV- groups (Figure 1D) | — | not stated |
| Univariate Cox proportional-hazards regression | association between key subgroup marker-gene expression and overall survival in TCGA | — | not stated |
| LASSO Cox regression | selection of prognostic genes and construction of the PLP2+ Tumor EPCs risk score | — | not stated |
| Kaplan-Meier survival analysis (log-rank implied) | survival difference between high vs low risk-score groups, and high vs low TMB groups | — | not stated |
| Time-dependent ROC (timeROC) | 1-, 3-, 5-year predictive accuracy of the risk score | — | na |
| DESeq2 (Wald test) | differential expression between high- and low-risk groups in TCGA bulk data (threshold |logFC|>2, p<0.05) | — | not stated |
| Hypergeometric/over-representation test (ClusterProfiler GO/KEGG/GSEA) | enrichment of cell-type and risk-group differential genes (adjusted p<0.05) | — | na |
-
Marker genes were identified with the Wilcoxon rank-sum test in Seurat's FindAllMarkers.↳ Could also: Single-cell-specific differential-expression frameworks such as MAST (which models the bimodal expression distribution) or a pseudobulk approach with DESeq2/edgeR aggregating counts per patient could also be used. — Pseudobulk or mixed-effects methods account for the limited number of biological replicates (4 patients) and treat the patient as the unit of replication, which can complement cell-level tests.
-
Patients were dichotomized into high and low groups at the median risk score for survival comparison.↳ Could also: Treating the risk score as a continuous covariate in a Cox model, or selecting a cut-point via maximally-selected rank statistics, could also be reported. — Retaining the continuous score preserves information and avoids dependence on a single threshold, while a continuous hazard ratio conveys effect size directly.
-
A LASSO Cox model was built and evaluated with time-dependent ROC and a nomogram C-index on the TCGA cohort.↳ Could also: External validation in an independent cohort (e.g., GEO) and bootstrap or cross-validated calibration could also be presented. — Independent validation and resampling-based calibration help characterize how the signature generalizes beyond the training data.
-
Adjusted p-value thresholds were applied for enrichment and marker selection.↳ Could also: Explicitly naming the correction method (e.g., Benjamini-Hochberg FDR) and applying a defined multiplicity correction across the survival/clinical comparisons could also be stated. — Naming the procedure and defining the test family makes the error-rate control transparent and reproducible.
-
Cell-type proportions between HPV+ and HPV- groups were compared with a paired Wilcoxon test (Figure 1D).↳ Could also: Compositional-data methods such as scCODA or a Dirichlet-multinomial/centered-log-ratio approach could also be used for cell-proportion comparisons. — Compositional methods account for the sum-to-one constraint of proportions and the small number of samples, which standard rank tests do not explicitly model.
-
Drug response was inferred via pRRophetic-predicted IC50 values compared between risk groups.↳ Could also: Reporting effect sizes with confidence intervals alongside the group comparison, and corroborating with an independent pharmacogenomic resource, could also be included. — Effect-size reporting and a second data source provide additional context on the magnitude and robustness of predicted sensitivity differences.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
ATF6 knockdown reduces proliferation and migration in cervical cancer cell lines (SiHa, HeLa).other human cervical-cancer-cell-line down 2024×1papers★ This paper is the founder (earliest)
-
High PLP2+ tumor epithelial progenitor cell signature score is associated with significantly worse overall survival in cervical cancer.RNA-seq human cervical cancer down 2024×1papers★ This paper is the founder (earliest)
-
Proportions of major cell types do not differ significantly between HPV-positive and HPV-negative cervical cancer tumors.scRNA-seq human cervical cancer none 2024×1papers★ This paper is the founder (earliest)
-
PLP2+ tumor epithelial progenitor cells (C3 subcluster) constitute a progenitor subpopulation driving cervical cancer differentiation and progression.scRNA-seq human cervical cancer 2024×1papers★ This paper is the founder (earliest)
-
InferCNV-guided sub-clustering identifies five distinct tumor epithelial cell subgroups in cervical cancer.scRNA-seq human cervical cancer 2024×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-38482016
Paper: Lin et al. 2024, Front Immunol 14:1351287. "Decoding the tumor microenvironment and molecular mechanism: unraveling cervical cancer subpopulations and prognostic signatures through scRNA-Seq and bulk RNA-seq." PMID 38482016 · PMCID PMC10933018 · DOI 10.3389/fimmu.2024.1351287.
Named code: https://github.com/broadinstitute/inferCNV (a third-party tool,
not the authors' own repo — per brief rule P16 this is equally valid; reproduce
by running the described pipeline on the paper's data).
Data: GEO GSE171894 — 4 cervical-cancer scRNA-seq samples
(GSM5236544 HPV+.1, GSM5236545 HPV+.2, GSM5236546 HPV-.1, GSM5236547 HPV-.2),
each a gzipped expression matrix (*.txt.gz). Public, no restriction.
Pipeline behind the reported results
Standard Seurat v4.3.0 scRNA-seq workflow: read 4 matrices → QC filter (mito < 20 %, hemoglobin < 5 %, "extreme" nFeature/nCount removed, DoubletFinder doublet removal) → 2000 HVGs → 30 PCs → UMAP → graph clustering → manual cell-type annotation by canonical markers → inferCNV (NK/T cells as reference) to call malignant epithelial cells → tumor-EPC subclustering. Bulk arm: TCGA-CESC differential expression + LASSO-Cox prognostic signature + timeROC / CIBERSORT / pRRophetic.
IN SCOPE (clearly-specified, low-hanging — the 80 %)
| id | reported result | paper loc | pipeline |
|---|---|---|---|
| C1 | 13770 cells retained after QC | Results / Fig.1 | Seurat QC |
| C2 | 6 major cell types w/ counts: NK_T 6975, Epithelial 5434, Fibroblasts 59, pDCs 139, B_Plasma 707, Myeloid 456 (sum = 13770) | Results / Fig.1 | Seurat clustering + marker annotation |
These two are unusually explicit per-cell-type integer counts and sum exactly to the QC total — an internal-consistency anchor that makes a faithful 1:1 comparison possible.
STRETCH (the hard ~20 %, attempt only if the core runs cleanly)
- inferCNV malignant-epithelial call + 5 tumor-EPC subclusters with counts (C0 TMPRSS2+ 1266, C1 ANKRD36C+ 919, C2 HK2+ 489, C3 PLP2+ 440, C4 MKI67+ 39). Depends on subcluster resolution and inferCNV cutoffs that the paper does not state — not expected to land 1:1.
OUT OF SCOPE (not attempted — needs external data / many unstated params)
- LASSO 9-gene prognostic signature (PLAGL1, HIF1A, ERG, ELF1, ATF6, ATF1 / TBX21, SPIB, LHX2, JUND, ETV7, ATF5) — requires TCGA-CESC bulk + unstated gene-mapping/penalty seed.
- Survival ROC AUC (1 yr 0.762, 3 yr 0.733, 5 yr 0.781), multivariate Cox p = 0.001 — depend on the LASSO signature + TCGA clinical table.
- CIBERSORT immune infiltration, pRRophetic drug sensitivity, CellChat — out of the 80 % and tool-chain-heavy.
Honesty notes
- QC threshold "extreme nFeature/nCount values" is not quantified in the paper; the exact 13770 therefore depends on author-private cutoffs. We use defensible standard cutoffs and report the gap honestly rather than tuning to hit 13770.
- Cell-type annotation in the paper is manual; we annotate clusters by canonical markers and aggregate to the same 6 labels.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a solid partial reproduction of the in-scope descriptive scRNA-seq landscape: 6/7 pinned cell counts on public GEO GSE171894 reproduce within tolerance (NK_T 6975→7110, Epithelial 5434→5652), and every in-scope value is derivable from the shipped matrices — no fabrication concern. The only deviations are on our methodology / preprocessing side: omitting DoubletFinder explains the +3.7% total excess, and unstated clustering resolution/annotation drives the lone partial (Myeloid +20.6%). Severity is low-to-moderate with magnitude, direction and rank order preserved. However, the paper's actual central conclusions (LASSO prognostic signature, survival ROC AUC, Cox p=0.001) were out of scope and not attempted, so the headline claim is only limitedly confirmed → overall yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.