Single-cell RNA-sequencing of circulating tumour cells: A practical guide to workflow and translational applications.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
▸Reproduction agent’s raw note
DROP (non_pipeline) - honest category drop, no fabrication concern. PMID 41053409 ('Single-cell RNA-sequencing of circulating tumour cells: A practical guide to workflow and translational applications', Tieng/Lee/Ab Mutalib, Cancer Metastasis Rev 2025) is a NARRATIVE REVIEW (Europe PMC full-text JATS article-type=review-article), not a primary research paper. It reports NO original pipeline-derived result and so has no claim to reproduce: there is no Methods/Results/Data-Availability of an analysis the authors ran, zero self-analysis language in the full text (grep for 'we analysed/performed/applied/our (re)analysis' = 0 hits), all 4 figures are conceptual schematics (milestone timeline; the proposed 12-step CTC scRNA-seq workflow; a workflow+tools diagram; emerging frontiers), and all 7 tables are literature/tool compilations. The harvester's code_url (github.com/brwnj/bcl2fastq) is a tool-catalogue URL in Table 4 ('Commonly used tools for pre-/post-processing'), one of three resource links for the generic Illumina demultiplexer bcl2fastq - not the authors' code. The harvester's data_accession (GSE144494) is one of ~20 third-party GEO datasets surveyed in Table 1 ('Overview of published scRNA-seq studies on CTCs') - the Aceto-lab HR+ breast-cancer CTC dataset (135 single CTCs/clusters from 45 patients, CTC-iChip + Smart-seq), not the authors' own data. Running that demultiplexer on that dataset would reproduce nothing the review claims; there is no expected output to match (overlaps no_expected_result; non_pipeline is the precise root cause). This is a text-mining false positive: code+data are both publicly resolvable, so it is NOT a data/code-availability failure and NOT a fabrication case - the paper simply reports no computed value to check. Decision made from full-text inspection on the «host» control plane; correctly NO «our HPC» compute was spent (HARD RULE 2 scope-to-pipeline, RULE 3 80/20, RULE 6 drops-are-valid). NOT ATTEMPTED: any pipeline run - there is none to attempt. Artifacts: scope.md, AUDIT.md, original/claims.tsv (empty w/ rationale), reproduction/agreement.json.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-14 ⛓ 6d2aee2bcd63
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThis review addresses the knowledge gap created by unstandardised protocols and fragmented resources in circulating tumour cell (CTC) scRNA-seq research by proposing a standardised 12-step CTC-specific scRNA-seq workflow and evaluating analytical tools for their suitability in CTC research.
- ★ A 12-step CTC-specific scRNA-seq workflow spanning enrichment, single-cell sorting, sequencing, data pre-processing and downstream analysis is proposed to overcome methodological inconsistencies. method
- ★ scRNA-seq enables deep transcriptomic profiling of CTCs, re-stratifying subtypes and detecting rare new subpopulations beyond bulk sequencing. finding
- ★ Hybrid cells, fusion products of tumour and normal/immune cells, represent a novel research frontier complicating tumour heterogeneity and immune response understanding. finding
- ★ Integration of machine learning into scRNA-seq workflows enhances CTC clustering, cell identification and heterogeneity analysis. method
- ★ A compiled comparison of scRNA-seq computational tools evaluates their pros, cons and suitability for CTC research. resource
- ★ CTC scRNA-seq reveals molecular mechanisms of EMT, metastasis, immune evasion and therapy resistance across cancer types. mechanism
- Table 1 provides a comprehensive overview of published CTC scRNA-seq studies across cancer types, enrichment methods, sequencing technologies and findings. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| scRNA-seq (Smart-seq2) | Breast cancer CTCs (single CTCs and CTC clusters, multi-dataset SRP131647/SRP133387/SRP066632/SRP186111) | none | alternative splicing events and 3' UTR length | Smart-seq2 (Takara Bio); FastQC, STAR, salmon, outrigger |
| scRNA-seq (10X Chromium) | 42,225 CTCs from 81 non-metastatic breast cancer patients | none | integrin expression profiles | Chromium system (10X Genomics) |
| scRNA-seq | NSCLC CTCs (3363 single-cell transcriptomes) | none | phenotypic cluster transcriptomes | — |
| scRNA-seq | Neuroblastoma CTCs | none | subgroup gene expression (FOS, RHOA, MIF) and CTC number vs stage | — |
| scRNA-seq (molecular characterisation) | 59 single CTCs from colorectal cancer | none | epithelial, EMT and stem cell gene expression | — |
| scRNA-seq (QIAseq UPX 3' transcriptome) | 24 pooled CTCs from two breast cancer patients | none | apoptotic vs non-apoptotic CTC discrimination | QIAseq UPX 3' transcriptome kit (QIAGEN); CLC Genomics Workbench |
| scRNA-seq (Fluidigm Polaris) | 72 CTCs from 6 breast cancer patients (3 subtypes); control GSE144494 | none | CTC cluster identification, CNV patterns | Fluidigm Polaris system; unCTC, FastQC, RSEM, bowtie, limma, inferCNV |
| scRNA-seq with ex vivo co-culture and pathway blockade | PDAC CTCs co-cultured with myeloid fibroblasts | drug/pathway blockade (CSF1R, CXCR2) | CTC proliferation, clustering, metastatic potential | — |
- ▲ 994 and 836 alternative splicing events identified in single breast cancer CTCs and CTC clusters respectively, with global 3' UTR lengthening in clusters governed by 14 core polyadenylation factors (esp. PPP1CA) 994 vs 836 events; 14 factors
- – Nine distinct integrin expression profiles identified across 42,225 breast cancer CTCs from 81 patients 9 groups
- – Three breast cancer CTC clusters (ER+, HER2+, triple-negative) identified with distinct integrin, platelet degranulation and oncogene expression 3 clusters
- – NSCLC CTC analysis identified distinct phenotypic clusters including epithelial/proliferative, cancer stem cell-like, and mesenchymal immune-evasive/glycolytic populations 3363 transcriptomes
- ▼ Blockade of CSF1R and CXCR2 pathways impaired PDAC CTC proliferation, clustering and metastatic potential
- ▲ Neuroblastoma patients with advanced-stage disease showed higher CTC numbers; two CTC subgroups differing in proliferation vs neuronal injury genes 2 subgroups
- – 18 genes (incl. novel GARS oncogene) linked to breast cancer CTC epithelial phenotype; risk score correlated with high metastasis, poor survival, and sensitivity to AKT-mTOR/CDK inhibitors 18 genes
- – Single-cell profiling of 55 ER+ mBC women revealed ESR1 mutations correlated with time to metastatic relapse and aromatase inhibitor therapy duration n=55
- count 42,225 CTCs (non-metastatic breast cancer patients, integrin profiling)
- count 3363 single-cell CTC transcriptomes (NSCLC phenotypic heterogeneity study)
- count 994 and 836 alternative splicing events (single CTCs vs CTC clusters in breast cancer)
- count 59 single CTCs (CRC molecular characterisation (Kozuka et al.))
- count 55 (ER+ metastatic breast cancer women cohort, ESR1 mutation detection)
- count 72 CTCs from 6 patients (breast cancer 3-subtype scRNA-seq (Fluidigm Polaris))
- count 18 genes (breast cancer CTC epithelial phenotype risk score)
- count nine distinct integrin expression profiles (breast cancer CTC grouping (143/81/8/2/11/38/33/5/154 CTCs per group))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a narrative review article that does not conduct original statistical analyses; it synthesises published primary studies on CTC scRNA-seq workflows and translational applications across cancer types. Quantitative data (cell counts, patient numbers, identified clusters) are drawn entirely from the cited literature. The paper proposes a 12-step CTC-specific scRNA-seq workflow and compiles computational tools reported by primary studies; no inferential statistics, hypothesis tests, or pooled effect estimates are generated by the review authors themselves.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma (linear models for differential expression of RNA-seq counts) | Cited in Table 1 as the differential-expression tool used in the Fluidigm Polaris-based BC CTC study (ER+, HER2+, triple-negative cluster analysis) | 72 CTCs from 6 patients, as quoted from the cited study | not stated |
| DDLK clustering | Cited in Table 1 as a clustering method used in the same Fluidigm Polaris-based BC CTC study | 72 CTCs from 6 patients, as quoted from the cited study | not stated |
| inferCNV (copy-number variation inference from scRNA-seq) | Cited in Table 1 as an analysis tool used in the Fluidigm Polaris-based BC CTC study | 72 CTCs from 6 patients, as quoted from the cited study | not stated |
| WGCNA (weighted gene co-expression network analysis) | Cited in the body text as part of a multi-omics approach (scRNA-seq + WGCNA + CIBERSORT) applied to metastatic PDAC CTCs | — | not stated |
| CIBERSORT (immune cell-type deconvolution) | Cited in the body text as part of the same multi-omics approach in metastatic PDAC CTCs revealing TME immune infiltrate composition | — | not stated |
-
The review is structured as a narrative synthesis without a formally documented, reproducible search strategy or inclusion/exclusion criteria↳ Could also: A PRISMA-guided systematic review with pre-registered eligibility criteria, and where data allow, a meta-analysis of quantitative outcomes (e.g., CTC count versus survival hazard ratios) — Formal study selection criteria and quantitative pooling would increase reproducibility and allow estimation of aggregate effect sizes across the heterogeneous primary studies summarised here
-
Differential expression across CTC subpopulations in the cited studies was assessed with limma applied to log-transformed scRNA-seq counts↳ Could also: Negative-binomial–based methods such as DESeq2 or edgeR, or pseudo-bulk aggregation of per-patient counts before differential testing — Negative-binomial models directly accommodate the discrete, overdispersed count nature of scRNA-seq data; pseudo-bulk approaches additionally account for patient-level correlation when CTCs are drawn from multiple individuals, reducing false-positive rates
-
CTC clustering in the cited studies used a variety of algorithms (DDLK, graph-based methods in 10X pipelines) without reported evaluation of cluster stability↳ Could also: Bootstrap-based stability assessment (e.g., SC3 consensus clustering, clusterCons) applied alongside the primary clustering method — Stability metrics help distinguish robust biologically meaningful CTC subtypes from artefactual partitions, which is especially relevant in rare-cell datasets where n per patient is low
-
Immune cell composition in the TME was estimated via CIBERSORT, a deconvolution method originally developed for bulk RNA-seq reference signatures↳ Could also: Single-cell-aware deconvolution methods such as MuSiC, SCDC, or BayesPrism that use scRNA-seq reference profiles directly — Single-cell-aware methods can leverage the high-resolution CTC and immune-cell reference atlases now available from scRNA-seq datasets, potentially improving specificity of immune infiltrate estimates in the CTC TME context
-
Heterogeneity across studies is described qualitatively, with no quantification of between-study variability in CTC subtype proportions or marker expression levels↳ Could also: Random-effects meta-regression relating CTC subtype frequency or prognostic marker expression to clinical covariates (disease stage, treatment line) across studies — Meta-regression could quantify the degree of between-study heterogeneity (I²) and identify clinical moderators, providing a quantitative complement to the narrative synthesis presented
-
Alternative splicing events and 3′ UTR lengthening in the cited BC study were characterised using outrigger and salmon without reported correction for the large number of splicing events tested (~994 events in single CTCs)↳ Could also: Applying a false discovery rate (FDR) correction such as Benjamini-Hochberg across all tested splicing events before reporting significant events — With nearly 1000 candidate splicing events, an FDR correction would clarify the expected proportion of false positives among the reported findings and facilitate comparison across studies
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope assessment — pmid-41053409
- Title: Single-cell RNA-sequencing of circulating tumour cells: A practical guide to workflow and translational applications.
- Authors: Tieng FYF, Lee LH, Ab Mutalib NS — Cancer Metastasis Rev 2025
- PMID: 41053409 · PMCID: PMC12500777 · DOI: 10.1007/s10555-025-10293-z
- JATS
article-type:review-article(from Europe PMC full-text XML)
Verdict: DROP — non_pipeline (text-mining false positive)
This is a narrative review / "practical guide", not a primary research paper. It performs no original bioinformatic analysis and reports no pipeline-derived numeric result, figure or table value that could be regenerated and compared. There is therefore nothing in scope to reproduce.
Evidence
- Article type is
review-article(full-text JATS XML). - Abstract self-describes as a review: "We address this gap by proposing a 12-step CTC-specific scRNA-seq workflow… with a detailed compilation of data analysis tools… This review supports these goals by guiding methods, informing tool selection and promoting data sharing for reproducibility."
- No self-analysis language anywhere in the full text — grep for
we analysed/analyzed/performed/applied,our analysis/reanalysis,we re-analysed,in this study wereturns zero hits. There is no Methods, Results, or Data-Availability statement describing an analysis the authors ran. - All four figures are conceptual schematics, not data plots:
- Fig. 1 — Timeline of CTC scRNA-seq milestones (2012–2024)
- Fig. 2 — A practical twelve-step workflow (diagram)
- Fig. 3 — Workflow and commonly used tools (diagram; this is where the string "bcl2fastq" appears, as the named demultiplexing step)
- Fig. 4 — Emerging research frontiers (diagram)
- All seven tables are literature / tool compilations, not result tables:
- Table 1 — Overview of published scRNA-seq studies on CTCs → lists ~20 third-party GEO accessions (GSE109761, GSE114704, GSE118389, GSE123904, GSE125449, GSE132257, GSE139829, GSE143791, GSE144494, GSE144561, GSE157745, GSE178341, GSE255299, GSE38495, GSE51372, GSE51827, GSE67980, GSE72056, GSE75367 …). GSE144494 (Aceto-lab HR+ breast-cancer CTC dataset, "135 single CTCs or CTC clusters from 45 patients", CTC-iChip + Smart-seq) is one row of a survey table — not the authors' own data.
- Table 4 — Commonly used tools for pre-/post-processing → lists the tool
bcl2fastq with three resource URLs, one of which is
https://github.com/brwnj/bcl2fastq(a generic Snakemake wrapper around Illuminabcl2fastq2). This is a tool-catalogue URL, not the paper's code.
Why the harvester flagged it (and why that is a false positive)
The link-miner picked up the first GitHub URL in the tool table
(brwnj/bcl2fastq) as code_url and the first GEO accession that resolves
(GSE144494) as data_accession. Both are real and public, but neither is tied
to any reported result: the review neither produced brwnj/bcl2fastq nor analysed
GSE144494. Running that demultiplexer on that dataset would reproduce nothing
the paper claims — there is no expected output to match against. This is exactly
the non_pipeline class in SCREENING.md ("not actually a computational-pipeline
reproduction (text-mining false positive)"); it overlaps with no_expected_result
(no pinnable reported value) — non_pipeline is the more precise root cause.
In-scope results
None. No pipeline-derived result is reported by this paper.
Out-of-scope (not attempted, with reason)
- Everything. The paper is a review; its contributions are a proposed workflow, tool/dataset compilations, and discussion — all manual/expository, none computational. Per HARD RULE 2 (scope to pipeline-derived results) and HARD RULE 6 (drops are valid), the room stops here without spending «our HPC» compute.
No individual results have been recorded for this entry yet.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
PMID 41053409 is a narrative review (Cancer Metastasis Rev 2025, JATS article-type=review-article) that proposes a 12-step CTC scRNA-seq workflow and compiles existing tools/datasets — it reports no original computed value. The harvested code_url (brwnj/bcl2fastq, a generic demultiplexer in Table 4) and data_accession (GSE144494, a third-party dataset in Table 1) are catalogue references, not the authors' analysis — a classic text-mining false positive. There is nothing to reproduce and no fabrication concern; both artifacts are publicly resolvable, so this is a clean category drop (non_pipeline), not a data/code-availability or authors'-side defect. Severity is nil and the (proposed-workflow) conclusion is untouched by reproduction.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.