Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Single-cell RNA-sequencing of circulating tumour cells: A practical guide to workflow and translational applications.

Cancer Metastasis Rev · 2025
L1 No computation 2/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5
✓ What held up
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
Reproduction agent’s raw note

DROP (non_pipeline) - honest category drop, no fabrication concern. PMID 41053409 ('Single-cell RNA-sequencing of circulating tumour cells: A practical guide to workflow and translational applications', Tieng/Lee/Ab Mutalib, Cancer Metastasis Rev 2025) is a NARRATIVE REVIEW (Europe PMC full-text JATS article-type=review-article), not a primary research paper. It reports NO original pipeline-derived result and so has no claim to reproduce: there is no Methods/Results/Data-Availability of an analysis the authors ran, zero self-analysis language in the full text (grep for 'we analysed/performed/applied/our (re)analysis' = 0 hits), all 4 figures are conceptual schematics (milestone timeline; the proposed 12-step CTC scRNA-seq workflow; a workflow+tools diagram; emerging frontiers), and all 7 tables are literature/tool compilations. The harvester's code_url (github.com/brwnj/bcl2fastq) is a tool-catalogue URL in Table 4 ('Commonly used tools for pre-/post-processing'), one of three resource links for the generic Illumina demultiplexer bcl2fastq - not the authors' code. The harvester's data_accession (GSE144494) is one of ~20 third-party GEO datasets surveyed in Table 1 ('Overview of published scRNA-seq studies on CTCs') - the Aceto-lab HR+ breast-cancer CTC dataset (135 single CTCs/clusters from 45 patients, CTC-iChip + Smart-seq), not the authors' own data. Running that demultiplexer on that dataset would reproduce nothing the review claims; there is no expected output to match (overlaps no_expected_result; non_pipeline is the precise root cause). This is a text-mining false positive: code+data are both publicly resolvable, so it is NOT a data/code-availability failure and NOT a fabrication case - the paper simply reports no computed value to check. Decision made from full-text inspection on the «host» control plane; correctly NO «our HPC» compute was spent (HARD RULE 2 scope-to-pipeline, RULE 3 80/20, RULE 6 drops-are-valid). NOT ATTEMPTED: any pipeline run - there is none to attempt. Artifacts: scope.md, AUDIT.md, original/claims.tsv (empty w/ rationale), reproduction/agreement.json.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-14 ⛓ 6d2aee2bcd63
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

This review addresses the knowledge gap created by unstandardised protocols and fragmented resources in circulating tumour cell (CTC) scRNA-seq research by proposing a standardised 12-step CTC-specific scRNA-seq workflow and evaluating analytical tools for their suitability in CTC research.

Core claims
  • A 12-step CTC-specific scRNA-seq workflow spanning enrichment, single-cell sorting, sequencing, data pre-processing and downstream analysis is proposed to overcome methodological inconsistencies. method
  • scRNA-seq enables deep transcriptomic profiling of CTCs, re-stratifying subtypes and detecting rare new subpopulations beyond bulk sequencing. finding
  • Hybrid cells, fusion products of tumour and normal/immune cells, represent a novel research frontier complicating tumour heterogeneity and immune response understanding. finding
  • Integration of machine learning into scRNA-seq workflows enhances CTC clustering, cell identification and heterogeneity analysis. method
  • A compiled comparison of scRNA-seq computational tools evaluates their pros, cons and suitability for CTC research. resource
  • CTC scRNA-seq reveals molecular mechanisms of EMT, metastasis, immune evasion and therapy resistance across cancer types. mechanism
  • Table 1 provides a comprehensive overview of published CTC scRNA-seq studies across cancer types, enrichment methods, sequencing technologies and findings. resource
Experimental setups
Assay System Perturbation Readout Platform
scRNA-seq (Smart-seq2) Breast cancer CTCs (single CTCs and CTC clusters, multi-dataset SRP131647/SRP133387/SRP066632/SRP186111) none alternative splicing events and 3' UTR length Smart-seq2 (Takara Bio); FastQC, STAR, salmon, outrigger
scRNA-seq (10X Chromium) 42,225 CTCs from 81 non-metastatic breast cancer patients none integrin expression profiles Chromium system (10X Genomics)
scRNA-seq NSCLC CTCs (3363 single-cell transcriptomes) none phenotypic cluster transcriptomes
scRNA-seq Neuroblastoma CTCs none subgroup gene expression (FOS, RHOA, MIF) and CTC number vs stage
scRNA-seq (molecular characterisation) 59 single CTCs from colorectal cancer none epithelial, EMT and stem cell gene expression
scRNA-seq (QIAseq UPX 3' transcriptome) 24 pooled CTCs from two breast cancer patients none apoptotic vs non-apoptotic CTC discrimination QIAseq UPX 3' transcriptome kit (QIAGEN); CLC Genomics Workbench
scRNA-seq (Fluidigm Polaris) 72 CTCs from 6 breast cancer patients (3 subtypes); control GSE144494 none CTC cluster identification, CNV patterns Fluidigm Polaris system; unCTC, FastQC, RSEM, bowtie, limma, inferCNV
scRNA-seq with ex vivo co-culture and pathway blockade PDAC CTCs co-cultured with myeloid fibroblasts drug/pathway blockade (CSF1R, CXCR2) CTC proliferation, clustering, metastatic potential
Key results
  • 994 and 836 alternative splicing events identified in single breast cancer CTCs and CTC clusters respectively, with global 3' UTR lengthening in clusters governed by 14 core polyadenylation factors (esp. PPP1CA) 994 vs 836 events; 14 factors
  • Nine distinct integrin expression profiles identified across 42,225 breast cancer CTCs from 81 patients 9 groups
  • Three breast cancer CTC clusters (ER+, HER2+, triple-negative) identified with distinct integrin, platelet degranulation and oncogene expression 3 clusters
  • NSCLC CTC analysis identified distinct phenotypic clusters including epithelial/proliferative, cancer stem cell-like, and mesenchymal immune-evasive/glycolytic populations 3363 transcriptomes
  • Blockade of CSF1R and CXCR2 pathways impaired PDAC CTC proliferation, clustering and metastatic potential
  • Neuroblastoma patients with advanced-stage disease showed higher CTC numbers; two CTC subgroups differing in proliferation vs neuronal injury genes 2 subgroups
  • 18 genes (incl. novel GARS oncogene) linked to breast cancer CTC epithelial phenotype; risk score correlated with high metastasis, poor survival, and sensitivity to AKT-mTOR/CDK inhibitors 18 genes
  • Single-cell profiling of 55 ER+ mBC women revealed ESR1 mutations correlated with time to metastatic relapse and aromatase inhibitor therapy duration n=55
Key statistics
  • count 42,225 CTCs (non-metastatic breast cancer patients, integrin profiling)
  • count 3363 single-cell CTC transcriptomes (NSCLC phenotypic heterogeneity study)
  • count 994 and 836 alternative splicing events (single CTCs vs CTC clusters in breast cancer)
  • count 59 single CTCs (CRC molecular characterisation (Kozuka et al.))
  • count 55 (ER+ metastatic breast cancer women cohort, ESR1 mutation detection)
  • count 72 CTCs from 6 patients (breast cancer 3-subtype scRNA-seq (Fluidigm Polaris))
  • count 18 genes (breast cancer CTC epithelial phenotype risk score)
  • count nine distinct integrin expression profiles (breast cancer CTC grouping (143/81/8/2/11/38/33/5/154 CTCs per group))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a narrative review article that does not conduct original statistical analyses; it synthesises published primary studies on CTC scRNA-seq workflows and translational applications across cancer types. Quantitative data (cell counts, patient numbers, identified clusters) are drawn entirely from the cited literature. The paper proposes a 12-step CTC-specific scRNA-seq workflow and compiles computational tools reported by primary studies; no inferential statistics, hypothesis tests, or pooled effect estimates are generated by the review authors themselves.

Replicationunclear Sample sizeNo original power calculation or sample-size justification is provided by the review; sample sizes are quoted verbatim from cited primary studies and range widely (e.g., 24 pooled CTCs from 2 patients to 42,225 CTCs from 81 patients) GroupsNo original group comparisons performed; review synthesises CTC subtype and cluster comparisons across multiple cancer types as reported in primary literature Pairingna Randomization/blindingna Dispersionnone
Statistical tests used
Test Applied to n Assumptions
limma (linear models for differential expression of RNA-seq counts) Cited in Table 1 as the differential-expression tool used in the Fluidigm Polaris-based BC CTC study (ER+, HER2+, triple-negative cluster analysis) 72 CTCs from 6 patients, as quoted from the cited study not stated
DDLK clustering Cited in Table 1 as a clustering method used in the same Fluidigm Polaris-based BC CTC study 72 CTCs from 6 patients, as quoted from the cited study not stated
inferCNV (copy-number variation inference from scRNA-seq) Cited in Table 1 as an analysis tool used in the Fluidigm Polaris-based BC CTC study 72 CTCs from 6 patients, as quoted from the cited study not stated
WGCNA (weighted gene co-expression network analysis) Cited in the body text as part of a multi-omics approach (scRNA-seq + WGCNA + CIBERSORT) applied to metastatic PDAC CTCs not stated
CIBERSORT (immune cell-type deconvolution) Cited in the body text as part of the same multi-omics approach in metastatic PDAC CTCs revealing TME immune infiltrate composition not stated
Approaches that could also have been used
  • The review is structured as a narrative synthesis without a formally documented, reproducible search strategy or inclusion/exclusion criteria
    Could also: A PRISMA-guided systematic review with pre-registered eligibility criteria, and where data allow, a meta-analysis of quantitative outcomes (e.g., CTC count versus survival hazard ratios) — Formal study selection criteria and quantitative pooling would increase reproducibility and allow estimation of aggregate effect sizes across the heterogeneous primary studies summarised here
  • Differential expression across CTC subpopulations in the cited studies was assessed with limma applied to log-transformed scRNA-seq counts
    Could also: Negative-binomial–based methods such as DESeq2 or edgeR, or pseudo-bulk aggregation of per-patient counts before differential testing — Negative-binomial models directly accommodate the discrete, overdispersed count nature of scRNA-seq data; pseudo-bulk approaches additionally account for patient-level correlation when CTCs are drawn from multiple individuals, reducing false-positive rates
  • CTC clustering in the cited studies used a variety of algorithms (DDLK, graph-based methods in 10X pipelines) without reported evaluation of cluster stability
    Could also: Bootstrap-based stability assessment (e.g., SC3 consensus clustering, clusterCons) applied alongside the primary clustering method — Stability metrics help distinguish robust biologically meaningful CTC subtypes from artefactual partitions, which is especially relevant in rare-cell datasets where n per patient is low
  • Immune cell composition in the TME was estimated via CIBERSORT, a deconvolution method originally developed for bulk RNA-seq reference signatures
    Could also: Single-cell-aware deconvolution methods such as MuSiC, SCDC, or BayesPrism that use scRNA-seq reference profiles directly — Single-cell-aware methods can leverage the high-resolution CTC and immune-cell reference atlases now available from scRNA-seq datasets, potentially improving specificity of immune infiltrate estimates in the CTC TME context
  • Heterogeneity across studies is described qualitatively, with no quantification of between-study variability in CTC subtype proportions or marker expression levels
    Could also: Random-effects meta-regression relating CTC subtype frequency or prognostic marker expression to clinical covariates (disease stage, treatment line) across studies — Meta-regression could quantify the degree of between-study heterogeneity (I²) and identify clinical moderators, providing a quantitative complement to the narrative synthesis presented
  • Alternative splicing events and 3′ UTR lengthening in the cited BC study were characterised using outrigger and salmon without reported correction for the large number of splicing events tested (~994 events in single CTCs)
    Could also: Applying a false discovery rate (FDR) correction such as Benjamini-Hochberg across all tested splicing events before reporting significant events — With nearly 1000 candidate splicing events, an FDR correction would clarify the expected proportion of false positives among the reported findings and facilitate comparison across studies
Software: FastQC · STAR · salmon · outrigger · CLC Genomics Workbench · unCTC · RSEM · bowtie · limma · DDLK clustering · inferCNV · WGCNA · CIBERSORT

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
10
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope assessment — pmid-41053409

  • Title: Single-cell RNA-sequencing of circulating tumour cells: A practical guide to workflow and translational applications.
  • Authors: Tieng FYF, Lee LH, Ab Mutalib NS — Cancer Metastasis Rev 2025
  • PMID: 41053409 · PMCID: PMC12500777 · DOI: 10.1007/s10555-025-10293-z
  • JATS article-type: review-article (from Europe PMC full-text XML)

Verdict: DROP — non_pipeline (text-mining false positive)

This is a narrative review / "practical guide", not a primary research paper. It performs no original bioinformatic analysis and reports no pipeline-derived numeric result, figure or table value that could be regenerated and compared. There is therefore nothing in scope to reproduce.

Evidence

  1. Article type is review-article (full-text JATS XML).
  2. Abstract self-describes as a review: "We address this gap by proposing a 12-step CTC-specific scRNA-seq workflow… with a detailed compilation of data analysis tools… This review supports these goals by guiding methods, informing tool selection and promoting data sharing for reproducibility."
  3. No self-analysis language anywhere in the full text — grep for we analysed/analyzed/performed/applied, our analysis/reanalysis, we re-analysed, in this study we returns zero hits. There is no Methods, Results, or Data-Availability statement describing an analysis the authors ran.
  4. All four figures are conceptual schematics, not data plots:
    • Fig. 1 — Timeline of CTC scRNA-seq milestones (2012–2024)
    • Fig. 2 — A practical twelve-step workflow (diagram)
    • Fig. 3 — Workflow and commonly used tools (diagram; this is where the string "bcl2fastq" appears, as the named demultiplexing step)
    • Fig. 4 — Emerging research frontiers (diagram)
  5. All seven tables are literature / tool compilations, not result tables:
    • Table 1 — Overview of published scRNA-seq studies on CTCs → lists ~20 third-party GEO accessions (GSE109761, GSE114704, GSE118389, GSE123904, GSE125449, GSE132257, GSE139829, GSE143791, GSE144494, GSE144561, GSE157745, GSE178341, GSE255299, GSE38495, GSE51372, GSE51827, GSE67980, GSE72056, GSE75367 …). GSE144494 (Aceto-lab HR+ breast-cancer CTC dataset, "135 single CTCs or CTC clusters from 45 patients", CTC-iChip + Smart-seq) is one row of a survey table — not the authors' own data.
    • Table 4 — Commonly used tools for pre-/post-processing → lists the tool bcl2fastq with three resource URLs, one of which is https://github.com/brwnj/bcl2fastq (a generic Snakemake wrapper around Illumina bcl2fastq2). This is a tool-catalogue URL, not the paper's code.

Why the harvester flagged it (and why that is a false positive)

The link-miner picked up the first GitHub URL in the tool table (brwnj/bcl2fastq) as code_url and the first GEO accession that resolves (GSE144494) as data_accession. Both are real and public, but neither is tied to any reported result: the review neither produced brwnj/bcl2fastq nor analysed GSE144494. Running that demultiplexer on that dataset would reproduce nothing the paper claims — there is no expected output to match against. This is exactly the non_pipeline class in SCREENING.md ("not actually a computational-pipeline reproduction (text-mining false positive)"); it overlaps with no_expected_result (no pinnable reported value) — non_pipeline is the more precise root cause.

In-scope results

None. No pipeline-derived result is reported by this paper.

Out-of-scope (not attempted, with reason)

  • Everything. The paper is a review; its contributions are a proposed workflow, tool/dataset compilations, and discussion — all manual/expository, none computational. Per HARD RULE 2 (scope to pipeline-derived results) and HARD RULE 6 (drops are valid), the room stops here without spending «our HPC» compute.

No individual results have been recorded for this entry yet.

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 75/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5

PMID 41053409 is a narrative review (Cancer Metastasis Rev 2025, JATS article-type=review-article) that proposes a 12-step CTC scRNA-seq workflow and compiles existing tools/datasets — it reports no original computed value. The harvested code_url (brwnj/bcl2fastq, a generic demultiplexer in Table 4) and data_accession (GSE144494, a third-party dataset in Table 1) are catalogue references, not the authors' analysis — a classic text-mining false positive. There is nothing to reproduce and no fabrication concern; both artifacts are publicly resolvable, so this is a clean category drop (non_pipeline), not a data/code-availability or authors'-side defect. Severity is nil and the (proposed-workflow) conclusion is untouched by reproduction.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

48.8 k
tokens (I/O) · 2.3 M incl. cache
4 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.