Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Single-cell multiomics analysis of chronic myeloid leukemia links cellular heterogeneity to therapy response.

Elife · 2024
L1 75/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Same input data as the authors
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
75/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 45% of all assessed papers rank 612 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction, described well enough to reproduce, compute ran on «our HPC» («job»: GEOquery + GSVA ssGSEA + t-test on GSE14671). The paper's validation claim on the public bulk cohort reproduces well: C1 (cohort split 18 non-responders / 41 responders of 59) is an EXACT match from the GSE14671 phenoData; C2 ('primitive cells enriched in non-responders') reproduces in DIRECTION and SIGNIFICANCE -- all 6 HSC/primitive signatures score higher in non-responders and 4/6 are significant at p<0.05 (top JAATINEN_HSC t=3.248 p=0.0027). C2 is graded partial because we used an open ssGSEA surrogate (brief P16) rather than the authors' CIBERSORTx, so the exact '>threefold' cell-fraction ratio is not directly reproduced -- a fraction ratio and an enrichment-score difference are different quantities. NOT attempted (the harder ~20%): exact CIBERSORTx with the authors' 11-cluster CITE-seq signature matrix (needs gated Stanford server + processing GSE236233/GSE173076), and all single-cell results from the GSE236233 CITE-seq data (UMAP/11 Leiden clusters, 14274 Lin-CD34+ cells, 71-gene pan-CML signature, 244/180 BCR::ABL1+/- DEGs, CD26/CD35 LSC surface markers). Note: the room-pinned code_url github.com/KarlssonG/cellradar is a small JS gene-list widget, NOT the paper's pipeline; the real code is on OSF osf.io/ns2j8 (Scarf/CellRanger/DESeq2/CIBERSORTx) -- corrected in scope.md. All grades provisional; human reviewer signs off in AUDIT.md.

💻 Code ↗ 🗄 Data: GSE14671

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 75
    assessed: 2026-06-20 ⛓ f8816958a826
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-20
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether single-cell multiomic (CITE-seq) profiling of CML stem/progenitor cell heterogeneity at diagnosis can reveal molecularly distinct leukemic stem cells (LSCs) versus residual normal hematopoietic stem cells (HSCs), and whether this heterogeneity predicts subsequent response to tyrosine kinase inhibitor (TKI) therapy.

Core claims
  • CITE-seq multiomics reveals that each CML patient harbors a unique composition of stem and progenitor cells at diagnosis finding
  • Patients with treatment failure at 12 months have a markedly higher abundance of molecularly defined primitive cells at diagnosis than optimal responders finding
  • BCR::ABL1+ LSCs and BCR::ABL1- HSCs can be distinctly separated by cell surface markers as CD26+CD35- and CD26-CD35+, respectively finding
  • The LSC/HSC ratio is higher in patients with prospective treatment failure than optimal responders, both at diagnosis and after 3 months of TKI therapy finding
  • A 71-gene pan-CML signature (50 up-regulated, 21 down-regulated) is shared across CML stem and progenitor clusters relative to normal bone marrow finding
  • CML Lin-CD34+ cells are enriched for basophil/mast cell, MEP, and MkP clusters and depleted for primitive, MPP1/2, and ly/pDC/mono clusters compared to normal bone marrow finding
  • Scarf-based projection of CML cells onto a normal bone marrow reference enables label-transfer annotation of cell heterogeneity without confounding from BCR::ABL1 transformation or batch effects method
  • A CITE-seq reference map of Lin-CD34+ normal bone marrow (4,696 cells, 11 clusters) from two age-matched donors resource
Experimental setups
Assay System Perturbation Readout Platform
CITE-seq (scRNA-seq + ADT surface proteomics, >40 antibodies) Lin-CD34+ (and CD34+CD38-/low) bone marrow cells from 9 CML patients at diagnosis none (diagnostic baseline, retrospectively stratified by later TKI response) global gene expression profile and surface marker (ADT) abundance, cluster composition
CITE-seq (scRNA-seq + ADT) Lin-CD34+ normal bone marrow (nBM) cells from 2 age-matched healthy donors none gene expression and ADT expression used to build reference UMAP with Leiden clusters
FACS sorting CML bone marrow LSC/progenitor populations none isolation of Lin-CD34+/CD38-/low populations prior to CITE-seq
Differential gene expression (bioinformatic) analysis CML vs nBM Lin-CD34+ clusters (all 9 patients) none cluster-specific and shared differentially expressed genes (adjusted p<.01, log2FC>1/<-1)
Cell projection / label transfer (Scarf) CML single cells mapped onto nBM reference none cluster identity assignment per CML cell Scarf
Key results
  • Primitive cluster abundance at diagnosis is fourfold higher in treatment-failure patients than optimal responders ~4-fold
  • MEP and erythroid progenitor clusters are more abundant in optimal responders than treatment failures ~3-fold
  • BCR::ABL1+ LSCs and BCR::ABL1- HSCs separate as CD26+CD35- and CD26-CD35+ populations respectively
  • LSC/HSC ratio is elevated in treatment-failure patients at diagnosis and at 3 months of TKI therapy
  • 384 genes differentially expressed uniquely between CML and nBM primitive clusters, the largest change among all cluster comparisons 384 genes
  • 71-gene pan-CML signature identified as shared across all cluster-wise CML vs nBM comparisons (50 up, 21 down) 71 genes
  • CML samples show enrichment of basophil/mast cell, MEP, MkP clusters and depletion of primitive, MPP1/2, ly/pDC/Mono clusters relative to nBM
  • Warning-category patients (CML5-7) show heterogeneous profiles, with CML5 resembling treatment failures and CML6/7 resembling optimal responders
Key statistics
  • count 14,274 cells (Total Lin-CD34+ CML cells profiled by CITE-seq across 9 patients at diagnosis)
  • count 4,696 cells (Normal bone marrow Lin-CD34+ reference cells from 2 donors)
  • count 11 clusters (Number of Leiden clusters identified in the nBM reference UMAP)
  • fold_change fourfold higher (Primitive cluster abundance in treatment-failure vs optimal responder patients at diagnosis)
  • fold_change ~threefold higher (MEP and erythroid cluster abundance in optimal responders vs treatment failures)
  • count 384 differentially expressed genes (Unique DEGs between CML and nBM primitive clusters)
  • count 71 genes (50 up, 21 down) (Pan-CML gene signature shared across all cluster DEG comparisons)
  • pvalue adjusted p-value <.01, log2 fold change >1/<-1 (Threshold used to define differentially expressed genes)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used single-cell CITE-seq (combined scRNA-seq and antibody-derived tag protein profiling) on bone marrow cells from nine CML patients and two normal bone marrow donors. Cellular heterogeneity was characterized via Leiden clustering and UMAP dimensionality reduction after projecting CML cells onto a normal bone marrow reference (label transfer), and differential gene expression was assessed between corresponding CML and normal clusters using an adjusted p-value and log2 fold-change threshold. Patient-level cluster abundance was compared descriptively (e.g., fold-change bar plots with error bars) across therapy-response groups (optimal, warning, treatment failure) without a stated inferential test for these group comparisons in the excerpted text.

Replicationbiological Sample sizeSample sizes given per analysis (e.g., 9 CML patients, 2 nBM donors, cell counts per UMAP), but no formal power calculation is described in the provided text GroupsCML patient cells vs normal bone marrow reference cells; and CML patients stratified by therapy outcome (optimal, warning, treatment failure) Pairingunpaired Randomization/blindingnot stated DispersionSD Effect sizesyes Multiplicity correctionAdjusted p-value (specific correction method, e.g., Benjamini-Hochberg, not named in the excerpted text)
Statistical tests used
Test Applied to n Assumptions
Differential expression testing (specific statistical test/model not named in the provided text) Pairwise comparisons of gene expression between each CML cluster and its corresponding normal bone marrow (nBM) cluster (Figure 2C, 2D; Supplementary files 5, 7) 9 CML patients (14,274 cells) vs 2 nBM donors (4,696 cells), pooled by cluster not stated
Leiden clustering Identification of 11 cell clusters within Lin-CD34+ nBM reference (Figure 1B) and CML samples 4,696 nBM cells; 14,274 CML cells not stated
UMAP (dimensionality reduction, not an inferential test) Visualization of cell heterogeneity (Figures 1B, 2A) na na
Descriptive fold-change comparison with standard deviation error bars (no named inferential test) Cluster distribution (%) of CML Lin-CD34+ cells vs nBM (Figure 2B) n=9 CML patients not stated
Approaches that could also have been used
  • Differential expression between CML and nBM clusters was reported using an adjusted p-value and log2 fold-change cutoff, without the underlying test being named in the available text.
    Could also: A model-based single-cell DE method such as a Wilcoxon rank-sum test (as implemented in Seurat/Scanpy), MAST, or a pseudobulk approach with DESeq2/edgeR — Naming the specific DE test and correction method (e.g., Benjamini-Hochberg FDR) explicitly would let readers evaluate the assumptions and error-rate control being applied to the multiple cluster comparisons.
  • Cluster abundance differences between therapy-outcome groups (optimal vs. treatment failure) were described qualitatively (e.g., 'fourfold higher,' 'threefold higher') without an accompanying statistical test.
    Could also: A non-parametric test such as Mann-Whitney U (given small group sizes: n=4 optimal vs n=2 failure) or a mixed-effects model accounting for patient-level variability — A formal test with an accompanying p-value or confidence interval would help quantify the uncertainty around these fold differences, which is particularly relevant given the small per-group patient numbers.
  • Cluster distribution comparisons (CML vs nBM) were summarized as log2 fold change with standard deviation error bars (Figure 2B).
    Could also: Reporting a 95% confidence interval or using a bootstrap resampling approach for the fold-change estimates — A confidence interval directly conveys the precision of the estimated fold change and can be more readily interpreted alongside the effect size than SD alone, especially with n=9 patients.
  • No explicit mention of correction for multiple comparisons across the many pairwise cluster-vs-cluster DE analyses beyond an adjusted p-value threshold.
    Could also: Explicitly reporting the FDR method (e.g., Benjamini-Hochberg) and threshold used across all cluster-pair tests, or a global permutation-based multiple-testing correction — Specifying the exact multiplicity control approach clarifies how the family-wise or false-discovery rate was managed across the many simultaneous cluster comparisons performed.
  • Patient-level heterogeneity in cluster composition was described narratively (e.g., CML5 resembling failures, CML6/7 resembling optimal responders) rather than through a formal clustering or classification statistic.
    Could also: A quantitative similarity/distance metric (e.g., hierarchical clustering with a distance measure, or a discriminant analysis) applied to patient-level cluster-composition vectors — A quantitative grouping metric could complement the visual/narrative comparison and provide a reproducible measure of how closely each patient's profile aligns with the optimal or failure groups.
Software: Scarf (for reference projection/label transfer) · CITE-seq analysis pipeline (specific package, e.g., Seurat/Scanpy, not named in excerpted text)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39503729

Paper: Single-cell multiomics analysis of chronic myeloid leukemia links cellular heterogeneity to therapy response. PMID 39503729 · PMCID PMC11540304 · DOI 10.7554/eLife.92074 · eLife 2024.

Code & data, as actually published (corrects the room manifest)

The room was instantiated with code_url = github.com/KarlssonG/cellradar and data_accession = GSE14671. After reading the paper:

  • KarlssonG/cellradar is NOT the paper's analysis pipeline. It is a small JavaScript/HTML browser widget ("identify and visualize cell type enrichment in your gene lists", 88% JS, 9 commits, hosted at karlssonG.github.io/cellradar). The paper's actual analysis code lives on OSF (osf.io/ns2j8) and uses Scarf v0.18.12, Cell Ranger 3.0.2, DESeq2, and CIBERSORTx.
  • Data accessions in the paper:
    • GSE236233 — the paper's own CITE-seq single-cell data (Lin−CD34+ CML, 9 patients).
    • GSE173076 — normal bone-marrow CITE-seq reference (the 11-cluster Lin−CD34+ nBM map).
    • GSE14671bulk CD34+ Affymetrix HG-U133 Plus 2.0 microarray of 59 chronic-phase CML patients (Radich/McWeeney imatinib-response cohort). This is the independent validation cohort the paper deconvolves with CIBERSORTx. This is the accession the room pinned, and it is public and downloadable.

In scope (pipeline-derived, attempted)

The validation analysis that is anchored on the public GSE14671 bulk data:

  • C1 — cohort composition. Paper: imatinib non-responders n=18/59 vs responders n=41/59 (Results & Fig 3E). GSE14671 GEO record: training set 12 NR + 24 R (n=36) + validation set 6 NR + 17 R (n=23) = 18 NR / 41 R / 59. Directly checkable from the GSE14671 phenoData → expect EXACT.
  • C2 — "primitive cells enriched in non-responders". Paper (Fig 3E): CIBERSORTx deconvolution of GSE14671 using their 11-cluster CITE-seq signature matrix shows a >threefold, statistically significant (Student t-test p<0.05) enrichment of "primitive" cells in non-responders vs responders. We reproduce the direction and significance of this claim by scoring each of the 59 bulk profiles for a primitive/HSC transcriptional program (ssGSEA, GSVA) using published HSC/primitive gene signatures, then comparing non-responders vs responders (t-test + fold change). This is the brief's P16 "third-party tool/approach on the paper's own data" — an equally valid reproduction of the deconvolution-derived clinical claim.

Out of scope / NOT attempted (the hard ~20%)

  • Exact CIBERSORTx reproduction with the paper's own signature matrix. That needs (a) processing GSE236233/GSE173076 CITE-seq into the 11-cluster signature matrix and (b) the Stanford CIBERSORTx web server (registration-gated, non-scriptable). We use an open ssGSEA primitive score instead → so C2 is graded on direction + significance, not the exact 3.0× fold change. Flagged accordingly.
  • All single-cell results (UMAP/Leiden 11 clusters, 14,274 Lin−CD34+ cells, 71-gene pan-CML signature, 244/180 BCR::ABL1± DEGs, CD26/CD35 surface-marker LSC findings): derive from GSE236233 CITE-seq + wet-lab flow/sorting — not attempted in this 80/20 pass.
  • cellradar JS widget: not a reproduction target (no quantitative claim depends on it).

Pipeline named per in-scope result

  • C1: GEO metadata parse (GEOquery phenoData).
  • C2: bulk microarray normalization (GEOquery series matrix, GPL570) + ssGSEA primitive scoring (GSVA) + t-test — a deconvolution surrogate for CIBERSORTx.
Figures / tables: Fig 3D
C1
Reported
imatinib non-responders n=18/59 vs responders n=41/59 (Results; Fig 3D/E)
Reproduced
GSE14671 series matrix 54675 probes x 59 samples; phenoData group counts NR=18, R=41, total=59
exact
C2
Reported
>threefold, statistically significant (Student t-test p<0.05) enrichment of 'primitive' cells in non-responders vs responders (CIBERSORTx deconvolution of GSE14671; Fig 3D/E)
Reproduced
ssGSEA primitive/HSC scoring (CIBERSORTx surrogate): all 6 HSC/primitive signatures higher in NR; 4/6 significant at p<0.05 (curated primitive t=2.564 p=0.0166; JAATINEN_HSC t=3.248 p=0.0027; EPPERT_HSC_R t=2.503 p=0.0198; EPPERT_CE_HSC_LSC t=2.492 p=0.0201). Direction + significance match; exact 3x fold not reproduced (surrogate).
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 75/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

On the public bulk cohort (GSE14671) the reproduction is strong: C1 (non-responders n=18/41 responders of 59) is an exact metadata match, and C2 ('primitive' cells enriched in non-responders) reproduces in direction and significance (6/6 signatures higher in NR, 4/6 p<0.05, JAATINEN_HSC t=3.248 p=0.0027). The only gap — the exact '>threefold' CIBERSORTx fold change — is on our side: we used an open ssGSEA surrogate because the Stanford CIBERSORTx server is gated and the authors' signature matrix was not rebuilt, so a cell-fraction ratio and an enrichment-score difference are simply different quantities. No fabrication concern; the limitation is documented, not hidden. Overall yellow because the deviation is explainable and the paper's central single-cell conclusions were out of scope this pass.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

266.8 k
tokens (I/O) · 16.7 M incl. cache
122 min
runtime · 0.03 CPU-h
3 GB
peak RAM
1
HPC jobs
hummel
machine