Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Synergism between IL7R and CXCR4 drives BCR-ABL induced transformation in Philadelphia chromosome-positive acute lymphoblastic leukemia.

Nat Commun · 2020
L1 92/100 PQI 97
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -4
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
How its reproducibility compares
92/100
Reproducibility score
1.0 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 83% of all assessed papers rank 179 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough; clean 1:1. The repo (medhaniea/pca-and-heatmap @fbc8d77, by paper co-author Mulaw) is a self-contained R script that runs PCA + Heatplus heatmap on its shipped GSE150784-derived normalized matrix and ships its own reference output pca_3d.pdf. Running the author's script on the author's shipped data reproduced the PCA variance percentages EXACTLY (PC1=57.12, PC2=19.8, PC3=9.18 %), the PC1 control-vs-BCR-ABL separation, and the heatmap column clustering that splits the two arms (C1-C4 all exact). As a stretch anti-fabrication check (C5) I independently regenerated the normalized matrix from GSE150784 deposited raw counts with edgeR: per-gene per-sample profiles for all 38 mappable selected genes correlate r=0.95-1.00 with the shipped values, confirming the matrix is genuine and not fabricated; but the absolute values are NOT byte-reproducible because the exact gene filter is under-specified (deposited matrix has 23997 genes, not regenerable from the stated '>2 CPM in >=6 samples' filter which yields 10882; closest candidates 21856/23519). NOT attempted: all wet-lab biology (IL7R-CXCR4, xenotransplant, antibody therapy) = out of scope; FASTQ->counts alignment (only count csv deposited).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 92
    assessed: 2026-06-17 ⛓ 12a3a844c9c2
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-17
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests whether IL7R and CXCR4 interact on the cell surface to recruit BCR-ABL1 and JAK kinases into close proximity, serving as a molecular platform required for BCR-ABL1-induced malignant transformation and development of Philadelphia chromosome-positive (Ph+) ALL.

Core claims
  • IL7R interacts/colocalizes with CXCR4 on the cell surface, recruiting BCR-ABL1 and JAK kinases into close proximity to form a platform for transformation mechanism
  • BCR-ABL1 kinase inhibition (imatinib) elevates IL7R and CXCR4 expression, and added IL7 (but not CXCL12 or TSLP) rescues transformed cells from inhibitor-induced death finding
  • IL7R expression is specifically required for initiation and maintenance of BCR-ABL1-induced pre-B cell transformation and ALL development finding
  • CXCR4 is required for BCR-ABL1-induced transformation and synergizes with IL7R to prevent pre-B cell differentiation and direct migration finding
  • Anti-IL7R antibody eliminates Ph+ ALL cells, including TKI-resistant cells, and prevents leukemia development in patient-derived xenograft models finding
  • BCR-ABL1 transformation upregulates JAK/STAT, interleukin, and cytokine signaling gene sets relative to wildtype pre-B cells finding
  • BCR-ABL1 is recruited to CXCR4 and can be activated by CXCL12, enabling CXCL12-induced Ca2+ flux dependent on BCR-ABL1 kinase activity mechanism
  • Anti-IL7R antibody may provide alternative treatment for drug-resistant Ph+ ALL resource
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq with PCA, GO and GSEA analysis mouse BM-derived pre-B cell lines (6 BCR-ABL1-transformed vs 6 empty-vector control) BCR-ABL1 overexpression vs empty vector global transcriptome / differential gene expression ReliaPrep RNA Miniprep System
gene expression correlation analysis 68 pediatric Ph+/BCR-ABL+ BCP-ALL patient cohort none IL7R vs CXCR4 expression correlation
flow cytometry (surface staining) BCR-ABL1-transformed mouse pre-B cells; human SUP-15 Ph+ ALL cells imatinib treatment (1 µM, 15 h) surface IL7R and CXCR4 protein expression
quantitative RT-PCR BCR-ABL1-transformed pre-B cells imatinib treatment Il7r, Cxcr4, Jak1, Stat5a transcript levels
cell viability / cell death and cell cycle assay BCR-ABL1-transformed WT pre-B cells imatinib + IL7 / CXCL12 / AMD3100 / TSLP cell survival and cell cycle progression
inducible gene deletion + viability (Cre-ERT2) and PCR BCR-ABL1-transformed Il7rα fl/fl and Cxcr4 fl/fl mouse pre-B cells tamoxifen-induced Cre deletion of Il7rα or Cxcr4 cell viability/survival, deletion confirmation, colony formation
xenotransplantation with luciferase bioimaging and survival NOD-SCID mice injected with BCR-ABL1-transformed Il7rα fl/fl pre-B cells; patient-derived Ph+ ALL xenografts in vivo tamoxifen-induced Il7r deletion; anti-IL7R antibody leukemic burden and mouse survival
proximity ligation assay (PLA), Ca2+ flux flow cytometry, immunoprecipitation, Western blot WT and BCR-ABL1-transformed mouse pre-B cells; human BCR-ABL+ pre-B ALL cells BCR-ABL1 transformation; imatinib/dasatinib; AMD3100; inducible IL7R/CXCR4 deletion IL7R/CXCR4/JAK3 association, CXCL12-induced Ca2+ mobilization, JAK phosphorylation
Key results
  • IL7R and CXCR4 gene expression significantly correlated in 68 Ph+ BCP-ALL patients Spearman r=0.6264
  • Five of eight IL7R-related gene sets significantly upregulated in BCR-ABL1-transformed vs control cells FDR<0.25
  • Imatinib increases surface IL7R and CXCR4 plus Jak1/Stat5a in transformed cells
  • IL7 counteracts imatinib-induced cell death and restores cell cycle; CXCL12, AMD3100 and TSLP do not
  • Inducible Il7rα deletion causes death of BCR-ABL1-transformed pre-B cells in vitro
  • In vivo Il7r deletion reduces leukemic burden and significantly prolongs xenograft mouse survival
  • BCR-ABL1-transformed cells show robust CXCL12-induced Ca2+ flux (absent in WT), blocked by imatinib/dasatinib or Cxcr4 deletion/AMD3100
  • BCR-ABL1-transformed cells show increased IL7R/CXCR4 foci and IL7R-dependent JAK3-CXCR4 association; deletion of CXCR4 or IL7R reduces JAK phosphorylation
Key statistics
  • correlation Spearman r=0.6264, p=0.000000011054191 (IL7R vs CXCR4 expression in Ph+ BCP-ALL patients)
  • pvalue p<0.000000000000001 (viable CD19+GFP+ cells after IL7 withdrawal, BCR-ABL1 vs control)
  • pvalue p=0.0002 (Mantel–Cox log-rank) (survival of NOD-SCID mice with Cre-ERT2 vs ERT2 Il7rα-BCR-ABL1 cells)
  • pvalue p=0.00000000006167 (PLA IL7R/CXCR4 association increase in BCR-ABL1-transformed cells (Fig 4c))
  • pvalue adjusted p=0.0003 / 0.0002 / 0.0002 (Dunnett) (Vehicle vs 1.25/2.5/5 ng/ml IL7 rescue at day 6)
  • pvalue qRT-PCR p=0.000035569652955 (imatinib-induced Il7r/Cxcr4 transcription change)
  • other 5-year disease-free survival 70 ± 12% (children with BCR-ABL1+ leukemia on TKI plus chemotherapy)
  • count 1223 BCP-ALL patients RNA-seq dataset (IL7R/CXCR4 expression comparison across BCP-ALL entities)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study combined RNA-sequencing (with GSEA for pathway enrichment) and a series of in vitro and in vivo cell-biology experiments in mouse and human models. Continuous outcomes across two or more groups were analyzed with unpaired two-sided t-tests or one-way ANOVA with Dunnett's post-hoc correction; correlation between gene-expression levels in patient samples was assessed by Spearman's ρ; and survival in xenograft experiments was compared with the Mantel–Cox log-rank test. Results were reported as exact p-values alongside mean ± SD or mean ± SEM, without formal effect-size estimates or confidence intervals.

Replicationbiological Sample sizeStated per figure (n = 3–21 depending on experiment); no formal a priori power calculation mentioned GroupsBCR-ABL1-transformed vs. EV-transduced pre-B cells; Ph+ patient cohort; conditional knockout vs. control; xenograft survival groups Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionFDR (GSEA, scope: 8 MSigDB gene sets); Dunnett's post-hoc correction (scope: pairwise comparisons vs. vehicle/control within each ANOVA)
Statistical tests used
Test Applied to n Assumptions
Gene Set Enrichment Analysis (GSEA) with signal-to-noise ranking metric and FDR correction Fig. 1b and Supplementary Fig. 1b – IL7R-related KEGG/REACTOME MSigDB gene sets in BCR-ABL1 vs. EV-transduced pre-B cells 6 BCR-ABL1 and 6 EV-transduced cell lines not stated
Two-tailed Spearman rank correlation Fig. 2a – correlation between IL7R and CXCR4 expression levels in Ph+ ALL patients 68 pediatric Ph+ BCP-ALL patients not stated
Unpaired t-test, two-sided Fig. 2c – quantitative RT-PCR of Il7r, Cxcr4, and associated factors under imatinib n = 3 independent samples per group not stated
One-way ANOVA with Dunnett's multiple comparisons (vs. vehicle control) Fig. 2d – cell number after imatinib + varying IL7 concentrations over 6 days n = 3 independent samples per group not stated
Unpaired t-test, two-sided Fig. 2e – cell viability under imatinib ± IL7 or CXCL12 n = 3 independent samples per group not stated
Unpaired t-test, two-sided Fig. 3b – viable CD19+GFP+ cells after IL7 withdrawal (EV vs. BCR-ABL1) n = 7 per group not stated
One-way ANOVA with Dunnett's multiple comparisons (vs. EREt control) Fig. 3e – percentage of living cells at day 5 across four Cre/Tam conditions n = 3 independent samples per group not stated
Mantel–Cox log-rank test Fig. 3g – survival of NOD-SCID xenograft mice injected with Cre-ERT2 vs. ERT2 BCR-ABL1 cells n = 21 per group not stated
Unpaired t-test, two-sided Fig. 4b, 4c, 4d – PLA signal quantification (IL7R/CXCR4 foci and JAK3/CXCR4 association per cell) representative of three independent experiments; per-cell n not stated not stated
Approaches that could also have been used
  • Dispersion was reported as SEM in some figures and SD in others within the same paper
    Could also: Report SD (or 95% CI) uniformly throughout — SD describes the spread of the actual data distribution, whereas SEM reflects uncertainty in the mean estimate and shrinks with larger n; for small-n experiments (n = 3–7) a consistent switch to SD or 95% CI would make the variability of individual observations easier to compare across figures
  • Multiple pairwise comparisons across different figures were each evaluated with a separate unpaired t-test
    Could also: Apply a single ANOVA (or mixed-effects model) per experiment, followed by a family-wise post-hoc correction (e.g., Tukey HSD or Dunnett's), or use a Bonferroni adjustment across related t-tests — Consolidating comparisons within each experiment controls the family-wise error rate and reduces the chance that any individual comparison appears significant by chance when many contrasts are made
  • PLA signal counts per cell (Figs. 4b–d) were compared with unpaired t-tests, treating signals as independent observations
    Could also: Use a linear mixed-effects model with experiment (biological replicate) as a random effect — Cells measured within the same experimental replicate are not fully independent; a mixed-effects model accounts for within-replicate clustering and gives more accurate standard errors when the outcome is counts nested within experiments
  • The GSEA significance threshold was set at FDR < 0.25
    Could also: Apply the more stringent conventional FDR threshold of < 0.05 or < 0.10 and report all tested gene sets — A threshold of 0.25 means up to 25% of significant gene sets could be false positives; stricter thresholds or reporting all eight tested sets with their scores would allow readers to gauge robustness of the enrichment signal
  • The correlation between IL7R and CXCR4 expression was assessed by Spearman's ρ alone
    Could also: Complement with a simple linear regression or scatter-plot with a regression line and 95% CI band — Regression would quantify the magnitude of the association (slope) and its uncertainty, providing additional information beyond the direction and rank-order correlation captured by ρ
  • The RNA-seq differential expression analysis underlying the heatmap (Fig. 1a) does not explicitly state the method used (e.g., DESeq2, edgeR, limma-voom)
    Could also: Explicitly name the differential expression package, model formula, and normalization method (e.g., DESeq2 Wald test with variance-stabilizing transformation) — Stating the exact tool and model makes the analysis reproducible and allows readers to assess how variance estimation, library-size normalization, and dispersion shrinkage were handled for n = 6 vs. 6 samples
Software: GSEA (Broad Institute / MIT / UC) · R2 genomics analysis and visualization platform (r2.amc.nl)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
21
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

NCT00287105 NCT in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
NCT00430118 NCT in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-32581241

Paper: Abdelrasoul et al., Synergism between IL7R and CXCR4 drives BCR-ABL induced transformation in Ph+ ALL, Nat Commun 2020. PMID 32581241. Code: https://github.com/medhaniea/pca-and-heatmap (author: Medhanie A. Mulaw, a paper co-author). Data: GEO GSE150784 — 6 control WT pre-B cell lines (empty vector, "pMIG_plus") vs 6 BCR-ABL-transformed counterparts ("pMIG_BCR_ABL"), bulk RNA-seq, mouse GRCm38.

Pipeline / data linkage

  • GEO ships raw counts: GSE150784_meta.feature.level.counts.csv.gz.
  • Processing (per GEO data_processing): TopHat2/Bowtie2 align → featureCounts → edgeR DGEList, filter genes with >2 CPM in >=6 samples, TMM normalization (R 3.6.1, edgeR 3.18.1).
  • The repo ships the resulting normalized log-expression matrix as Data/exprs_example.txt (genes x 12 samples) and runs PCA + heatmap on it.
  • Sample names in Pheno.txt (6 pMIG_plus + 6 pMIG_BCR_ABL) match GSE150784's 12 samples (GSM4558717–GSM4558728). The repo data IS this paper's data.

In scope (pipeline-derived, deterministic)

id result pipeline how to compare
C1 PCA PC1 variance = 57.12% prcomp(t(exprs), center=TRUE) + summary importance exact
C2 PCA PC2 variance = 19.8% same exact
C3 PCA PC3 variance = 9.18% same exact
C4 Heatmap clustering of Selected_genes (Heatplus annHeatmap2, corrdist + average linkage, scale="row", Row cuth=1.6 / Col cuth=0.1) — column dendrogram separates the two groups; row cluster count deterministic clustering structure/order
C5 (stretch) Regenerate normalized matrix from GEO raw counts via edgeR TMM and confirm it reproduces exprs_example.txt values (provenance / anti-fabrication) edgeR filterByExpr/TMM → log-CPM within-tol

Reference values C1–C3 are read from the repo's shipped output pca_3d.pdf (axis labels). This is a faithful 1:1 of the author's own script on the author's own shipped data (deterministic) — the canonical reproduction for this RU.

Out of scope (wet-lab / not pipeline / not deposited)

  • All biology claims (IL7R–CXCR4 interaction, xenotransplant, antibody therapy): wet-lab, not computational. NOT attempted.
  • Upstream alignment (TopHat2/Bowtie2/featureCounts from FASTQ): raw FASTQ not the deposited artifact here (only count csv deposited); enormous compute, not the repo's scope. NOT attempted (C5 starts from deposited raw counts, not FASTQ).
  • The 3D rendering geometry of pca_3d.pdf (rgl/OpenGL viewing angle): cosmetic, not a quantitative result; we reproduce the PC variance numbers it displays, not pixels.
C1
Reported
PCA PC1 variance 57.12%
Reproduced
57.12%
exact
C2
Reported
PCA PC2 variance 19.8%
Reproduced
19.8%
exact
C3
Reported
PCA PC3 variance 9.18%
Reproduced
9.18%
exact
C1b
Reported
PCA separates control vs BCR-ABL on PC1
Reproduced
all 6 control PC1<0, all 6 BCR-ABL PC1>0
exact
C4
Reported
Heatmap column clustering separates the two arms
Reproduced
column dendro: 6 controls left, 6 BCR-ABL right; 2 row clusters @cuth1.6
exact
C5
Reported
shipped normalized matrix = edgeR/TMM of GSE150784 raw counts
Reproduced
per-gene profiles r=0.95-1.00 vs independent edgeR regen (provenance confirmed, no fabrication); absolute values not byte-reproducible (23997 genes not regenerable from stated filter; gene-varying offset)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 92/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -4

All reported computational claims reproduce exactly 1:1 on the author's shipped GSE150784-derived matrix — PCA variances (57.12/19.8/9.18%), the PC1 control-vs-BCR-ABL separation, and the heatmap column clustering (C1-C4 exact). An anti-fabrication stretch check confirms the shipped matrix is genuine (per-gene r=0.95-1.00 vs independent edgeR), ruling out fabrication. The only wrinkle is on the deposit/authors' side: the exact gene filter is under-specified in GEO, so the absolute normalized values are not byte-reproducible from raw counts (regen PCA 64.97/18.16/7.51%) — a preprocessing documentation gap that does not affect any reproduced figure. Overall a solid, near-clean reproduction.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

156 k
tokens (I/O) · 9.3 M incl. cache
17 min
runtime · 0.02 CPU-h
1.3 GB
peak RAM
3 (1 failed)
HPC jobs
hummel
machine