Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Intrinsic suppression of type I interferon production underlies the therapeutic efficacy of IL-15-producing natural killer cells in B-cell acute lymphoblastic l

J Immunother Cancer · 2023
L1 76/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Reported values are derivable from the shared data
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
76/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 48% of all assessed papers rank 586 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL 1:1 reproduction of the paper's PUBLIC-DATA computational claim. Sudan et al. (JITC 2023) is mainly a wet-lab + mass-cytometry (CyTOF) immunology paper; its primary pipeline (CyTOF immune profiling, Figs 1-2) is normalized with the brief's cited repo github.com/nolanlab/bead-normalization and gated in Cytobank, but the raw FCS data is 'available on reasonable request' (not deposited) -> NOT reproducible. The Fig 1A survival analysis needs COG P9906 outcome/RFS data which GEO states cannot be provided (restricted). EGAS00001003266 in Fig 4B is EGA controlled-access. The reproducible public-data result is Fig 4B/4C: IL-15 transcript inversely correlates with MYC in B-ALL. Reproduced on «our HPC»/«infra» from the cited PUBLIC GEO series matrices using canonical HG-U133 probes (MYC 202431_s_at[+244089_at], IL15 205992_s_at): (C1a) GSE11877 (COG P9906, n=207) Pearson r=-0.175 p=0.012 / Spearman -0.184 p=0.008 -> significant NEGATIVE, matches paper. (C1b) GSE13159 (MILE, n=576 B-ALL) r=-0.111 p=0.0075 / Spearman -0.129 p=0.0019 -> significant NEGATIVE. (C2) subtype contrast: MYC-driven+MLL group n=83 (EXACTLY the paper's stated 'MYC driven+MLL n=83') has lower IL-15 (median 0.195) than ETV6::RUNX1+Ph+ n=180 (0.336), MWU p=7.5e-15 -> aggregate claim reproduced; HONESTY NUANCE: at single-subtype resolution the t(8;14) MYC-translocation class alone (n=13) has the HIGHEST IL-15 (0.42), so the low signal is carried by KMT2A/MLL t(11q23) + hyperdiploid, and MILE lacks BCL2-tx/hypodiploid classes -> graded partial. (C3) GSE132929 Burkitt's IL-15 lower than non-Burkitt's, MWU p=7.5e-22. NOT attempted, and why: CyTOF normalization/gating (the repo's real use) raw FCS restricted; Fig 1A survival COG outcome data restricted; EGA arm controlled-access; all mouse/flow/qPCR/CRISPRa NK-92 wet-lab (non-pipeline). Fabrication check: no discrepancy -- every reproduced Fig 4B/4C value is derivable from the cited public GEO data and holds in the stated direction with significance; the non-reproduced parts are out of scope due to restricted data, not due to any detected inconsistency. All grades provisional; human audit sheet in AUDIT.md.

💻 Code ↗ 🗄 Data: GSE11877

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 76
    assessed: 2026-06-15 ⛓ 49d8651ad5ee
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator headless) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper hypothesizes that an intrinsic block in type I interferon (IFN-I) production during primary leukemogenesis suppresses IFN-I-driven anti-leukemia immune surveillance (notably IL-15-dependent NK-cell responses) in B-cell acute lymphoblastic leukemia (B-ALL), and that restoring this pathway via IL-15-producing NK cells is therapeutically effective.

Core claims
  • High expression of IFN-I signaling/response genes predicts favorable clinical outcome (longer relapse-free survival) in B-ALL patients finding
  • Human and mouse B-ALL microenvironments harbor an intrinsic defect in autocrine (B-cell) and paracrine (pDC) IFN-I production and IFN-I-driven immune responses finding
  • Suppression of IFN-I production most markedly lowers IL-15 transcription, reducing NK-cell number and effector maturation in B-ALL mechanism
  • Reduced IFN-I production is sufficient to suppress immunity and promote MYC-driven leukemogenesis in mice mechanism
  • MYC overexpression sensitizes B-ALL cells to NK cell-mediated killing while IL-15 suppression is most severe in MYC-high subtypes finding
  • CRISPRa-engineered IL-15-secreting human NK cells kill high-grade B-ALL in vitro and block leukemia progression in vivo more effectively than non-IL-15-producing NK cells resource
  • Adoptive transfer of healthy NK cells and administration of IFN-Is prolong survival/reduce leukemia progression in B-ALL-prone mice finding
  • Ex vivo IFN-I treatment of primary mouse B-ALL cells restores proximal IFN-I signaling and partially restores IL-15 production finding
Experimental setups
Assay System Perturbation Readout Platform
Relapse-free survival / gene expression outcome analysis COG P9906 high-risk pediatric B-ALL patient cohort (207 children) none RFS probability stratified by IFN-I signaling gene expression (IFNAR1, IFNAR2, STAT1, MX1, OAS1)
Flow cytometry (intracellular IFNα2b, pSTAT1, surface staining) Human B-ALL patient BMMCs/PBMCs vs healthy donors class C CpG ODN stimulation; IFNβ stimulation IFNα2b+ cells, pSTAT1, pDC/cDC frequencies, CXCR4/HLA-DR MFI BD FACSymphony cytometer; FlowJo V.10.7.1
Mass cytometry (CyTOF) Human and mouse B-ALL samples PMA/ionomycin stimulation vs unstimulated surface proteins and intracellular cytokines across immune subsets Standard BioTools Cell-ID Intercalator-Ir; Cytobank
Transgenic mouse leukemogenesis / survival Eμ-Myc and Eμ-Myc/IFNAR1-/- mice germline IFNAR1 knockout leukemia development, leukemia-free survival, immune cell frequencies
Adoptive NK-cell transfer Eμ-Myc B-ALL-bearing mice IV injection of 7×10^5 syngeneic UBC-GFP healthy NK cells vs PBS leukemia-free survival
In vivo IFNβ administration Eμ-MYC B-ALL-prone mice (~7–20 weeks) IP 50,000 IU IFNβ vs PBS for 9 days leukemia progression and circulating NK/NK-effector frequencies flow cytometry
NK cytotoxicity and proliferation assay CRISPRa control-sgRNA vs IL-15-sgRNA NK-92 cells co-cultured with B-ALL targets CRISPRa IL-15 overexpression; IL-2; PMA/ionomycin specific cytotoxicity (7-AAD), live cell counts, IL-15 secretion BD FACSymphony; FlowJo; PBL high-sensitivity human IL-15 ELISA (#41702)
Cell line-derived xenograft / bioluminescence imaging Luciferase-labeled P493-6 B-ALL in NSG (NOD-SCID-IL2Rγ-/-) mice IV CRISPRa control vs IL-15-producing NK-92 cells (7×10^6/mouse) disease progression by BLI Lago-X spectral instruments imaging
Key results
  • IFN-I response high B-ALL patients (n=15) had significantly longer RFS than IFN-I response low patients (n=14)
  • IFN-I response high patients were ~3 times less likely to have WBC ≥100,000/μL (13% vs 43%) ~3-fold (13% vs 43%)
  • IFN-I response low patients had reduced IRF7 and CD123 transcript expression
  • B-ALL patients show reduced IFNα2b+ cells in CpG-stimulated BMMCs/PBMCs and HLA-DR+ non-B/non-T/non-NK fraction vs healthy donors
  • Suppression of IFN-I production most markedly lowers IL-15 transcription and reduces NK-cell number/effector maturation
  • Adoptive transfer of healthy NK cells significantly prolonged survival of ALL-bearing transgenic mice
  • CRISPRa IL-15-secreting NK cells killed B-ALL in vitro and blocked leukemia progression in vivo better than non-IL-15 NK cells
  • IL-15 suppression most severe in MYC-high B-ALL subtypes; MYC overexpression increases sensitivity to NK killing
Key statistics
  • count 207 children with high-risk B-ALL (COG P9906) (cohort for RFS analysis)
  • count High IFN-I Response n=15, Low IFN-I Response n=14 (RFS stratification groups)
  • percent 13% vs 43% with WBC ≥100,000/μL (high vs low IFN-I response patients)
  • count n=7 BMMC and n=10 PBMC per group (B-ALL vs healthy donor IFNα2b flow cytometry)
  • count n=13 high vs n=9 low (IFN-I+IRF7+CD123) (concomitant expression WBC outcome analysis)
  • pvalue p<0.05 (significant); 0.05<p<0.1 (trending) (two-tailed statistical thresholds reported)
  • count 7×10^5 NK cells injected (adoptive transfer dose into Eμ-Myc mice)
  • count 7×10^6 NK cells per mouse; effector:target 10:1 (xenograft NK dose and cytotoxicity assay ratio)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper employs high-dimensional flow and mass cytometry, murine survival analyses, qPCR, and in vitro/in vivo NK-cell efficacy assays to characterize IFN-I suppression in B-ALL and evaluate IL-15-producing CRISPRa-engineered NK cells. Pairwise group comparisons between mouse cohorts used two-tailed Mann-Whitney U tests; Kaplan-Meier curves were compared with the log-rank test; and exact p values were reported for significant (p<0.05) and trending (0.05<p<0.1) results. Sample sizes were determined a priori with the 'cpower' function in R, and qPCR biological replicates were each run in three technical replicates.

Replicationmixed Sample sizeSample size calculated a priori using 'cpower' function in R; per-figure n stated for patient and healthy donor cohorts (n=7–15 per group); cell-line experiments repeated three times; each biological qPCR sample run in three technical replicates GroupsB-ALL patients vs age-matched healthy donors; IFN-I-response-high vs IFN-I-response-low B-ALL patients; Eμ-Myc vs Eμ-Myc/IFNAR1−/− mice; NK-cell-treated vs PBS-treated mice; CRISPRa IL-15-sgRNA vs control-sgRNA NK-92 cells Pairingunpaired Randomization/blindingnot stated Dispersionunclear Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
two-tailed Mann-Whitney U pairwise comparisons between mouse cohorts (explicitly stated in Methods); statistical test used for human patient vs healthy donor flow cytometry comparisons is not named in the provided methods text mouse cohort sizes not stated in methods text; patient and healthy donor flow cytometry groups n=7–10 per group as stated per figure not stated
log-rank Kaplan-Meier relapse-free survival comparison (COG P9906 IFN-I-high n=15 vs IFN-I-low n=14); also used for leukemia-free survival in NK-cell adoptive transfer mouse experiment n=15 vs n=14 (COG P9906 RFS); mouse cohort sizes for survival not specified in provided text not stated
Approaches that could also have been used
  • IFN-I response genes were dichotomized at their median for Kaplan-Meier survival analysis, producing two groups of n=15 and n=14 from 207 total patients
    Could also: Cox proportional hazards regression using IFN-I gene expression as a continuous covariate, optionally adjusted for known prognostic variables such as WBC count, age, and cytogenetic subtype — Continuous modeling preserves all variation in gene expression rather than compressing it into two levels, and produces a hazard ratio with a confidence interval that facilitates comparison with other studies; covariate adjustment can clarify whether the survival association is independent of established risk factors
  • Multiple pairwise comparisons were made across many figures and endpoints without a stated multiplicity correction
    Could also: A false discovery rate correction (e.g., Benjamini-Hochberg) or family-wise error rate correction (e.g., Bonferroni) applied across related comparisons within each figure or experimental unit — When many tests are performed simultaneously, controlling the expected proportion of false positives helps distinguish reproducible signals from chance findings; making the correction approach explicit adds interpretive transparency
  • Survival between two patient subgroups was compared with the log-rank test alone
    Could also: A multivariable Cox proportional hazards model including IFN-I response status and established clinical prognostic covariates — Adjusted Cox regression quantifies the independent prognostic contribution of IFN-I response status and yields a hazard ratio with a confidence interval, enabling effect-size estimation and cross-study comparison beyond a binary significant/not-significant conclusion
  • Pairwise group comparisons used the Mann-Whitney U test (stated explicitly for mouse cohorts); where three or more groups are compared the analysis framework is not described
    Could also: A Kruskal-Wallis omnibus test followed by Dunn's post-hoc pairwise test (or one-way ANOVA with Tukey HSD if normality is met) when more than two groups are compared simultaneously — Testing all groups jointly before pairwise comparisons controls the experiment-wise error rate within that family; Dunn's post-hoc procedure accounts for the number of pairwise contrasts naturally
  • Dispersion around central tendency values is not specified in the statistical methods section
    Could also: Explicit reporting of SD (for approximately normal data), IQR (paired with nonparametric tests), or 95% confidence intervals alongside group medians or means — Dispersion measures are particularly informative at small n (n=7–15 per group throughout), allowing readers to assess variability and effect magnitude independently of p values; SD and IQR also signal whether parametric or nonparametric methods are more appropriate
  • Effect sizes are not reported alongside p values for any comparison
    Could also: Rank-biserial correlation (for Mann-Whitney U), Cohen's d (for t-tests), or hazard ratios with 95% CIs (for survival analyses) — Effect size estimates convey biological or clinical magnitude of differences separately from sample-size-dependent statistical significance, and are increasingly expected by journals and meta-analyses for quantitative synthesis
Software: FlowJo 10.7.1 · MATLAB · Cytobank (Beckman Coulter) · R (cpower function)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
15
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37217248

Paper. Sudan, Kandarpa, et al. Intrinsic suppression of type I interferon production underlies the therapeutic efficacy of IL-15-producing natural killer cells in B-cell acute lymphoblastic leukemia. J Immunother Cancer 2023;11:e006649. PMID 37217248 · PMCID PMC10231005 · DOI 10.1136/jitc-2022-006649. Open access.

Brief metadata sanity check (resolved). The room brief paired this paper with github.com/nolanlab/bead-normalization (a MATLAB CyTOF bead-normalization tool) and GEO:GSE11877 (a 2009 COG pediatric B-ALL microarray set whose GEO record lists other PMIDs). Both are genuinely cited by the paper, not a mismining:

  • The repo is used as a black-box: "Data were normalized using MATLAB (https://github.com/nolanlab/bead-normalization/releases)" — i.e. mass-cytometry (CyTOF) bead normalization. The raw FCS data are "available on reasonable request" (not deposited) → CyTOF normalization is not reproducible from public data → data_restricted, out of scope.
  • GSE11877 (COG P9906) and GSE13159 (MILE) are used in Figure 4B for the gene-expression analysis (MYC vs IL-15 transcript correlation across B-ALL subtypes). Their expression matrices are public → in scope.

Pipeline-derived results and their data sources

Result (paper) Pipeline Data Public? In scope
Fig 4B IL-15 transcript correlates negatively with MYC across B-ALL; IL-15 lower in MYC/BCL2-tx, KMT2A-r, hypodiploid vs Ph/Ph-like/ETV6::RUNX1 microarray expr → per-gene probe extraction → linear regression / group comparison GSE11877, GSE13159, EGAS00001003266 GSE11877 ✅, GSE13159 ✅, EGA ❌ controlled YES (on the 2 public GEO sets)
Fig 4C Burkitt's lymphoma (MYC-driven) expresses lower IL-15 than non-Burkitt's B-NHL microarray expr → group comparison GSE132929 YES (secondary)
Fig 1A COG P9906 RFS by IFN-I-signature (IFNAR1/2,STAT1,OAS1,MX1) high vs low (Kaplan-Meier, log-rank) expr stratification → survival GSE11877 expr + COG outcome/RFS data expr ✅ but RFS/outcome COG-restricted ("cannot be provided… contact COG") NO — data_restricted (outcome not public)
CyTOF immune profiling (Figs 1,2; pDC/cDC/NK frequencies) CyTOF: bead-normalize (nolanlab repo) → Cytobank gating raw FCS ❌ "on reasonable request" NO — data_restricted
Mouse Eµ-Myc experiments, qPCR, flow, CRISPRa NK-92, in-vivo survival (Figs 2,3,5,6) wet-lab NO — non_pipeline (wet-lab)

What we attempt (80/20)

Primary (C1): Reproduce the Fig 4B headline — IL-15 (IL15) transcript expression is negatively correlated with MYC transcript expression in B-ALL — directly from the public series matrices of GSE11877 and GSE13159, using the canonical HG-U133 probes (MYC 202431_s_at, IL15 205992_s_at). Pearson + Spearman; expect r<0, p<0.05.

Secondary (C2): Fig 4B subtype contrast — IL-15 lower in MYC-driven/KMT2A-r subtypes vs ETV6::RUNX1/Ph — using GSE13159 (MILE) leukemia-class sample labels.

Tertiary (C3): Fig 4C — Burkitt's lymphoma lower IL-15 than non-Burkitt's B-NHL — GSE132929 (only if cheap).

What we do NOT attempt, and why

  • CyTOF normalization/gating (the bead-normalization repo's actual use): raw FCS not public → data_restricted.
  • Fig 1A survival on COG P9906: RFS/outcome data COG-restricted → data_restricted.
  • EGAS00001003266 arm of Fig 4B: EGA controlled-access → data_restricted.
  • All mouse / flow / qPCR / CRISPRa NK-cell wet-lab: not a pipeline.

The paper's core mechanistic computational claim that is checkable on public data is the MYC↔IL-15 inverse relationship (Fig 4B). That is our reproduction target.

Figures / tables: Fig 4BFig 4C
C1a
Reported
IL-15 transcript correlates negatively with MYC in B-ALL (GSE11877/COG P9906); qualitative
Reproduced
Pearson r=-0.175 (p=0.012), Spearman rho=-0.184 (p=0.0078), n=207
within tolerance
C1b
Reported
IL-15 transcript correlates negatively with MYC in B-ALL (GSE13159/MILE); qualitative
Reproduced
Pearson r=-0.111 (p=0.0075), Spearman rho=-0.129 (p=0.0019), n=576 B-lineage ALL
within tolerance
C2
Reported
IL-15 reduced in MYC/BCL2-tx, KMT2A-rearr, hypodiploid vs Ph+/Ph-like/ETV6::RUNX1
Reproduced
GSE13159 MYC+MLL n=83 median 0.195 vs ETV6::RUNX1+Ph+ n=180 median 0.336; Mann-Whitney p=7.5e-15 (low<high); but t(8;14) MYC-tx alone n=13 is highest (0.42)
partial
C3
Reported
MYC-driven Burkitt's lymphoma expresses lower IL-15 than non-Burkitt's B-NHL (Fig 4C)
Reproduced
GSE132929 Burkitt n=59 median 3.89 vs non-Burkitt n=231 median 6.04; Mann-Whitney p=7.5e-22 (Burkitt lower)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator headless) · v1.0 L1 76/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

175.5 k
tokens (I/O) · 15.1 M incl. cache
23 min
runtime · 0.02 CPU-h
2.5 GB
peak RAM
2
HPC jobs
hummel
machine