Blood and tissue correlates of steroid non-response in checkpoint inhibition-induced immune-related adverse events.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH for the in-scope public part; outcome = PARTIAL (one clean 1:1 data point + a documented public-data gap). The paper has two data classes: (1) its OWN UNICIT-cohort data (spectral flow cytometry, IMC, bulk RNA-seq; Figs 2-4) which is 'upon reasonable request' with EU restrictions = out of scope (data_restricted); (2) a re-analysis of THREE PUBLIC scRNA-seq cohorts (Fig 5). We reproduced the cleanly-public, fully-specified Thomas2024 component on «our HPC» («job»): downloaded the exact public GSE206299 cd4 h5ad to «infra», ran the deposited script's patient selection + QC, and reconciled the paper's reported N=37 re-analyzed patients EXACTLY (12 Thomas recomputed from public data + 17 Gupta hardcoded + 8 Luoma = 37; a naive read of the 13 hardcoded Thomas IDs would give 38, but the script's own remove-list drops SIC_32, yielding 12). All 13 selected SIC IDs are real entries in the public deposit -> no fabrication signal. We did NOT attempt the headline integrated finding (Th17/Th22 abundance vs response): it is BLOCKED because the Gupta2024 input the script consumes (CD3_scRNAseq_data.RDS, an annotated/clustered Seurat object) is ABSENT from the cited GSE189040 (which ships only raw Pool*.zip), and the result is additionally stochastic and under-specified ('clusters not individually annotated') = the hard 20%, deliberately not chased and not faked. QC counts are pure matrix arithmetic, so they are tool-independent (no need for the authors' fragile Seurat-v5 + zellkonverter stack). Honest 1:1 on the one clear, public, pinnable claim; honest documentation of what cannot be reproduced from the cited public sources.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 71assessed: 2026-06-15 ⛓ 20cdb625517b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan response to first-line steroids in immune checkpoint inhibitor-induced immune-related adverse events (especially gastro-intestinal irAEs) be predicted, and what blood- and tissue-based immune mechanisms drive steroid non-response?
- ★ An enhanced type 1/type 17 (Th1/Th17, TC1/TC17) immune response in blood and tissue is associated with steroid non-response in irAEs finding
- ★ Steroid non-responders show elevated TC1/TC17 CD8+ T cells and Th1/Th17-associated interleukins in blood before steroid initiation and persistent (CD8+) T cell activation after steroid initiation finding
- ★ Cross-sectional colitis tissue analysis suggests lower lymphocyte infiltration within 24h in steroid responders, indicating rapid steroid effects on irAE-affected tissue finding
- ★ Non-responders' colitis tissue is enriched with activated CD4+ memory T cells and a pronounced type 1/17 immune response finding
- ★ Peripheral T cell PD-1 receptor occupancy is unrelated to steroid response finding
- A multi-omics approach (flow cytometry, serum multiplex, bulk RNA-seq, histology) can identify blood- and tissue-based correlates of steroid response in irAEs method
- Both high peak-dose steroids and second-line immunosuppression are independently associated with worse overall and cancer-specific survival in melanoma patients treated for irAEs finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Spectral flow cytometry (immune cell subsets, cytokines, transcription factors) | PBMCs from irAE patients (13 steroid responders, 11 non-responders) and 6 healthy donors | systemic steroids (paired pre/post); ex vivo PMA/ionomycin restimulation | abundance and functional profiles of immune cell subsets | Cytek Aurora 5L spectral flow cytometer |
| PD-1 receptor occupancy assay | PBMCs / peripheral T cells from irAE patients | anti-PD-1 therapy (nivolumab/pembrolizumab); none ex vivo | T cell fraction completely/partially bound by PD-1-blocking antibodies; CXCL13 | BD LSR Fortessa |
| ICI-bound PD-1 internalization assay (imaging cytometry) | Pre-steroids PBMCs from 4 patients (2 highest, 2 lowest CD4+ T cell PD-1 RO) | none (ex vivo time course 0, 30, 60 min) | internalization feature (intracellular-to-total PE/IgG4 intensity ratio) | ImageStream X Mark II; IDEAS software |
| Serum multiplex immunoassay (26 analytes) | Extended serum cohort (14 steroid responders, 17 non-responders) | systemic steroids (pre-steroids and paired samples) | serum concentrations of cytokines/chemokines/MMPs (e.g. IL-17, IFN-γ, CXCL9/10) | Luminex xMAP / Biorad FlexMAP3D; in-house multiplex immunoassay |
| Bulk RNA-sequencing with CIBERSORT deconvolution | Snap-frozen colon biopsies from UNICIT patients with histologically confirmed ICI colitis | ICI colitis requiring immunosuppression; none ex vivo | gene expression (TPM/counts) and estimated abundance of 22 immune cell subsets (LM22) and γδ T cells (LM7) | Illumina NovaSeq 6000; NEBNext Ultra II Directional RNA Library Prep Kit |
| Histology (H&E digital image analysis) | Colon tissue slides from ICI colitis patients (RNA-seq cohort plus retrospective registry) | steroid treatment timing (cross-sectional, within/after 24h) | lymphocyte infiltration in colitis tissue | — |
- ▲ Elevated TC1/TC17 CD8+ T cells and Th1/Th17-associated interleukins in blood before steroid initiation in non-responders
- ▲ Persistent (CD8+) T cell activation after steroid initiation in blood of steroid non-responders
- ▼ Lower lymphocyte infiltration within 24h in colitis tissue of steroid responders
- ▲ Non-responders' colitis tissue enriched with activated CD4+ memory T cells and pronounced type 1/17 immune response
- – Peripheral T cell PD-1 receptor occupancy unrelated to steroid response
- count 22 million cells remained for analysis after pre-gating and quality control (spectral flow cytometry dataset)
- count 26 analytes measured (serum multiplex immunoassay panel)
- count 22 immune cell subsets deconvoluted (LM22 signature) (CIBERSORT bulk RNA-seq deconvolution)
- pvalue Pperm < 0.05 (CIBERSORT deconvolution inclusion threshold)
- count irAEs occur in up to 60% of patients on combined anti-PD-1 plus anti-CTLA-4 (introduction background)
- count Flow cytometry cohort n=24 patients (13 responders, 11 non-responders) plus 6 healthy donors (flow cytometry PBMC cohort)
- count Extended serum cohort n=31 (14 responders, 17 non-responders) (serum cohort)
- count Twenty-million paired-end (150 bp) reads per sample (RNA sequencing depth)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This prospective observational multi-omics study enrolled ICI-treated patients who developed irAEs, comparing steroid responders and non-responders across three data types: spectral flow cytometry of PBMCs collected longitudinally (paired pre- and post-steroid samples; N=24 patients), serum multiplex immunoassay for 26 analytes (N=31), and bulk RNA-sequencing of snap-frozen colon biopsies in a cross-sectional tissue design. Flow cytometry data were processed with archsinh transformation, peacoQC quality control, landmark-based batch normalization, and unsupervised clustering via FlowSOM plus Seurat UMAP; tissue immune composition was estimated by CIBERSORT deconvolution (absolute mode, LM22 matrix, P_perm < 0.05 inclusion threshold). The provided text was truncated mid-methods before the inferential statistical tests, multiplicity correction strategy, and full result-reporting details were described, so those fields are marked null below.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| CIBERSORT permutation test (P_perm < 0.05 used as sample-inclusion threshold for deconvolution) | Bulk RNA-seq immune cell deconvolution of colitis tissue biopsies | — | not stated |
-
Tissue immune cell composition was deconvoluted with CIBERSORT in absolute mode using the LM22 reference signature matrix derived from peripheral blood immune cell types↳ Could also: Single-cell RNA-seq-informed deconvolution methods such as MuSiC, DWLS, or BayesPrism using gut-specific or ICI-colitis-matched reference atlases could also be applied — A gut-tissue-specific or disease-matched single-cell reference may better represent the cellular heterogeneity of inflamed colonic mucosa (e.g., resident macrophages, ILC subsets, epithelial cells) than a blood-derived pan-tissue matrix, potentially improving deconvolution accuracy in this tissue compartment
-
Batch effects in spectral flow cytometry were corrected with a custom landmark-based normalization method analogous to fdaNorm, using the downslope inflection point as landmarks↳ Could also: Established cytometry-specific batch correction methods such as CytoNorm or batchelor (fastMNN) could also be used — Published and benchmarked methods provide external reference points for assessing correction adequacy and enable direct comparison with other studies; they also produce correction diagnostics that are familiar to reviewers
-
Unsupervised clustering was performed with FlowSOM (nClus=50) followed by UMAP visualization in Seurat on a randomly drawn 1% cell subset (~220,000 cells)↳ Could also: Clustering could also be applied to the full dataset, or marker-guided semi-supervised gating could leverage the well-characterized PBMC marker panel used here — Subsampling for UMAP may undersample rare immune populations; full-data clustering or supervised gating on known PBMC markers (CD3, CD4, CD8, CD14, CD56, FoxP3) would ensure rare subsets are represented and assignments are directly interpretable from biology
-
Tissue analysis used a cross-sectional design, comparing biopsies from different patients at varying times relative to steroid initiation (e.g., within vs. beyond 24 h)↳ Could also: Serial paired biopsies from the same patients before and shortly after steroid initiation could also provide within-subject longitudinal tissue data — A paired within-subject design eliminates inter-patient variability in baseline tissue composition, increases statistical power for detecting steroid-induced changes, and directly mirrors the paired blood design already used; the tradeoff is additional invasive procedures for patients
-
The variability of the ImageStream internalization feature (n=4 samples, 3 time points each) was approximated as 1.25 × MAD as a surrogate for standard deviation↳ Could also: Bootstrap confidence intervals or permutation-based uncertainty estimates could also characterize spread for this feature at very small n — The 1.25 × MAD approximation assumes approximate normality; with n=4, distributional assumptions are unverifiable, and resampling-based intervals make no such assumption while remaining interpretable as uncertainty bounds
-
The study enrolls small cohorts (n=13–17 per group) without a stated a priori power calculation in the available text, and the abstract describes findings as 'clear trends'↳ Could also: A formal power or sample size analysis anchored to a prespecified primary endpoint and effect size estimate could also accompany a study of this design — Reporting a power analysis contextualizes the study's ability to detect clinically meaningful differences and aids interpretation of non-significant or trend-level findings; for exploratory multi-omics work, an estimation of detectable effect sizes at the achieved n serves a similar purpose
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
13 downstream papers · 3 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Molecular Pathways of Colon Inflammation Induced by... 2020 · 385 cites
- Single-cell transcriptomic analyses reveal distinct... 2024 · 46 cites
- Improvement of PD-1 Blockade Efficacy and Eliminatio... 2021 · 22 cites
- Immune checkpoint inhibitor-induced colitis is media... 2023 · 21 cites
- Single-cell profiling reveals unique features of dia... 2023 · 19 cites
- Highly multiplexed spatial analysis identifies tissu... 2023 · 10 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41254329 (UNICIT: steroid non-response in ICI-induced irAEs)
- Paper: van Eijs et al., Commun Med 2025. DOI 10.1038/s43856-025-01164-3. PMC12627558.
- Code: https://github.com/mickvaneijs/UNICIT (commit
19bf3bb00e8a6419de9674ea51f923c8fd467331, main, pushed 2025-04-02; no license). - Data availability (verbatim): Source data in Supplementary Data 1; data underlying Figs 2a, 4f/g/i, Supp Figs 2a/4 at Zenodo 10.5281/zenodo.16948925; "additional data on the cohorts… upon reasonable request" (the UNICIT own cohort — flow/IMC/bulk).
The paper has TWO data classes
A. Authors' own UNICIT-cohort data → OUT OF SCOPE (not public)
Generated at UMC Utrecht; "upon reasonable request", EU data-sharing restrictions.
- Spectral flow cytometry (Aurora) of PBMC/tissue — Figs 2,3; scripts
Aurora_script_QC_preprocessing_v4.R,Aurora_script_umap-FlowSOM_v4.R,final_script_pred_response_blood.R,final_script_pred_response_tissue.R. Out of scope: raw flow data not deposited →data_restricted. - Imaging mass cytometry (IMC) —
IMC_analysis_pipeline.Rmd. Raw images not deposited →data_restricted. - Bulk RNA-seq of 24 colon biopsies + CIBERSORT/GSVA (Fig 4a-e) — own cohort, on request →
data_restricted. - Zenodo 16948925 holds only the small source-data tables behind a few panels, not raw inputs.
B. Re-analysis of THREE PUBLIC scRNA-seq cohorts → IN SCOPE (Fig 5)
Script final_script_scRNAseq_tissue.R. Pipeline: per-dataset patient selection →
QC (200<nFeature<3000, %MT<10, silence TCR genes) → per-patient SCTransform(glmGamPoi)
→ merge → common-feature subset → RPCA IntegrateLayers (Seurat v5) → cluster →
FindAllMarkers → per-patient cell fractions → Th17/Th22 identified by IL17A/IL22 →
GSVA on AverageExpression pseudobulk → DESeq2 (adjust for GSE). Headline result
(Fig 5): "Tissue abundance of IL17A- and IL22-expressing CD4+ clusters was NOT
associated with steroid response" (a negative finding); 37 patients re-analyzed.
Public inputs (all GEO):
| cohort | accession | what GEO actually ships | script reads | reproducible? |
|---|---|---|---|---|
| Thomas 2024 | GSE206301 / sub GSE206299 | processed GSE206299_ircolitis-tissue-{cd4,cd8,b,myeloid}.h5ad.gz |
GSE206299_ircolitis-tissue-cd4.h5ad |
YES — exact file public |
| Luoma 2020 | GSE144469 | GSE144469_RAW.tar (per-sample 10x mtx) |
C1..C8/{matrix,features,barcodes} |
partial — raw public, but NR/R + IFX sample labels are hardcoded clinical metadata |
| Gupta 2024 | GSE189040 | only raw Pool*.zip (10x) |
CD3_scRNAseq_data.RDS (clustered+annotated Seurat w/ Donor+celltype) |
NO — annotated object NOT in the cited GEO deposit |
Reproducibility verdict on the in-scope part
- The integrated tri-cohort Fig 5 result cannot be faithfully reproduced from the
cited public deposits, because the Gupta2024 input the script consumes
(
CD3_scRNAseq_data.RDS, with per-cellDonorand annotated T-helper clusters incl. "Th17 PD1+/-") is absent from GSE189040, which ships only raw pool matrices. Cluster annotations / donor mapping were held locally by the authors. - The Thomas2024 (GSE206299) component IS cleanly reproducible: the exact h5ad is public and the script's selection + QC is fully specified. This is the chosen 1:1 target. QC cell/patient counts are pure arithmetic on the count matrix (thresholds on per-cell detected-genes and %MT) → tool-independent (identical whether computed in Seurat or anndata), so they are computed faithfully without needing the fragile zellkonverter/Seurat-v5 stack.
What we attempt (80/20)
- PRIMARY (clean 1:1, public data): download
GSE206299_ircolitis-tissue-cd4.h5adto «infra»; reproduce the script's Thomas2024 patient selection (13 hardcoded SIC IDs, minus 5 excluded samples) + QC; report deterministic #patients and #cells (pre/post-QC). Verify the hardcoded patients exist in the public data and
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean, honest partial reproduction. The one cleanly-public, fully-specified claim — N=37 re-analyzed ICI-colitis patients — reconciles exactly (12 Thomas recomputed 1:1 from public GSE206299 + 17 Gupta + 8 Luoma; SIC_32 correctly dropped), and all 13 selected SIC IDs are real entries in the public deposit, so there is no fabrication signal. The headline immunological finding (Th17/Th22 abundance vs steroid response, Fig 5) was not attempted because the Gupta2024 annotated input is absent from the cited GSE189040 and the integration is stochastic/under-specified — a data-availability gap on the authors'/deposit side, not an observed discrepancy. Net: solid where reproducible, with the central conclusion left limited/untested rather than disconfirmed.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.