Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

iBRIDGE: A Data Integration Method to Identify Inflamed Tumors from Single-cell RNA-Seq Data and Differentiate Cell Type-Specific Markers of Immune-Cell Infiltr

Cancer Immunol Res · 2023
L1 94/100 PQI 98
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
94/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 87% of all assessed papers rank 133 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> reproduced 1:1. Reproduced the iBRIDGE R package's own worked README vignette (authors' code, commit 6a33706) on the package-SHIPPED GSE176078 breast-cancer scRNA-seq malignant subset with TCGA-BRCA inflamed/cold signatures, end-to-end on «our HPC» SLURM (conda: Seurat+AUCell). All comparable expected outputs reproduced: matrix dim EXACT; inflamed & cold top-6 gene-feature lists EXACT (and EXACT under two independent Seurat versions 4.4.0 and 5.5.0); cell-class counts within <0.1% (8636/8637/10642 vs 8626/8624/10665); 22 per-patient inflammation scores within ~2% with the same ascending rank order; above/below-threshold patient split 13/9 IDENTICAL. Sub-percent drift in C4/C5 attributable to Seurat 5.5.0/sctransform 0.4.3 (current bioconda, solver-upgraded from intended 4.4.0 pin) vs the author's ~2022 build, not a pipeline disagreement. NO fabrication signal: every value derivable from shipped package data+code. NOT ATTEMPTED (out of scope, the hard ~20%): GSE131907 lung dataset named in the spawn brief (repo's documented runnable example is GSE176078/BRCA only; GSE131907 has no documented parameters and would need rebuilding the malignant subset from raw GEO); myeloid/fibroblast cell-type marker analyses, IFN-pathway enrichment, in-vitro cell-line application, and UMAP/cluster figures (stochastic, no pinnable numbers).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 94
    assessed: 2026-06-14 ⛓ e2605a8ef506
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The malignant cell population in scRNA-seq data—because it clusters by patient due to patient-specific somatic alterations—can be integrated with bulk TCGA RNA-seq markers of inflamed/cold tumors to identify patients with a T-cell inflamed tumor immune microenvironment from single-cell data.

Core claims
  • iBRIDGE integrates the malignant subset of scRNA-seq data with reference bulk TCGA RNA-seq data to identify patients with a T-cell inflamed vs cold TIME. method
  • iBRIDGE single-cell classifications correlate highly with matched bulk assessments (0.85 and 0.9 correlation coefficients). finding
  • Malignant cells cluster by patient whereas immune and stromal cells cluster by cell type, motivating use of malignant cells as the bridge to bulk data. mechanism
  • Type I and type II interferon pathways are dominant signals of the inflamed phenotype, especially in malignant and myeloid cells. finding
  • A TGFβ-driven mesenchymal phenotype marks the inflamed state not only in fibroblasts but also in malignant cells. finding
  • Per-patient average iBRIDGE scores combined with bulk ICR or CD3E RNAScope ground-truth enable threshold-based absolute classification via ROC/Youden index. method
  • iBRIDGE can be applied to in vitro cancer cell lines and identify lines adapted from inflamed/cold patient tumors. finding
  • iBRIDGE is released as an open-source R package with a workflow vignette. resource
Experimental setups
Assay System Perturbation Readout Platform
scRNA-seq (10X Genomics) NSCLC/LUAD patient tumors (GSE131907, Kim et al.) none raw UMI counts; malignant cell inflamed/cold classification 10X Genomics
scRNA-seq (10X Genomics) Colon adenocarcinoma patient tumors (GSE132465) none raw UMI counts; cell classification 10X Genomics
scRNA-seq (10X Genomics) Breast cancer patient tumors (GSE176078) none raw UMI counts; cell classification 10X Genomics
Bulk RNA-seq Breast cancer patients (GSE176078, matched to scRNA-seq) none raw counts; ICR GSE scores as ground-truth
scRNA-seq (10X Genomics) + CD3E RNAScope Colon adenocarcinoma patient tumors (GSE178341; 9-patient RNAScope subset) none raw UMI counts; CD3E percent positive cells as ground-truth 10X Genomics / RNAScope
Bulk RNA-seq TCGA multi-cohort tumors (10079 patients) none raw counts; ICR signature GSE, DE between inflamed/cold tertiles
Bulk RNA-seq (in vitro) Cancer cell lines, CCLE/DepMap (838 lines) none (cell line) raw counts; iBRIDGE inflamed/cold classification DepMap portal
Microarray (in vivo and in vitro) Multi-tumor patient and cell line samples (GSE85507) none RMA-normalized expression; in vivo correlates of ICR microarray
Key results
  • iBRIDGE single-cell inflamed/cold scores correlated highly with matched bulk assessments. r=0.85 and r=0.90
  • Malignant cells show the highest pair-wise inter-patient distances (Euclidean or 1-Pearson) versus immune/stromal cells, which cluster by type.
  • GO analysis of inflamed patients shows interferon pathways enriched, dominant in malignant and myeloid cells.
  • TGFβ-driven mesenchymal phenotype enriched in inflamed fibroblasts and malignant cells.
  • Within a patient cluster, malignant cells show minimal assignment of opposite (inflamed vs cold) classes, enabling unambiguous patient classification.
  • iBRIDGE applied to CCLE cell lines distinguishes lines derived from inflamed vs cold tumors.
Key statistics
  • correlation 0.85 (iBRIDGE vs matched bulk assessment correlation coefficient)
  • correlation 0.9 (iBRIDGE vs matched bulk assessment correlation coefficient)
  • count 20-gene ICR signature (ICR signature used to mark Th1 cytotoxic infiltration in bulk)
  • count top 3,000 highly variable features (variable genes extracted from malignant scRNA-seq cells)
  • count ~210K cells (NSCLC GSE131907 dataset cell count (58 patients))
  • other Padj < 0.05 and log2(FC) > 0.3 (scRNA-seq DE significance threshold (Seurat FindMarkers))
  • other Padj < 0.01 and log2(FC) > 1 (bulk RNA-seq/microarray DE significance threshold (limma))
  • other 5% CD3E positive cells; twofold and 20% difference cutoffs (thresholds for absolute (RNAScope) and relative patient classification)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

iBRIDGE is a computational integration method that overlaps highly variable malignant-cell features from scRNA-seq datasets with bulk RNA-seq differential expression markers derived from TCGA tertile comparisons. Inflamed/cold classification at the cell and patient level relies on tertile-based thresholding of AUCell gene-set enrichment (GSE) scores, with DESeq2 used for bulk DE, Seurat FindMarkers for scRNA-seq DE, and limma for microarray DE. Validation is performed via Pearson correlation against matched bulk ICR scores and ROC/Youden-index analysis for absolute classification thresholds. Results are reported with adjusted p-value and log2FC thresholds, and visualized with box-whisker plots and correlation coefficients.

Replicationbiological Sample sizeSample sizes listed in Table 1 by dataset and cohort; no a priori power analysis or sample size justification described GroupsInflamed vs. cold tumor immune microenvironment patients/cells; inflamed vs. cold in vitro cell lines Pairingunpaired Randomization/blindingnot stated DispersionIQR Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionNot named; paper reports 'P_adj' throughout (DESeq2 and limma default to Benjamini-Hochberg FDR; Seurat FindMarkers defaults to Bonferroni — none explicitly named in text)
Statistical tests used
Test Applied to n Assumptions
DESeq2 differential expression (Wald test by default) Bulk TCGA RNA-seq: top vs. bottom ICR tertile patients per cohort; CCLE cell lines inflamed vs. cold TCGA total n=10,079 across cohorts; per-cohort n for each DE run not stated not stated
Seurat FindMarkers differential expression scRNA-seq: inflamed vs. cold comparisons per cell type within each dataset Varies by dataset: 58 patients/208,506 cells (GSE131907), 23/47,285 (GSE132465), 26/100,064 (GSE176078), 62/371,223 (GSE178341) not stated
limma linear model Microarray dataset GSE85507: in vivo and in vitro ICR correlates 15 in vivo + 16 in vitro samples not stated
AUCell gene set enrichment scoring (top-10% feature ranking threshold) Per-cell inflamed and cold GSE scores for iBRIDGE classification across all scRNA-seq datasets Cell counts per dataset as listed in Table 1 na
yaGST single-sample gene set enrichment Bulk ICR GSE scores for TCGA patients and CCLE in vitro cell lines TCGA n=10,079; CCLE n=838 cell lines na
ROC curve and Youden index (ROCit R package) Absolute classification threshold identification: GSE176078 (matched scRNA-seq/bulk, n=24) and GSE178341 (matched RNAscope, 9-patient subset) n=24 (GSE176078); n=9 (GSE178341) not stated
Pearson correlation coefficient Validation of per-patient iBRIDGE scores against matched bulk ICR scores in two datasets (r=0.85 and r=0.90 reported) n=24 patients (GSE176078); second dataset n not stated in available text not stated
Approaches that could also have been used
  • Inflamed/cold patient classification in bulk data used fixed tertile cutoffs of the ICR GSE distribution, excluding the middle third of samples from DE marker derivation
    Could also: A data-driven continuous threshold (e.g., Gaussian mixture model, k-means on scores, or median split retaining all samples) could also be used to define inflamed/cold groups — Tertile-based dichotomization discards approximately one-third of observations; continuous or mixture-model approaches retain all data points and may better capture cases where the score distribution is unimodal rather than bimodal
  • scRNA-seq differential expression between inflamed and cold classes was performed with Seurat's FindMarkers, treating each cell as an independent observation
    Could also: Pseudobulk DE methods — e.g., aggregating counts per patient then applying DESeq2 or edgeR — could also be used for multi-patient scRNA-seq DE — Cell-level DE methods can underestimate variance by ignoring within-patient correlation among cells from the same individual; pseudobulk approaches account for this structure and are increasingly recommended for multi-patient scRNA-seq studies
  • AUCell with a top-10% ranking threshold was used for single-cell gene set enrichment scoring
    Could also: UCell, ssGSEA, or GSVA could also be applied for single-cell GSE scoring — Different scoring methods differ in handling of sparse/dropout counts and score scale; comparing results across methods (e.g., AUCell vs. UCell) can assess robustness of cell-level classifications
  • Pearson correlation was used to validate per-patient iBRIDGE scores against matched bulk ICR scores
    Could also: Spearman rank correlation could also be computed alongside Pearson for this validation — Pearson assumes linearity and is sensitive to outliers; Spearman is non-parametric and more robust when score distributions are skewed, which is relevant for small validation cohorts (n=24 and n=9)
  • Pairwise patient distances across cell types were visualized with box-whisker plots without a formal statistical test reported for the comparison
    Could also: A Kruskal-Wallis test or a mixed-effects model with post-hoc pairwise correction could also be applied to formally compare distance distributions across cell types — A formal test with multiple-comparison correction would quantify whether the observed differences in within-patient clustering between malignant and non-malignant compartments reach statistical significance beyond visual inspection
  • Separate DE analyses were run independently per cell type and per dataset without a unified family-wise correction across all comparisons
    Could also: A global BH FDR correction applied across all cell-type and cohort comparisons simultaneously could also be used — Running independent DE analyses per cell type inflates the total number of hypotheses tested across the experiment; a unified correction would control the experiment-wide false discovery rate more conservatively
Software: R 4.1.1 · Seurat · DESeq2 · limma · AUCell · scater · EDASeq · infercnv · SingleR · clusterProfiler · ROCit · yaGST

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
2
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

RRID:SCR_016341 RRID in Article (http://semanticscience.org/resource/SIO_001029)
also used by 2 papers:
GSE131907 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE132465 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE178341 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE85507 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
RRID:SCR_000154 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_00319313 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_005012 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_006751 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_013836 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_021140 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_021327 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

Downstream reach in the literature

98 downstream papers · 1 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: table
C1
Reported
3000 x 27915
Reproduced
3000 x 27915
exact
C2
Reported
inflamed top6: B2M,HLA-C,HLA-B,HLA-A,MUCL1,CALML5
Reproduced
B2M,HLA-C,HLA-B,HLA-A,MUCL1,CALML5
exact
C3
Reported
cold top6: NEAT1,COX6C,AZGP1,AGR2,PIP,SLC39A6
Reproduced
NEAT1,COX6C,AZGP1,AGR2,PIP,SLC39A6
exact
C4
Reported
Cold=8626,Inflamed=8624,Unassigned=10665
Reproduced
Cold=8636,Inflamed=8637,Unassigned=10642
within tolerance
C5a
Reported
CID4067 ave score 0.2844366 (below)
Reproduced
0.2903 (below)
within tolerance
C5b
Reported
CID3963 ave score 8.0389367 (above)
Reproduced
7.9213 (above)
within tolerance
C5c
Reported
13 patients above / 9 below threshold
Reproduced
13 above / 9 below (identical membership)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 94/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

This is a clean, essentially 1:1 reproduction of the iBRIDGE authors' own README vignette on the package-shipped GSE176078/TCGA-BRCA data: matrix dimension and both inflamed/cold top-6 gene lists are EXACT, and the 13-above/9-below patient split is identical. The only deviations are sub-percent class counts (≤0.08%) and ~2% per-patient scores, fully attributable to Seurat 5.5.0/sctransform 0.4.3 vs the author's ~2022 build — a technical/version cause, not authors' or methodology fault, with no fabrication signal. The main caveat is scope: the brief named GSE131907 (lung) but only the documented BRCA vignette was reproducible, so cell-type-marker, IFN-enrichment and in-vitro claims remain unverified — yet every value that was compared is derivable and matches.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

97.7 k
tokens (I/O) · 7.4 M incl. cache
24 min
runtime · 0.16 CPU-h
7.5 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine