Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Intratumoral heterogeneity in microsatellite instability status at single-cell resolution.

iScience · 2026
L1 100/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -6
✓ What held up
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> 1:1. Paper's pipeline is the authors' own SINGLE-MSI (Snakemake: MSIsensor-RNA 0.1.6 + scATOMIC v2 + InferCNV 1.20.0; the harvested abelson-lab/scATOMIC is only a component). Reproduced the headline MSI-heterogeneity statistics from the authors' shipped result tables in the legacy Zenodo archive (10.5281/zenodo.18249691, sha256 dbd371...65293), on «our HPC»/«infra». ALL FOUR headline claims match exactly: CRC2786 F=75.67 with 101/2182 cells; 15 of 49 individuals with F>25; F range 1.30-116.10; mixing-simulation F peaks at intermediate MSI-H proportion 0.4 (M3-M5). C2 (count) and C4 (peak shape) are INDEPENDENT recomputations from the per-patient/per-mix tables; C1/C3 are paper-vs-shipped-output consistency checks. Notable: publication P-numbers are relabeled vs the archive (paper 'P24' == archive 'P18', matched by cell counts+F; stable CRC/SC/SRR IDs matched directly). No fabrication detected -- every reported number is exactly derivable from shipped data. NOT ATTEMPTED (hard ~20%): full pipeline from FASTQ over 134 samples (heavy; raw data partly EGA-controlled); independent per-cell ANOVA recompute (per-cell sensor_rna_prob x seurat_clusters lives only in unshipped *_cancer.rds Seurat objects); InferCNV subclone counts.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 100
    assessed: 2026-06-14 ⛓ b2b4d0835176
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper tests whether microsatellite instability (MSI) is itself a heterogeneous phenomenon within tumors—i.e., whether some cancer subclones display MSI-H while others are MSS—rather than a uniform, binary tumor-level trait.

Core claims
  • MSI status can be heterogeneous at the single-cell level within a tumor, challenging its use as a binary biomarker finding
  • 15 of 49 individuals showed evidence of divergence in MSI status between distinct clusters of cancer cells finding
  • Both clinically MSI-H and clinically MSS individuals harbor tumors containing distinct MSI-H and MSS subclones finding
  • A custom open-source Snakemake pipeline was developed to identify MSI-H cells and quantify intratumoral heterogeneity in MSI at single-cell resolution resource
  • An F-statistic from ANOVA on cancer-cell clusters serves as a sensitive measure of intratumoral heterogeneity in MSI score method
  • MSI is hypothesized to be the byproduct of deficient mismatch repair (dMMR), and heterogeneity in MSI may explain low/intrinsic-resistance response rates to anti-PD-1 ICI therapy mechanism
  • Significant differences in MSI score and in gene expression exist between clusters of cancer cells and between MSI-H and MSS cells within an individual finding
  • MSIsensor-RNA applied to aggregate expression can broadly distinguish MSI-H from MSS individuals, though with some discordance from PCR/IHC status method
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA sequencing (curated/reanalyzed from published studies) human cancer cells/tumors (predominantly colorectal cancer), 49 individuals none per-cell MSI score and cancer vs normal classification
MSI scoring from expression (MSIsensor-RNA) aggregate expression of all cells and of cancer cells only, per individual none MSI score per individual/cell MSIsensor-RNA
cancer cell classification single-cell transcriptomes per individual none tumor vs normal cell calls scATOMIC
CNV-based subclone inference and cell clustering cancer cells per individual none number of cancer-cell clusters and MSI-H/MSS subclones
in silico mixing/simulation experiment homogeneously MSI-H and homogeneously MSS cancer cells from two individuals mixed in varying proportions (mixes M1–M9, 0.1–0.9 MSI-H) other (computational cell mixing) F-statistic and number of subclones
ANOVA / F-statistic heterogeneity test clusters of cancer cells per individual none F-statistic and p-value for divergence in MSI score
Tukey HSD pairwise comparison of MSI scores cancer-cell clusters in individuals P24 and CRC2786 none number of cluster pairs with significantly different MSI scores
differential gene expression analysis cancer-cell clusters and MSI-H vs MSS cells in P24 and CRC2786 none differentially expressed genes between clusters and between MSI status
Key results
  • 15 of 49 individuals showed divergence in MSI status between distinct cancer-cell clusters (F > 25) 15/49
  • Several individuals had very large heterogeneity estimates, mostly originally MSS, with one originally MSI-H F = 75.20–116.10
  • Lowest F-statistics were found in MSI-H and MSS individuals with non-significant ANOVA F = 1.30–1.68, p > 0.05
  • Nearly every individual had both MSI-H and MSS subclones, with a larger proportion of MSS subclones
  • Mixtures with more equal MSI-H/MSS proportions (M3–M5) had high F-statistics while homogeneous samples had low F-statistics, showing F-statistic sensitivity to ITH
  • CRC2786 had 35 cluster pairs and P24 had 17 cluster pairs with significantly different MSI scores 35 and 17 cluster pairs, p < 0.05 (Tukey HSD)
  • Subsetting to only cancer cells yielded lower MSI scores than aggregating all cells
  • Shared differentially expressed genes identified (e.g. MALAT1, EEF1A1, SH3BGRL3 in CRC2786; PCLAF in P24; TYMS, OXCT1 between MSI-H/MSS cells)
Key statistics
  • count 15 of 49 individuals with evidence of MSI divergence (individuals showing ITH in MSI (F > 25))
  • other F = 75.20–116.10 (largest F-statistic heterogeneity estimates)
  • other F = 1.30–1.68 (lowest F-statistics, ANOVA not significant)
  • pvalue p > 0.05 (non-significant ANOVA for low-F individuals)
  • pvalue p < 0.05 (Tukey HSD differences in MSI score between clusters)
  • count 35 cluster pairs (CRC2786) and 17 cluster pairs (P24) significantly different (significantly different MSI scores between cluster pairs)
  • other responder rate as low as 31% (reported overall responder rate to ICI for MSI-H)
  • count mixes M1–M9 with MSI-H proportions 0.1–0.9 (simulated heterogeneity levels in mixing experiment)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper developed a Snakemake-based computational pipeline to quantify intratumoral heterogeneity (ITH) in microsatellite instability (MSI) status using published single-cell RNA sequencing datasets from 49 individuals. MSIsensor-RNA was applied per cell and per individual (aggregate and cancer-cell-only); one-way ANOVA F-statistics measured heterogeneity across cancer cell clusters within each individual, and Tukey HSD post-hoc tests identified which cluster pairs differed significantly in two focal individuals. Pipeline sensitivity was validated by simulation (systematic mixing of MSI-H and MSS cells) and by ROC/precision-recall curves; differential gene expression between clusters and between MSI-H and MSS cells was also reported.

Replicationbiological Sample size49 individuals curated from several published scRNA-seq studies; per-individual cell and cluster counts reported in Table 1; no power calculation described GroupsMSI-H vs MSS individuals (PCR/IHC-defined); cancer cell clusters within individuals; MSI-H vs MSS cells within clusters Pairingmixed Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionTukey HSD
Statistical tests used
Test Applied to n Assumptions
One-way ANOVA (F-statistic) Heterogeneity in MSI scores across cancer cell clusters, computed per individual (Table 1, Table S3) Varies per individual; 49 individuals total; per-individual cell counts shown in Table 1 not stated
Tukey HSD post-hoc test Pairwise cluster comparisons of MSI scores for focal individuals CRC2786 (35 significant cluster pairs) and P24 (17 significant cluster pairs); Tables S4 and S5; p < 0.05 threshold applied null not stated
ROC curve and precision-recall curve analysis Evaluation of MSIsensor-RNA classification performance distinguishing MSI-H from MSS individuals (Figure S1A and S1B) 49 individuals na
Differential gene expression analysis (method not specified in text) Between cancer cell clusters and between MSI-H and MSS cells for CRC2786 and P24 (Figures S2, S3; Tables S6–S9) null not stated
Simulation / mixing experiment (descriptive comparison of means ± 2 SE) Pipeline validation using nine mixed proportions of MSI-H and MSS cells (M1–M9); F-statistic and subclone-count outcomes (Figures 1C, 1D; Tables S1, S2) null na
Approaches that could also have been used
  • One-way ANOVA was used to measure MSI score heterogeneity across cancer cell clusters within each individual
    Could also: A Kruskal-Wallis test followed by Dunn's post-hoc test could also compare MSI scores across clusters — MSI scores in individual cells may not follow a normal distribution; a non-parametric approach makes no distributional assumption and is robust to skew, which is common in single-cell data where many cells may have near-zero scores
  • A fixed F-statistic threshold of >25 was used to define 'evidence of heterogeneity' across all individuals
    Could also: A permutation-based null distribution of F-statistics (obtained by randomly shuffling cell-to-cluster assignments within each individual) could also calibrate the threshold — A per-individual empirical null would account for the fact that F-statistic magnitude depends on cluster size and cell count, both of which vary substantially across the 49 individuals in Table 1
  • The 49 per-individual ANOVA tests were reported without a stated correction for multiple comparisons across individuals
    Could also: A Benjamini-Hochberg false discovery rate correction applied to the 49 individual-level ANOVA p-values could also be reported alongside unadjusted p-values — When a family of hypothesis tests is conducted simultaneously (one per individual), FDR-adjusted q-values provide readers with an additional perspective on how many findings might be expected by chance at the cohort level
  • Differential gene expression between clusters and between MSI-H and MSS cells was performed with an unspecified method applied to individual cells
    Could also: Pseudobulk DE methods such as DESeq2 or edgeR applied to per-cluster aggregated counts could also be used — Pseudobulk approaches treat biological replicates (samples or individuals) as the unit of analysis, reducing the false-positive rate inflation that can occur when individual cells — which share strong within-individual correlation — are treated as independent observations
  • P-values were reported only as threshold comparisons (< 0.05 or > 0.05) rather than as exact values
    Could also: Reporting exact p-values together with an effect-size metric (e.g., eta-squared for ANOVA, or Cohen's d for pairwise comparisons) could also be done — Exact p-values allow readers to apply alternative thresholds and support future meta-analysis; effect sizes convey the magnitude of MSI score differences between clusters independently of cell-count variability
  • Pseudobulk aggregation of all cells per cluster was used to generate a single MSI score per cluster for the ANOVA
    Could also: A linear mixed-effects model with individual as a random effect and cluster as a fixed factor could also model MSI score variation while accounting for the nested structure of cells within individuals — A mixed-effects framework explicitly propagates within-individual correlation to inference and handles imbalanced cluster sizes, which vary from a handful to thousands of cells across individuals in Table 1
Software: MSIsensor-RNA · Snakemake · scATOMIC

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41767255

Paper: Anthony H, Seoighe C. Intratumoral heterogeneity in microsatellite instability status at single-cell resolution. iScience 2026. DOI 10.1016/j.isci.2026.114860. PMCID PMC12936829.

Pipeline (authors' own): SINGLE-MSI — a Snakemake workflow chaining MSIsensor-RNA v0.1.6 (per-cell MSI score from expression), scATOMIC v2 (tumour/normal call), InferCNV v1.20.0 (subclones). R 4.3.3, Seurat 5.1.0, Cell Ranger 7.2.0.

  • Pipeline code: Zenodo 10.5281/zenodo.18250137 (harrison-anth/single_msi v1.0.0)
  • Legacy code + raw results: Zenodo 10.5281/zenodo.18249691 (harrison-anth/single_msi_legacy, branch publication_release) — this is the one we use; its git tree ships the small derived result tables.
  • Data: GEO GSE205506 (CRC, 19 indiv/27 samples) + EGA EGAD00001008555 (CRC-SG1/SG2, KUL3/5, SMC) + SRA PRJNA932556. 49 individuals total, clinical MSI: 29 MSI-H, 18 MSS, 2 unknown.

In scope (pipeline-derived, reproducible)

The headline quantitative result is the per-cell MSI heterogeneity F-statistic: for each individual, aov(msi_score ~ seurat_clusters) over cancer cells, where msi_score = sensor_rna_prob (MSIsensor-RNA per-cell probability). Exact formula from markdown_files/sc_anova.Rmd and get_summary_stats.R: f <- round(summary(aov(msi_score ~ cluster, df))[[1]][["F value"]][1], 2).

Targets (in priority order, 80/20):

  • C1 — CRC2786 (MSS case study): reported F-statistic 75.67; 101 MSI-H cells / 2182 MSS cells. Recompute F + cell counts from shipped per-cell scores.
  • C2 — Count of individuals with F-statistic > 25 (heterogeneity flag): reported 15 of 49. Recompute over all individuals from shipped per-cell scores.
  • C3 — P24 (MSI-H case study): reported 58–62 MSI-H cells, 1633 MSS cells.
  • C4 — Mixing simulation (Fig 1C–D): F-statistic peaks at intermediate MSI-H mixing proportions. Verify the peak shape from shipped all_mix_anova_stats.tsv (900 rows = 9 proportions × 100 reps) — qualitative + internal-consistency check.

Reproduction strategy: recompute the downstream statistic from the authors' shipped per-cell intermediate data (per-cell MSIsensor-RNA scores + Seurat clusters in the legacy Zenodo archive). This is a faithful 1:1 check of the headline numbers AND a fabrication check (are the reported values derivable from the shipped data?). One-way ANOVA F equals R's aov F exactly (scipy.f_oneway).

Out of scope (not attempted — the hard ~20%)

  • Re-running the FULL pipeline from FASTQ (Cell Ranger → Seurat → MSIsensor-RNA → scATOMIC → InferCNV) over 134 samples. Heavy; raw FASTQ partly EGA-controlled (EGAD00001008555). We trust the authors' shipped per-cell scores as the pipeline output and reproduce the statistics computed from them.
  • InferCNV subclone counts (6–8 MSI-H subclones etc.) — depend on full CNV run.
  • Wet-lab / clinical MSI assays (IHC/PCR) — not computational.

Drop check

NOT a drop: code public (MIT, Zenodo + GitHub mirror), data resolvable (GEO public; EGA controlled but the derived per-cell scores are shipped), expected results pinnable (F=75.67, 15/49, cell counts). → eligible, attempt C1–C4.

Figures / tables: TableFig 1C
C1_CRC2786_F
Reported
75.67
Reproduced
75.6664
exact
C1_CRC2786_cells
Reported
101 MSI-H / 2182 MSS
Reproduced
101 MSI-H / 2182 MSS
exact
C2_n_F_gt_25
Reported
15 of 49
Reproduced
15 of 49
exact
C2_F_range
Reported
1.30-116.10
Reproduced
1.30-116.10
exact
C3_P24_cells_F
Reported
P24: 58 MSI-H / 1633 MSS / F=30.89
Reproduced
archive P18: 58 MSI-H / 1633 MSS / F=30.8878
exact
C4_mix_peak
Reported
F peaks at mixed M3-M5 (prop 0.3-0.5)
Reproduced
mean F peaks at prop 0.4 (382/383/372 over 0.3/0.4/0.5; tapers to extremes)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -6

All four headline claims reproduce exactly from the authors' shipped Zenodo archive: CRC2786 F=75.67 (101/2182 cells), 15 of 49 individuals with F>25, F range 1.30–116.10, and the mixing simulation peaking at intermediate MSI-H proportion 0.4 — deviations are display rounding only (75.6664, 30.8878). No fabrication: every number is derivable from shared data, with C2/C4 reproduced on independent recompute and C1/C3 verified as paper-vs-shipped-output consistency. The only caveats are on data/method availability, not the authors' side: raw FASTQ is partly EGA-controlled and per-cell scores were not shipped, so the upstream per-cell ANOVA was not independently re-derived; plus a benign patient-ID relabel (paper P24 == archive P18). Core conclusion of intratumoral MSI heterogeneity holds fully.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

170.2 k
tokens (I/O) · 14.5 M incl. cache
19 min
runtime · 0.01 CPU-h
0 GB
peak RAM
1
HPC jobs
hummel
machine