Smart spatial omics (S2-omics) optimizes region of interest selection to capture molecular heterogeneity in diverse tissues.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓Any deviation was negligible
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH: yes, excellently. S2-omics (PMID 41298871) is the authors' own histology->ROI-selection pipeline (UNI ViT-L/16 features -> PCA+KMeans morphology segmentation 20->merge 15 -> rectangle ROI scoring), fully deterministic (setup_seed(42) everywhere + pinned requirements.txt) with a rendered Tutorial-1 reference notebook. RESULT = 1:1 EXACT (reproduced). I ran run_roi_selection_single.py on «our HPC» (H100, «job») at repo commit a25634c with the documented Tutorial-1 params (--roi_size 6.5 6.5 --num_roi 1 --prior_preference 1, down_samp_step 10, k=20->15) on the authors' shipped demo image + UNI checkpoint (Google Drive). All 8 pipeline-derived outputs reproduced BIT-EXACTLY: 49490 patches; 20->15 clusters; best ROI [[76,8],[146,49],[105,119],[35,78]]; ROI/scale/coverage/balance scores 0.7923578790248927 / 0.6245726221007392 / 0.9455532908082276 / 0.8423550596065301 (identical to 16 sig figs). Minimal documented env deviations (hardware/headless only): torch 2.0.1 from cu118 (H100 sm_90) instead of the cu117 pin, and opencv-python-headless (no libGL on node); everything else exactly per requirements.txt. NOT ATTEMPTED (out of scope, see scope.md): wet-lab spatial-omics generation; the multi-tissue biological heterogeneity-capture main-figure claims (need the non-shipped full cohort + expert annotations + downstream spatial-omics readouts); the stochastic per-run label-broadcasting accuracy (soft target); Tutorials 2-4 (same pipeline, additional cases). Data note: the brief's Zenodo 10.5281/zenodo.15164980 is a DIFFERENT paper (iSCALE) - the real S2-omics demo data is the repo's Google Drive folder. HPC quirks worked around (all our-side, now documented in kartei): HOME group-quota hard-full -> HOME+caches redirected to «infra»; git absent on compute nodes -> installed into env; short TMPDIR for torch DataLoader AF_UNIX sockets.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-14 ⛓ 88dc744354cc
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetBecause histology image patterns correlate with underlying spatial molecular profiles, H&E-based automated selection of regions of interest (ROIs) can replace manual, subjective ROI selection while still maximizing the molecular information content captured in cost- and area-limited spatial omics experiments.
- ★ S2-omics is an end-to-end workflow that automatically selects ROIs from H&E histology images to maximize molecular information content for spatial omics profiling. method
- ★ S2-omics introduces an ROI score metric, computed from histology-based clustering, to quantify the representativeness of candidate regions. method
- ★ Similar histological patterns correspond to similar spatial molecular profiles, justifying histology-guided ROI selection. mechanism
- ★ S2-omics includes a whole-slide molecular information recovery module that broadcasts cell type and cell community labels from profiled ROIs to unsampled tissue regions using histology features. method
- ★ Across gastric and colorectal cancer samples and multiple spatial omics platforms (Xenium, Visium HD, CosMx), S2-omics-selected ROIs match or outperform manual pathologist selections. finding
- ★ Suboptimal ROI selection can miss key molecular signals (e.g., missing enteroendocrine cells and specific cell communities), reducing downstream discovery potential. finding
- S2-omics uses adjacent-section H&E imaging (rather than the target section) for ROI selection to preserve tissue integrity for molecular profiling. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Xenium spatial transcriptomics | human gastric cancer tissue section (~10 mm x 24 mm) | none (histology-guided ROI selection) | cell type and cell community label prediction accuracy, tumour cell and TLS detection, ROI score benchmarking | 10x Genomics Xenium |
| Visium HD spatial transcriptomics | human colorectal cancer section P1 CRC (20 mm x 16 mm) | none (histology-guided ROI selection vs expert/pathologist selection) | cell type prediction accuracy, ROI overlap with expert-selected ROI, histology cluster composition | 10x Genomics Visium HD |
| Visium HD spatial transcriptomics | human colorectal cancer section P5 CRC | none (histology-guided ROI selection) | cell type prediction, histology cluster composition comparison | 10x Genomics Visium HD |
| Visium HD spatial transcriptomics | additional colon cancer samples and one healthy colon sample (different patients) | none (histology-guided ROI selection) | ROI overlap with pathologist-selected ROIs | 10x Genomics Visium HD |
| Histology image feature extraction / unsupervised clustering | human breast, colon, kidney, liver and stomach tissue (H&E images) | none | histology cluster segmentation used to derive ROI score | UNI pathology foundation model |
| Spatial transcriptomics/proteomics (platform validation) | multiple tissue types | none | ROI enrichment for biologically informative patterns vs manual selection | CosMx (in addition to Xenium and Visium HD) |
- ▲ S2-omics-selected ROI achieved the highest ROI score among 500 randomly sampled ROIs and ranked top three for cell type prediction accuracy and top ten for cell community prediction accuracy. ROI score 0.73 (vs benchmark ROI score 0.65 for a mediocre example)
- ▲ Cell type and cell community label broadcasting from the S2-omics ROI achieved strong prediction accuracy in the gastric cancer sample. 73.8% (cell type), 72.8% (cell community)
- ▲ Tumour cells were correctly predicted from the broadcasting model trained only on the selected ROI. 73.4%
- ▲ Of 40 tertiary lymphoid structures (TLSs) in the tissue, 4 were directly captured in the ROI and 28 of the remaining 36 were recovered via broadcasting. 32/40 (~80%) total TLS recovery
- ▲ S2-omics-selected ROI in CRC sample P1 showed strong concordance with the expert (10x Genomics)-selected ROI, covering more expert-selected cells and more valid tissue area. 89.3% coverage of expert-selected cells; 16.3% more valid superpixels
- ▼ A randomly selected ROI with a mediocre ROI score captured very few enteroendocrine cells and several cell communities, lowering downstream cell-label prediction accuracy. ROI score 0.65 (vs 0.73 optimal)
- other ROI score = 0.73 (highest-scoring ROI among 500 random 4mm x 4mm candidates in gastric cancer sample)
- other ROI score = 0.65 (mediocre-scoring randomly selected ROI used to illustrate suboptimal selection)
- other 73.8% accuracy (cell type label broadcasting accuracy from ROI to whole gastric tissue slide)
- other 72.8% accuracy (cell community label broadcasting accuracy from ROI to whole gastric tissue slide)
- other 73.4% correctly predicted (tumour cell prediction accuracy from broadcasting model)
- count 4 of 40 TLSs directly in ROI; 28 of remaining 36 recovered (tertiary lymphoid structure capture/recovery in gastric cancer sample)
- other 89.3% cell coverage; 16.3% more valid superpixels (S2-omics vs expert-selected 6.5 mm x 6.5 mm ROI overlap in CRC sample P1)
- other ~US $7,000 per sample for 6.5 mm x 6.5 mm capture area (cost/area constraint of Visium HD platform motivating ROI selection)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
S2-omics is a computational methods paper evaluated primarily through accuracy-based benchmarking rather than classical inferential statistical tests. The main evaluation strategy compares the method's output (selected ROIs and their downstream label-broadcasting predictions) against ground-truth spatial transcriptomics annotations, using classification accuracy (% correct predictions) as the primary performance metric. Benchmarking is conducted against a distribution of 500 randomly selected ROIs to contextualize the method's ranking, and qualitative comparisons are made between automated and expert-selected ROIs. Results are reported as proportions and percentages with no formal hypothesis tests or confidence intervals stated.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Classification accuracy (% correctly predicted labels) | Cell type and cell community label broadcasting on gastric cancer Xenium sample (73.8% and 72.8% respectively) | All superpixels in the tissue slice; 40 TLSs total; exact n of superpixels not stated | not stated |
| Rank-based benchmarking (ordinal ranking of ROI score and accuracy among 500 random candidates) | Extended Data Fig. 1c — S2-omics ROI ranked in top 3 for cell type accuracy and top 10 for cell community accuracy among 500 randomly selected ROIs | 500 randomly selected 4 mm × 4 mm ROIs | not stated |
| Custom ROI score (proprietary metric quantifying histological cluster representativeness) | ROI selection step across all tissue samples; S2-omics ROI score 0.73 vs mediocre ROI score 0.65 | — | not stated |
| Coverage overlap proportion | Comparison between S2-omics and 10x Genomics expert ROI on P1 CRC (89.3% coverage, 16.3% more valid superpixels) | — | not stated |
-
Prediction accuracy (single point estimate) is the sole quantitative measure of ROI quality in the label-broadcasting evaluation↳ Could also: Bootstrapped confidence intervals or k-fold cross-validation across superpixel splits could accompany each accuracy figure — A single accuracy estimate on one tissue section provides no measure of its variability; confidence intervals would convey whether differences in accuracy between ROIs (e.g., 73.8% vs. a competing method) are meaningful or within sampling noise
-
Comparison between S2-omics and expert selections is primarily visual and overlap-proportion-based↳ Could also: An inter-rater agreement statistic such as Cohen's kappa or Dice coefficient could quantify spatial overlap between the two ROI boundaries — Spatial overlap statistics formalize what visual comparison approximates, making the agreement between automated and expert selection directly comparable across samples and reproducible in future studies
-
Performance is contextualized by rank among 500 randomly sampled ROIs without a formal test against that null distribution↳ Could also: A one-sample permutation test or z-score against the empirical distribution of the 500 random ROI scores/accuracies could be reported — Rank alone (top 3 of 500) does not indicate whether the difference from the median random ROI is statistically distinguishable from chance; a permutation-based p-value or standardized effect size would quantify this
-
Accuracy is the single classification performance metric used throughout↳ Could also: Macro-averaged F1-score, balanced accuracy, or Cohen's kappa could also be reported, especially given potentially imbalanced cell-type class distributions — In tissues with rare cell types (e.g., enteroendocrine cells, TLSs), overall accuracy can be inflated by the majority class; metrics that weight classes equally would more directly capture recovery of rare populations, which is a stated goal of the method
-
Evaluation is conducted on a small number of independently selected tissue samples without a held-out test set formally separated from development↳ Could also: A leave-one-sample-out or nested cross-validation scheme across the full sample collection could estimate generalization performance — With few samples, performance on any individual sample can be strongly influenced by that sample's characteristics; cross-sample validation would provide a less optimistic estimate of how well the method generalizes to new tissues
-
The ROI score is reported as an absolute scalar value (e.g., 0.73 vs. 0.65) without a defined reference distribution↳ Could also: Reporting the ROI score as a percentile within the empirical distribution of the 500 random ROIs, or providing its mean and SD across the random distribution, would give interpretable context — An absolute score is difficult to interpret without knowing the range and spread of that score under random selection; normalizing or referencing against a null distribution would make the metric intuitive for readers comparing across experiments
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41298871 (S2-omics)
Paper: Yuan et al., "Smart spatial omics (S2-omics) optimizes region of interest selection to capture molecular heterogeneity in diverse tissues." Nat Cell Biol 2025. DOI 10.1038/s41556-025-01811-w. PMID 41298871.
Code: https://github.com/ddb-qiwang/S2Omics @ commit a25634ceffc8591d5cdc9ee0b60c09974ed2bad2
(authors' own tool). Demo data + UNI checkpoint: Google Drive folder
1z1nk0sF_e25LKMyHxJVMtROFjuWet2G_ (linked from README; not the brief's
Zenodo record — see note below).
What S2-omics does (pipeline, in scope)
A histology-image → ROI-selection pipeline. Steps in run_roi_selection_single.py:
p1_histology_preprocess— readhe-raw.jpg+pixel-size-raw.txt, makehe.jpg.p2_superpixel_quality_control— HistoSweep texture QC mask of tissue superpixels.p3_feature_extraction— UNI ViT-L/16 foundation-model features per 224px patch (GPU), 224-level + 16-level features concatenated + layernormed.p4_get_histology_segmentation— PCA(80) + KMeans(k=20) morphology clusters.p5_merge_over_clusters— hierarchical merge 20 → 15 clusters.p6_roi_selection_rectangle— sample candidate ROIs, score by scale/coverage/balance, fuse, pick best ROI(s).
Every stochastic step calls setup_seed(42) (torch+numpy+random seeded,
cudnn.deterministic=True); requirements.txt pins exact versions
(numpy 1.26.0, scikit-learn 1.7.0, torch 2.0.1, timm 1.0.9 …). => the pipeline
is deterministic, so the demo (Tutorial 1) outputs are an exact-match target.
IN SCOPE (attempted — Tutorial 1, VisiumHD colorectal cancer demo section)
The repo ships a fully-rendered tutorial notebook
(notebooks/Tutorial_1_VisiumHD_ROI_selection_colon.ipynb) whose printed outputs
are the reference values. We re-run the same pipeline with the documented command
(--roi_size 6.5 6.5 --num_roi 1, defaults k=20→15) on the same demo image and
compare numerically. See original/claims.tsv.
OUT OF SCOPE (not attempted, and why)
- Wet-lab spatial-omics generation (VisiumHD/CosMx/Xenium sequencing, library prep, sectioning) — experimental, not a pipeline.
- Biological validation / heterogeneity-capture claims in the main figures (comparison of S2-omics ROIs vs random/expert ROIs across many tissues) — these require the full multi-sample datasets, manual expert annotations, and downstream spatial-omics readouts that are not shipped with the demo. Out of the 80/20 core.
- Label broadcasting accuracy (Tutorial 1 step 7) — a stochastic NN trained per-run; reference accuracy 0.861@epoch100 is reported but is a soft target. Noted but not graded as a primary claim.
- Tutorials 2–4 (CosMx kidney FOV, consecutive breast ROI, TMA circle ROI) — same pipeline, additional cases; not needed for a clear data point.
Data-link note (auditability)
The brief lists data = zenodo:10.5281/zenodo.15164980. That Zenodo record is
actually the benchmark dataset for a different paper (iSCALE, Nat Methods —
a 31 GB gastric-cancer Xenium sample), not S2-omics demo data. The S2-omics demo
data that reproduces the paper's tutorials lives on the Google Drive folder linked
in the repo README. We reproduce against the Google Drive demo (the artifact the
authors actually ship for reproduction).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.