Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Smart spatial omics (S2-omics) optimizes region of interest selection to capture molecular heterogeneity in diverse tissues.

Nat Cell Biol · 2025
L1 100/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • Any deviation was negligible
What did not (or only partly)
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH: yes, excellently. S2-omics (PMID 41298871) is the authors' own histology->ROI-selection pipeline (UNI ViT-L/16 features -> PCA+KMeans morphology segmentation 20->merge 15 -> rectangle ROI scoring), fully deterministic (setup_seed(42) everywhere + pinned requirements.txt) with a rendered Tutorial-1 reference notebook. RESULT = 1:1 EXACT (reproduced). I ran run_roi_selection_single.py on «our HPC» (H100, «job») at repo commit a25634c with the documented Tutorial-1 params (--roi_size 6.5 6.5 --num_roi 1 --prior_preference 1, down_samp_step 10, k=20->15) on the authors' shipped demo image + UNI checkpoint (Google Drive). All 8 pipeline-derived outputs reproduced BIT-EXACTLY: 49490 patches; 20->15 clusters; best ROI [[76,8],[146,49],[105,119],[35,78]]; ROI/scale/coverage/balance scores 0.7923578790248927 / 0.6245726221007392 / 0.9455532908082276 / 0.8423550596065301 (identical to 16 sig figs). Minimal documented env deviations (hardware/headless only): torch 2.0.1 from cu118 (H100 sm_90) instead of the cu117 pin, and opencv-python-headless (no libGL on node); everything else exactly per requirements.txt. NOT ATTEMPTED (out of scope, see scope.md): wet-lab spatial-omics generation; the multi-tissue biological heterogeneity-capture main-figure claims (need the non-shipped full cohort + expert annotations + downstream spatial-omics readouts); the stochastic per-run label-broadcasting accuracy (soft target); Tutorials 2-4 (same pipeline, additional cases). Data note: the brief's Zenodo 10.5281/zenodo.15164980 is a DIFFERENT paper (iSCALE) - the real S2-omics demo data is the repo's Google Drive folder. HPC quirks worked around (all our-side, now documented in kartei): HOME group-quota hard-full -> HOME+caches redirected to «infra»; git absent on compute nodes -> installed into env; short TMPDIR for torch DataLoader AF_UNIX sockets.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.15164980

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-14 ⛓ 88dc744354cc
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Because histology image patterns correlate with underlying spatial molecular profiles, H&E-based automated selection of regions of interest (ROIs) can replace manual, subjective ROI selection while still maximizing the molecular information content captured in cost- and area-limited spatial omics experiments.

Core claims
  • S2-omics is an end-to-end workflow that automatically selects ROIs from H&E histology images to maximize molecular information content for spatial omics profiling. method
  • S2-omics introduces an ROI score metric, computed from histology-based clustering, to quantify the representativeness of candidate regions. method
  • Similar histological patterns correspond to similar spatial molecular profiles, justifying histology-guided ROI selection. mechanism
  • S2-omics includes a whole-slide molecular information recovery module that broadcasts cell type and cell community labels from profiled ROIs to unsampled tissue regions using histology features. method
  • Across gastric and colorectal cancer samples and multiple spatial omics platforms (Xenium, Visium HD, CosMx), S2-omics-selected ROIs match or outperform manual pathologist selections. finding
  • Suboptimal ROI selection can miss key molecular signals (e.g., missing enteroendocrine cells and specific cell communities), reducing downstream discovery potential. finding
  • S2-omics uses adjacent-section H&E imaging (rather than the target section) for ROI selection to preserve tissue integrity for molecular profiling. method
Experimental setups
Assay System Perturbation Readout Platform
Xenium spatial transcriptomics human gastric cancer tissue section (~10 mm x 24 mm) none (histology-guided ROI selection) cell type and cell community label prediction accuracy, tumour cell and TLS detection, ROI score benchmarking 10x Genomics Xenium
Visium HD spatial transcriptomics human colorectal cancer section P1 CRC (20 mm x 16 mm) none (histology-guided ROI selection vs expert/pathologist selection) cell type prediction accuracy, ROI overlap with expert-selected ROI, histology cluster composition 10x Genomics Visium HD
Visium HD spatial transcriptomics human colorectal cancer section P5 CRC none (histology-guided ROI selection) cell type prediction, histology cluster composition comparison 10x Genomics Visium HD
Visium HD spatial transcriptomics additional colon cancer samples and one healthy colon sample (different patients) none (histology-guided ROI selection) ROI overlap with pathologist-selected ROIs 10x Genomics Visium HD
Histology image feature extraction / unsupervised clustering human breast, colon, kidney, liver and stomach tissue (H&E images) none histology cluster segmentation used to derive ROI score UNI pathology foundation model
Spatial transcriptomics/proteomics (platform validation) multiple tissue types none ROI enrichment for biologically informative patterns vs manual selection CosMx (in addition to Xenium and Visium HD)
Key results
  • S2-omics-selected ROI achieved the highest ROI score among 500 randomly sampled ROIs and ranked top three for cell type prediction accuracy and top ten for cell community prediction accuracy. ROI score 0.73 (vs benchmark ROI score 0.65 for a mediocre example)
  • Cell type and cell community label broadcasting from the S2-omics ROI achieved strong prediction accuracy in the gastric cancer sample. 73.8% (cell type), 72.8% (cell community)
  • Tumour cells were correctly predicted from the broadcasting model trained only on the selected ROI. 73.4%
  • Of 40 tertiary lymphoid structures (TLSs) in the tissue, 4 were directly captured in the ROI and 28 of the remaining 36 were recovered via broadcasting. 32/40 (~80%) total TLS recovery
  • S2-omics-selected ROI in CRC sample P1 showed strong concordance with the expert (10x Genomics)-selected ROI, covering more expert-selected cells and more valid tissue area. 89.3% coverage of expert-selected cells; 16.3% more valid superpixels
  • A randomly selected ROI with a mediocre ROI score captured very few enteroendocrine cells and several cell communities, lowering downstream cell-label prediction accuracy. ROI score 0.65 (vs 0.73 optimal)
Key statistics
  • other ROI score = 0.73 (highest-scoring ROI among 500 random 4mm x 4mm candidates in gastric cancer sample)
  • other ROI score = 0.65 (mediocre-scoring randomly selected ROI used to illustrate suboptimal selection)
  • other 73.8% accuracy (cell type label broadcasting accuracy from ROI to whole gastric tissue slide)
  • other 72.8% accuracy (cell community label broadcasting accuracy from ROI to whole gastric tissue slide)
  • other 73.4% correctly predicted (tumour cell prediction accuracy from broadcasting model)
  • count 4 of 40 TLSs directly in ROI; 28 of remaining 36 recovered (tertiary lymphoid structure capture/recovery in gastric cancer sample)
  • other 89.3% cell coverage; 16.3% more valid superpixels (S2-omics vs expert-selected 6.5 mm x 6.5 mm ROI overlap in CRC sample P1)
  • other ~US $7,000 per sample for 6.5 mm x 6.5 mm capture area (cost/area constraint of Visium HD platform motivating ROI selection)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

S2-omics is a computational methods paper evaluated primarily through accuracy-based benchmarking rather than classical inferential statistical tests. The main evaluation strategy compares the method's output (selected ROIs and their downstream label-broadcasting predictions) against ground-truth spatial transcriptomics annotations, using classification accuracy (% correct predictions) as the primary performance metric. Benchmarking is conducted against a distribution of 500 randomly selected ROIs to contextualize the method's ranking, and qualitative comparisons are made between automated and expert-selected ROIs. Results are reported as proportions and percentages with no formal hypothesis tests or confidence intervals stated.

Replicationbiological Sample sizeMultiple independent tissue samples described per tissue type (gastric cancer n=1, CRC multiple patients including P1 and P5, healthy colon n=1); sample sizes stated qualitatively; no formal power calculation mentioned GroupsS2-omics-selected ROIs vs. expert/pathologist-selected ROIs vs. randomly selected ROIs; S2-omics predictions vs. spatial transcriptomics ground truth Pairingpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Classification accuracy (% correctly predicted labels) Cell type and cell community label broadcasting on gastric cancer Xenium sample (73.8% and 72.8% respectively) All superpixels in the tissue slice; 40 TLSs total; exact n of superpixels not stated not stated
Rank-based benchmarking (ordinal ranking of ROI score and accuracy among 500 random candidates) Extended Data Fig. 1c — S2-omics ROI ranked in top 3 for cell type accuracy and top 10 for cell community accuracy among 500 randomly selected ROIs 500 randomly selected 4 mm × 4 mm ROIs not stated
Custom ROI score (proprietary metric quantifying histological cluster representativeness) ROI selection step across all tissue samples; S2-omics ROI score 0.73 vs mediocre ROI score 0.65 not stated
Coverage overlap proportion Comparison between S2-omics and 10x Genomics expert ROI on P1 CRC (89.3% coverage, 16.3% more valid superpixels) not stated
Approaches that could also have been used
  • Prediction accuracy (single point estimate) is the sole quantitative measure of ROI quality in the label-broadcasting evaluation
    Could also: Bootstrapped confidence intervals or k-fold cross-validation across superpixel splits could accompany each accuracy figure — A single accuracy estimate on one tissue section provides no measure of its variability; confidence intervals would convey whether differences in accuracy between ROIs (e.g., 73.8% vs. a competing method) are meaningful or within sampling noise
  • Comparison between S2-omics and expert selections is primarily visual and overlap-proportion-based
    Could also: An inter-rater agreement statistic such as Cohen's kappa or Dice coefficient could quantify spatial overlap between the two ROI boundaries — Spatial overlap statistics formalize what visual comparison approximates, making the agreement between automated and expert selection directly comparable across samples and reproducible in future studies
  • Performance is contextualized by rank among 500 randomly sampled ROIs without a formal test against that null distribution
    Could also: A one-sample permutation test or z-score against the empirical distribution of the 500 random ROI scores/accuracies could be reported — Rank alone (top 3 of 500) does not indicate whether the difference from the median random ROI is statistically distinguishable from chance; a permutation-based p-value or standardized effect size would quantify this
  • Accuracy is the single classification performance metric used throughout
    Could also: Macro-averaged F1-score, balanced accuracy, or Cohen's kappa could also be reported, especially given potentially imbalanced cell-type class distributions — In tissues with rare cell types (e.g., enteroendocrine cells, TLSs), overall accuracy can be inflated by the majority class; metrics that weight classes equally would more directly capture recovery of rare populations, which is a stated goal of the method
  • Evaluation is conducted on a small number of independently selected tissue samples without a held-out test set formally separated from development
    Could also: A leave-one-sample-out or nested cross-validation scheme across the full sample collection could estimate generalization performance — With few samples, performance on any individual sample can be strongly influenced by that sample's characteristics; cross-sample validation would provide a less optimistic estimate of how well the method generalizes to new tissues
  • The ROI score is reported as an absolute scalar value (e.g., 0.73 vs. 0.65) without a defined reference distribution
    Could also: Reporting the ROI score as a percentile within the empirical distribution of the 500 random ROIs, or providing its mean and SD across the random distribution, would give interpretable context — An absolute score is difficult to interpret without knowing the range and spread of that score under random selection; normalizing or referencing against a null distribution would make the metric intuitive for readers comparing across experiments
Software: UNI pathology foundation model · 10x Genomics Xenium platform (data generation) · 10x Genomics Visium HD platform (data generation) · CosMx platform (data generation)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
8
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41298871 (S2-omics)

Paper: Yuan et al., "Smart spatial omics (S2-omics) optimizes region of interest selection to capture molecular heterogeneity in diverse tissues." Nat Cell Biol 2025. DOI 10.1038/s41556-025-01811-w. PMID 41298871.

Code: https://github.com/ddb-qiwang/S2Omics @ commit a25634ceffc8591d5cdc9ee0b60c09974ed2bad2 (authors' own tool). Demo data + UNI checkpoint: Google Drive folder 1z1nk0sF_e25LKMyHxJVMtROFjuWet2G_ (linked from README; not the brief's Zenodo record — see note below).

What S2-omics does (pipeline, in scope)

A histology-image → ROI-selection pipeline. Steps in run_roi_selection_single.py:

  1. p1_histology_preprocess — read he-raw.jpg + pixel-size-raw.txt, make he.jpg.
  2. p2_superpixel_quality_control — HistoSweep texture QC mask of tissue superpixels.
  3. p3_feature_extractionUNI ViT-L/16 foundation-model features per 224px patch (GPU), 224-level + 16-level features concatenated + layernormed.
  4. p4_get_histology_segmentation — PCA(80) + KMeans(k=20) morphology clusters.
  5. p5_merge_over_clusters — hierarchical merge 20 → 15 clusters.
  6. p6_roi_selection_rectangle — sample candidate ROIs, score by scale/coverage/balance, fuse, pick best ROI(s).

Every stochastic step calls setup_seed(42) (torch+numpy+random seeded, cudnn.deterministic=True); requirements.txt pins exact versions (numpy 1.26.0, scikit-learn 1.7.0, torch 2.0.1, timm 1.0.9 …). => the pipeline is deterministic, so the demo (Tutorial 1) outputs are an exact-match target.

IN SCOPE (attempted — Tutorial 1, VisiumHD colorectal cancer demo section)

The repo ships a fully-rendered tutorial notebook (notebooks/Tutorial_1_VisiumHD_ROI_selection_colon.ipynb) whose printed outputs are the reference values. We re-run the same pipeline with the documented command (--roi_size 6.5 6.5 --num_roi 1, defaults k=20→15) on the same demo image and compare numerically. See original/claims.tsv.

OUT OF SCOPE (not attempted, and why)

  • Wet-lab spatial-omics generation (VisiumHD/CosMx/Xenium sequencing, library prep, sectioning) — experimental, not a pipeline.
  • Biological validation / heterogeneity-capture claims in the main figures (comparison of S2-omics ROIs vs random/expert ROIs across many tissues) — these require the full multi-sample datasets, manual expert annotations, and downstream spatial-omics readouts that are not shipped with the demo. Out of the 80/20 core.
  • Label broadcasting accuracy (Tutorial 1 step 7) — a stochastic NN trained per-run; reference accuracy 0.861@epoch100 is reported but is a soft target. Noted but not graded as a primary claim.
  • Tutorials 2–4 (CosMx kidney FOV, consecutive breast ROI, TMA circle ROI) — same pipeline, additional cases; not needed for a clear data point.

Data-link note (auditability)

The brief lists data = zenodo:10.5281/zenodo.15164980. That Zenodo record is actually the benchmark dataset for a different paper (iSCALE, Nat Methods — a 31 GB gastric-cancer Xenium sample), not S2-omics demo data. The S2-omics demo data that reproduces the paper's tutorials lives on the Google Drive folder linked in the repo README. We reproduce against the Google Drive demo (the artifact the authors actually ship for reproduction).

n_patches
Reported
49490
Reproduced
49490
exact
n_clusters_init
Reported
20
Reproduced
20
exact
n_clusters_merged
Reported
15
Reproduced
15
exact
roi_coords
Reported
[[76, 8], [146, 49], [105, 119], [35, 78]]
Reproduced
[[76, 8], [146, 49], [105, 119], [35, 78]]
exact
roi_score
Reported
0.7923578790248927
Reproduced
0.7923578790248927
exact
scale_score
Reported
0.6245726221007392
Reproduced
0.6245726221007392
exact
coverage_score
Reported
0.9455532908082276
Reproduced
0.9455532908082276
exact
balance_score
Reported
0.8423550596065301
Reproduced
0.8423550596065301
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

400.6 k
tokens (I/O) · 46.2 M incl. cache
224 min
runtime · 0.54 CPU-h
29 GB
peak RAM
8 (4 failed)
HPC jobs
hummel
machine