Artificial intelligence-guided discovery of gastric cancer continuum.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH TO REPRODUCE; result is a 1:1 reproduction of the paper's core pipeline output (close, honest, not drop). BoNE (github.com/sahoo00/BoNE, GPL-3.0, master @ c950951) is the authors' general Boolean-network tool; the gastric-cancer analysis code is NOT shipped, but the GC-BoNE cluster gene lists ARE (Suppl. Online Resource 3) and the scoring algorithm is fully specified. We extracted the 6 cluster gene lists (sizes match the paper EXACTLY: 240/507/134/28/14/23), reimplemented BoNE's composite-score -> ROC-AUC verbatim from bone.py (getRanks2/mergeRanks) + MacUtils.py (StepMiner fitstep/getThrData), and applied it on «our HPC»/«infra» to the paper's own GEO datasets. RESULTS vs paper Fig 1c: C#11-2-4-14 on GSE37023/GPL97 (the platform whose 36-normal+29-tumor split EXACTLY equals the paper's n=65) reproduces ROC-AUC 0.936 vs reported 0.96 (within-tol, |delta|=0.024); C#7-13-14 on GSE122401 (RNA-seq) reproduces ~0.89 vs reported 0.98 (partial -- the dataset ships only RSEM ISOFORM/ENST data with no gene symbols, so symbols were mapped to transcripts via mygene.info and aggregated, an approximation of the paper's Hegemon gene-level pipeline that accounts for the gap). On the network-construction cohort GSE66229 the score separates normal vs tumor at AUC 0.969 with full gene coverage. A weight-scheme sensitivity (the exact weights are NOT published) confirms the paper's stated direction-based weighting rule is the operative one and is NOT a free parameter we tuned -- naive monotonic weights collapse to ~0.5. NO FABRICATION SIGNAL: every reproduced AUC is <= the reported value (never inflated) and all gene/sample counts match the paper exactly. NOT ATTEMPTED (80/20): the Boolean-network construction itself (StepMiner+BIR over the whole transcriptome on GSE66229), the Fig 2 21-dataset validation panel (avg AUC 0.933), Fig 3 progression, Fig 4 intestinal metaplasia, survival/HR, and all wet-lab/IHC results; and we deliberately did not tune cluster weights to close the residual 0.02-0.09 AUC gaps.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 71assessed: 2026-06-14 ⛓ 650c9187e214
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether an AI-guided Boolean implication network (BoNE), built from asymmetric Boolean implication relationships in gastric cancer transcriptomic data, can model the healthy mucosa → gastric cancer continuum and better predict pre-neoplastic progression (atrophic gastritis → intestinal metaplasia → low/high-grade neoplasia → GC) than existing gene signatures.
- ★ A Boolean implication network built from GSE66229 yields a GC-BoNE gene signature (Boolean paths C#11-2-4-14 and C#7-13-14) that classifies tumor vs normal/adjacent-normal gastric samples finding
- ★ GC-BoNE outperforms previously published gene signatures at distinguishing normal vs GC samples across 21 validation datasets (average ROC-AUC 0.933 vs 0.690–0.921) finding
- ★ GC-BoNE tracks progressively increasing risk along the metaplasia→dysplasia→neoplasia continuum despite not being trained on those datasets, outperforming other signatures (average ROC-AUC 0.828 vs 0.633–0.806) finding
- ★ GC-BoNE (C#11-2-4-14) can prognosticate risk of progression from incomplete intestinal metaplasia (IIM) to GC, distinguishing progressors from non-progressors (IIM-C vs IIM-GC ROC-AUC 0.95) where other signatures fail finding
- ★ GC-BoNE can objectively rank 38 mouse models (20 GEO datasets) by how well they recapitulate human GC gene expression changes, with H. felis infection and CDH1/SMAD4/CLDN18 knockout GEMMs ranking highest finding
- ★ Boolean Network Explorer (BoNE), using StepMiner Boolean thresholding and Boolean implication relationships (BIRs), is introduced as a computational method to construct disease continuum maps from gene expression data method
- Cluster 11-2-4 changes are associated with progression from healthy to IIM, while cluster 14 changes are associated with progression from IIM to GC mechanism
- Reactome pathway analysis links GC-BoNE clusters to muscle contraction (down), cell cycle, immune/neutrophil degranulation, ion channel transport (down), and extracellular matrix processes (up) finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| microarray transcriptomics | human gastric tumor and adjacent normal tissue (GSE66229) | none (disease state comparison) | gene expression used to build Boolean implication network | — |
| microarray transcriptomics | human gastric tissue, training dataset GSE37023 (GPL96 Affymetrix U133A) | none | ROC-AUC classification of normal vs GC via multivariate OLS regression | Affymetrix Human Genome U133A Array |
| microarray transcriptomics | human gastric tissue, training dataset GSE122401 | none | ROC-AUC classification of normal vs GC | — |
| microarray transcriptomics | 21 independent human GC validation datasets | none | ROC-AUC of GC-BoNE vs published gene signatures for tumor vs normal classification | — |
| microarray transcriptomics | human gastric mucosa, E-MTAB-8889 (NAG/CG/CAG/IM stages) | none (natural disease progression) | GC-BoNE composite score across gastritis-to-metaplasia progression | — |
| microarray transcriptomics | human gastric mucosa, GSE55696 (CG/LGIN/HGIN/EGC stages) | none (natural disease progression) | GC-BoNE composite score across dysplasia-to-neoplasia progression | — |
| transcriptomics | 38 mouse models from 20 NCBI GEO datasets (e.g., GSE13873, GSE103639, GSE45956, GSE16902, GSE93774) | H. felis infection or genetic knockout (CDH1, SMAD4, CLDN18) | ROC-AUC and Welch's t test ranking of similarity to human GC gene expression continuum | — |
| microarray transcriptomics | human prospective cohort, GSE78523 (HC, IIM-C, IIM-GC, CIM-C, CIM-GC) | none (long-term follow-up, mean 12±3.4 years) | ROC-AUC classification of progressors vs non-progressors using GC-BoNE clusters | — |
- ▲ C#11-2-4-14 classified normal vs GC with ROC-AUC 0.96 in training dataset GSE37023 ROC-AUC=0.96
- ▲ C#7-13-14 classified normal vs GC with ROC-AUC 0.98 in training dataset GSE122401 ROC-AUC=0.98
- ▲ GC-BoNE outperformed other signatures across 21 validation datasets avg ROC-AUC 0.933 vs 0.690–0.921
- ▲ GC-BoNE outperformed other signatures on progression (gastritis→metaplasia→dysplasia→neoplasia) datasets avg ROC-AUC 0.828 vs 0.633–0.806
- ▲ C#11-2-4-14 distinguished IIM non-progressors from progressors (IIM-C vs IIM-GC), unlike comparator signatures ROC-AUC=0.95
- – DEA (Li 2015) signature failed to separate IM progressors from non-progressors ROC-AUC IIM-C vs IIM-GC=0.38; CIM-C vs CIM-GC=0.47
- – H. felis infection mouse model (GSE13873) ranked #1 and CDH1/SMAD4/CLDN18 knockout GEMMs ranked #2-6 for recapitulating human GC gene expression
- – Cluster 14 alone distinguished IIM-C vs IIM-GC better than HC vs IIM-C, while clusters 11-2-4 did the opposite ROC-AUC 0.86 (IIM-C vs IIM-GC) vs 0.63 (HC vs IIM-C) for cluster 14
- fold_change RR=4.48 (95% CI 2.50–8.03) (pooled relative risk of cancer/dysplasia in IIM vs CIM patients (meta-analysis))
- fold_change RR=4.96 (95% CI 2.72–9.04) (relative risk of cancer specifically in IIM vs CIM)
- fold_change RR=4.82 (95% CI 1.45–16.0) (relative risk of dysplasia in IIM vs CIM)
- other average ROC-AUC 0.933 (GC-BoNE) vs 0.690–0.921 (other signatures) (normal vs GC classification across validation datasets)
- other average ROC-AUC 0.828 (GC-BoNE) vs 0.633–0.806 (other signatures) (GC progression dataset classification)
- other ROC-AUC 0.57–1.00 (C#11-2-4-14), 0.66–1.00 (C#7-13-14) (range of classification performance across validation datasets)
- count n=400 (300 GC tumor, 100 patient-matched normal) (GSE66229 dataset used for network construction)
- other mean follow-up 12 ± 3.4 years (GSE78523 prospective cohort follow-up duration for IM progression outcomes)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper applies a Boolean implication network framework (BoNE) to transcriptomic data from public GC datasets to model the healthy-to-cancer continuum. Classification performance across training and 21+ independent validation datasets was evaluated primarily by ROC-AUC. Group differences in composite Boolean path scores were tested with Welch's two-sample t-test, and Ordinary Least Squares regression was used for multivariate model selection. Results were reported with asterisk-coded p-value thresholds and ROC-AUC values; no explicit multiplicity correction for repeated testing across datasets and pairwise comparisons was stated.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Welch's two-sample t-test (two-tailed, unpaired, unequal variance) | Comparison of composite Boolean path scores between experimental groups throughout: tumor vs. adjacent normal, sequential pre-neoplastic stage pairs, mouse model rankings, and IM subtype/outcome groups | Varies by dataset; GSE66229 n=400, GSE37023 n=65, GSE122401 n=160; individual group sizes for E-MTAB-8889, GSE55696, GSE78523 not reported in main text | not stated |
| Ordinary Least Squares (OLS) multivariate regression | Model selection to determine which Boolean path score best distinguishes normal vs. GC in training datasets GSE37023 and GSE122401; coefficients with 95% CIs reported in Fig. 1c | GSE37023 n=65; GSE122401 n=160 | not stated |
| BooleanNet statistics (Boolean implication relationship significance test) | Assessment of significance of pairwise gene-gene Boolean implication relationships during network construction from GSE66229 | n=400 (GSE66229: 300 tumor, 100 patient-matched normal) | not stated |
| ROC-AUC (area under receiver operating characteristic curve) | Primary performance metric for all group classification comparisons across training, validation, progression-stage, mouse model, and IM outcome datasets | Varies by dataset | na |
-
Dozens of pairwise Welch t-tests were performed across 21+ validation datasets, multiple progression stage pairs, and 38 mouse models without a stated multiplicity correction↳ Could also: Apply a false discovery rate correction (e.g., Benjamini-Hochberg) or family-wise error rate correction (e.g., Bonferroni) across the full set of pairwise comparisons — With a large family of simultaneous tests, the expected number of false positives grows; a multiplicity correction is standard in high-throughput comparative genomics studies and would allow readers to assess which findings survive at a controlled error rate
-
Sequential pre-neoplastic stage comparisons (e.g., NAG→CG→CAG→IM; CG→LGIN→HGIN→EGC) were each analyzed as separate pairwise t-tests↳ Could also: Use a one-way ANOVA or linear mixed model with stage as an ordered factor, followed by a post-hoc test (e.g., Tukey HSD or Jonckheere-Terpstra trend test) — A single omnibus test controls the family-wise error rate for the set of stage contrasts; an ordered-alternatives trend test would additionally quantify whether the composite score increases monotonically across the cascade, which is a primary claim of the paper
-
Some training datasets (e.g., GSE122401) used patient-matched tumor and adjacent-normal pairs, but Welch's unpaired t-test was applied throughout↳ Could also: Apply a paired t-test (or Wilcoxon signed-rank test) for matched-pair datasets — Paired tests exploit the within-subject correlation to reduce error variance, generally increasing power when the pairing is informative; using an unpaired test on matched data is conservative but foregoes that efficiency gain
-
Classification performance was evaluated exclusively with ROC-AUC↳ Could also: Also report calibration metrics (e.g., Brier score, calibration curves) or precision-recall AUC, particularly for the prognostic IM progression comparisons where class sizes may be imbalanced — ROC-AUC measures rank-based discrimination but is insensitive to calibration and can be optimistic under class imbalance; precision-recall AUC and calibration curves would provide complementary information relevant to clinical utility
-
OLS regression was used for multivariate model selection between Boolean path scores in two training datasets↳ Could also: Use regularized regression (e.g., LASSO or elastic net) combined with cross-validation for model selection — Regularization penalizes model complexity and cross-validated selection provides an out-of-sample generalization estimate; with relatively modest n (65 and 160) and potentially correlated path scores, this approach would reduce overfitting risk compared to standard OLS
-
p-values are reported only as binned asterisk codes (≤0.05, ≤0.01, ≤0.001) rather than exact values↳ Could also: Report exact p-values (e.g., p = 0.023) for each comparison — Exact p-values allow readers to apply their own significance thresholds, facilitate future meta-analyses, and are recommended by many reporting guidelines (e.g., APA, ICMJE) to increase transparency and reproducibility
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36692601 (GC-BoNE: AI-guided discovery of gastric cancer continuum)
Vo D, Ghosh P, Sahoo D. Gastric Cancer 2023. PMID 36692601 · PMCID PMC9871434 ·
DOI 10.1007/s10120-022-01360-3. Tool: BoNE (Boolean Network Explorer),
https://github.com/sahoo00/BoNE (GPL-3.0, default branch master).
What BoNE does (pipeline)
BoNE builds a Boolean-implication network from a reference cohort, clusters genes into Boolean-equivalent groups, and selects a path of clusters whose weighted, StepMiner-normalised, averaged expression yields a per-sample composite score that orders samples along a biological continuum (here: normal → gastric cancer). Classification performance is reported as ROC-AUC.
In scope (pipeline-derived, attempted)
The paper's central quantitative pipeline output is the GC-BoNE composite score → ROC-AUC for discriminating non-malignant vs gastric-cancer samples, using two Boolean paths over the clusters defined in Suppl. Online Resource 3 (sheet "GC-BoNE"):
| Claim | Path (clusters) | Dataset | Reported ROC-AUC | Location |
|---|---|---|---|---|
| C1 | C#11-2-4-14 | GSE37023 (training) | 0.96 | Fig. 1c |
| C2 | C#7-13-14 | GSE122401 (training) | 0.98 | Fig. 1c |
| C3 | both paths | GSE66229 (network-construction, 300 GC + 100 normal) | not a printed number; expected very high | Fig. 1a context |
Cluster gene lists (exact, from Suppl. 3 — counts verified): Cluster 11 (n=240), Cluster 2 (n=507), Cluster 4 (n=134), Cluster 7 (n=28), Cluster 13 (n=14), Cluster 14 (n=23).
Scoring algorithm reimplemented verbatim from bone.py/SMaRT/MacUtils.py:
- per gene g: StepMiner threshold
thr_g = (m1+m2)/2(best 1-step SSE split,fitstep);getThrData -> t[3] = thr_g + 0.5. - per gene, per sample s:
z = (expr[g,s] - t[3]) / 3 / std_g(expr over samples). - per cluster c:
S_c[s] = Σ_{g∈c} z(a sum, pergetRanks2). - composite:
score[s] = Σ_c weight_c · S_c[s](mergeRanks). - ROC-AUC =
sklearn.roc_curve/auc(label, score). - NB: the per-gene
t[3]term is a sample-independent constant ⇒ it shifts every sample's score equally and does not change ROC-AUC. AUC depends only on the per-gene 1/std normalisation, cluster membership, and the cluster weights.
Out of scope / not attempted (80/20)
- The Boolean-network construction itself on GSE66229 (StepMiner+BIR over the whole transcriptome → the 14 clusters). We take the published cluster definitions (Suppl. 3) as given and reproduce the scoring/AUC step. Rebuilding the network from scratch is the hard last ~20% and is not required to test the reported AUCs.
- The 21-dataset validation panel (Fig. 2), progression (Fig. 3), intestinal metaplasia prognostication (Fig. 4), survival/HR, and all wet-lab/IHC results.
- Exact cluster weight vector: the paper states only the rule (disease-high ⇒ positive graded weight, healthy-high ⇒ negative graded weight). We use the BoNE-canonical monotonic weights along path order and report a small weight sensitivity, rather than claiming a single exact vector.
Data
- GSE66229, GSE37023, GSE122401 — all public GEO, fetched inside the «our HPC» job onto «infra» (series matrix + GPL annotation via GEOparse). Normal vs tumour labels read from GEO sample characteristics.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The paper's central claim — the GC-BoNE composite score discriminates non-malignant from gastric cancer at high ROC-AUC — reproduces faithfully (C1 0.936 vs 0.96, C3 0.969, full gene/sample-count matches and every AUC ≤ reported, so no fabrication signal). The one notable shortfall, C2 0.98→0.89, lies on the input/preprocessing side: GSE122401 deposits only RSEM ENST isoform data, forcing an isoform→gene aggregation that differs from the authors' Hegemon gene-level pipeline, compounded by the unpublished exact weight vector (only the sign+grading rule is given). These deviations are small, explainable, and split between our self-chosen mapping steps and a mild paper underspecification — not an authors' defect or a non-derivable value. Overall a solid reproduction with explainable deviations rather than a pixel-perfect 1:1.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.