AI-assisted discovery of an ethnicity-influenced driver of cell transformation in esophageal and gastroesophageal junction adenocarcinomas.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH -> 1:1 REPRODUCED. The paper (Sahoo lab BoNE) classifies normal esophagus (NE) vs Barretts esophagus (BE) with a composite score over two published gene clusters (SLC44A4 24-up, SPINK7 220-down), reporting ROC-AUC 0.88-1.00 across 7 independent validation cohorts (Fig 2D). The authors ship their own analysis notebook (be-eac/be-eac.ipynb) in github.com/sahoo00/BoNE; the exact gene clusters and weights [+1,-1] were recovered verbatim from its cell outputs. We re-implemented the BoNE scoring (StepMiner-z composite, ported from bone.py getRanks2/mergeRanks + MacUtils.fitstep) independently and applied it to the papers own 8 public GEO cohorts on «our HPC». Result: 7 validation cohorts gave AUC 0.876-1.00 (== reported 0.88-1.00; lower bound rounds to 0.88), the GSE100843 training set reproduced the exact 36N/40BE split at AUC 0.9993, and cluster sizes (220 down / 24 up) match exactly. No fabrication detected; all values derivable from shipped code + public data. CAVEAT: we reproduced the downstream scoring/classification using the published fixed clusters; we did NOT re-run the upstream Boolean-network construction (needs the authors non-redistributed Hegemon databases) - so GSE100843 is in-sample while the 7 validation cohorts are genuine out-of-sample. NOT ATTEMPTED: the EAC classifier ROC-AUC (C3 - LNX1 down list incomplete, E-MTAB-4054 ArrayExpress not parsed), wet-lab IHC/organoid experiments, DARC/ACKR1 SNP case-control genotyping, and survival/Cox analyses (out of scope - see scope.md).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 84assessed: 2026-06-16 ⛓ 0fb3c78a3b91
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests what drives cellular transformation from Barrett's esophagus (BE) to esophageal/gastroesophageal junction adenocarcinoma (EAC), and whether some of these drivers are racially/ethnically influenced, explaining the lower EAC incidence in African Americans versus White individuals.
- ★ An AI-guided Boolean network approach (BoNE) models transcriptomic continuum states of normal esophagus, BE, and EAC to derive classifier gene signatures method
- ★ Loss of the SPT6→TP63 axis drives keratinocyte-to-intestinal transcommitment underlying BE formation mechanism
- ★ Boolean invariant logic confirms all EACs must originate via the BE metaplastic intermediate finding
- ★ A CXCL8/IL8-neutrophil immune microenvironment is a driver of cellular transformation in EAC and GEJ adenocarcinoma finding
- ★ This IL-8/neutrophil-driven immune signature is prominent in White individuals but notably absent in African Americans finding
- ★ Absolute neutrophil count (ANC) and neutrophil-related gene signatures track and prognosticate risk of BE to EAC progression finding
- ★ SNPs associated with ethnicity-linked ANC differences (e.g., benign ethnic neutropenia) modify risk of BE to EAC progression finding
- Low GSTT2 expression combined with high neutrophil-centric inflammation may synergize as risk factors in White individuals finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Boolean network transcriptomic modeling (BoNE) | public gene expression datasets (e.g., GSE100843, GSE39491) of normal esophagus, BE, and EAC | none | gene expression classifier clusters (SPINK7/SLC44A4, LNX1/IL10RA/LILRB3) | — |
| RNA-seq/gene expression profiling | human keratinocyte-derived BE organoid model | SPT6 knockdown (siRNA) | differentially expressed genes overlapping network signature | — |
| Immunohistochemistry (IHC) | human FFPE esophageal biopsy specimens from patients with/without BE | none (disease-state comparison) | SPT6 and TP63 protein expression | — |
| ATAC-seq | SPT6-depleted BE organoid model | SPT6 knockdown | chromatin accessibility at SPINK7-cluster genes | — |
| Complete blood cell count | human patient cohort (NDBE, DBE, EAC; Brazil) | none (disease-stage comparison) | absolute neutrophil count, lymphocyte count, platelet count, leukocyte count | — |
| Transcriptomic prognostic signature analysis | TCGA EAC, ESCC, and gastric adenocarcinoma datasets | none | prognostic association of EAC/neutrophil/CXCL8 signatures with outcome | — |
| Race/sex-stratified transcriptomic analysis | GSE77563 esophageal squamous mucosa from White and African American patients | none (race/diagnosis comparison) | EAC signature, tumor inflammation signature (TIS), neutrophil signatures, GSTT2 expression | — |
| Genomic case-control study | human BE progression patient cohort | none | SNPs associated with ethnicity-linked ANC changes (e.g., benign ethnic neutropenia) | — |
- – BE gene signature classified samples across 7 independent validation cohorts ROC AUC 0.88-1.00
- – Network-derived BE signature overlaps significantly with organoid model differentially expressed genes (up and down) P = 1.37e-4 and 8.65e-63
- ▼ SPT6 and TP63 protein levels are suppressed in esophageal squamous lining of patients with BE versus without P = 0.8e-9 and 0.9e-7
- – CXCL8 high => SLC44A4 high is an invariant Boolean implication relationship across NE, BE, and EAC samples
- ▲ IL-8 and neutrophil process signatures are induced in two waves: NE to NDBE and BE-dysplasia to EAC
- ▲ Protumor N2 tumor-associated neutrophil (TAN) signature is induced in EAC and GEJ-AC
- ▲ ANC is the most significant variable tracking risk of NDBE to DBE to EAC progression in univariate and multivariate analyses
- – EAC signature and TIS are induced in histologically normal squamous lining of White patients with BE but not African American patients with BE
- correlation r range 0.8-0.99 (TIS versus EAC signature correlation across EAC and GEJ-AC datasets)
- pvalue P = 1.37 x 10^-4 (overlap of upregulated genes between network BE signature and organoid model)
- pvalue P = 8.65 x 10^-63 (overlap of downregulated genes between network BE signature and organoid model)
- pvalue P = 0.8 x 10^-9 (SPT6 protein suppression in BE vs non-BE squamous lining (IHC))
- pvalue P = 0.9 x 10^-7 (TP63 protein suppression in BE vs non-BE squamous lining (IHC))
- pvalue P = 1.59 x 10^-10 (overlap of 274 aberrantly methylated genes with SLC44A4 cluster)
- count n = 932 (samples across public gene expression datasets used for model validation)
- count n = 113 (cross-sectional cohort of patients with BE and EAC)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper uses an AI-guided Boolean network approach (BoNE) trained on transcriptomic data sets to derive gene signatures for esophageal metaplasia (NE→BE) and neoplastic transformation (BE→EAC), then validates those signatures in up to 7 independent public cohorts (total n=932 samples) using ROC AUC. Key biological predictions are further examined in human organoid models, patient biopsy IHC, a retrospective clinical cohort (n=113), and TCGA prognostic analyses. Supporting statistical approaches include significance tests for gene-set overlaps, pairwise comparisons of signature induction across sequential disease stages, Pearson correlations among signatures, and univariate/multivariate analyses of absolute neutrophil count (ANC) as a predictor of disease progression.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Machine-learning classifier evaluated by ROC AUC | Classification of NE vs. BE (7 independent validation cohorts) and BE vs. EAC (4 independent validation cohorts) | n=932 samples cited across all cohorts; individual cohort sizes not all enumerated in provided text | not stated |
| Overlap significance test (specific method not named; likely hypergeometric or Fisher's exact) | Overlap of organoid DEGs with SPINK7/SLC44A4 clusters (P=1.37×10^-4 and 8.65×10^-63); overlap of 274 aberrantly methylated genes with SLC44A4 cluster (P=1.59×10^-10) | null | not stated |
| Group comparison, specific test not named (IHC) | SPT6 and TP63 protein expression in BE vs. non-BE biopsy specimens (P=0.8×10^-9 and 0.9×10^-7) | null | not stated |
| Pairwise comparisons with ROC AUC and P values (specific test not named) | Gene signature induction at sequential disease stages: NE vs. NDBE, NDBE vs. DBE, DBE vs. EAC, and EAC vs. GEJ-AC (Figure 4F) | null | not stated |
| Pearson correlation | TIS vs. EAC signatures across EAC and GEJ-AC data sets (r range 0.8–0.99) | null | not stated |
| Univariate and multivariate analysis (specific model not named; likely logistic regression or ordinal regression) | ANC as predictor of NDBE→DBE→EAC progression in retrospective Brazilian cohort | n=113 (NDBE n=72, DBE n=11, EAC n=30) | not stated |
-
Multiple pairwise comparisons of gene signatures across sequential disease stages (NE→NDBE→DBE→EAC) were performed without a stated multiplicity correction↳ Could also: A single omnibus test (e.g., Kruskal-Wallis or one-way ANOVA) on the ordered groups followed by post-hoc pairwise tests with Benjamini-Hochberg FDR or Bonferroni correction would also be standard — Omnibus testing followed by post-hoc correction explicitly controls the family-wise error rate or FDR across the family of comparisons, which is a common expectation when the same outcome is tested at multiple sequential stages
-
Gene-set overlap significance was reported with exact p-values but without naming the specific statistical test or the gene universe used↳ Could also: A hypergeometric test or Fisher's exact test with an explicitly defined gene universe (e.g., all genes expressed in the data set or all annotated human genes) is standard practice for overlap analysis — Naming the test and the universe size allows readers to reproduce the calculation and evaluate how sensitive the p-value is to the choice of background universe, which can have a large effect on enrichment results
-
Pearson correlation was used to assess relationships between TIS and EAC signatures (r 0.8–0.99)↳ Could also: Spearman rank correlation could also be applied, particularly if score distributions are skewed or individual cohort sizes are small — Spearman correlation requires no assumption about distributional form and is more robust to outliers; it is often preferred when the normality of composite gene-signature scores has not been verified
-
Classifier performance was summarized using point-estimate ROC AUC values across validation cohorts without confidence intervals↳ Could also: Bootstrap resampling or the DeLong method for confidence intervals around each AUC, combined with a meta-analytic pooling of AUC across cohorts, would also be standard — Confidence intervals around AUC estimates quantify uncertainty in classifier performance, which is especially informative when individual validation cohorts vary in size and composition
-
The multivariate analysis of ANC and BE→EAC progression is described without naming the specific regression model or listing all covariates included↳ Could also: A proportional odds logistic regression (appropriate for the ordered NDBE→DBE→EAC outcome) or a Cox proportional hazards model (if time-to-progression data were available) are standard named approaches, each requiring explicit covariate lists — Naming the model family and listing all covariates allows assessment of model assumptions (e.g., proportional odds assumption) and enables reproducibility; the choice also determines what effect-size metric is reported (odds ratio vs. hazard ratio)
-
Dispersion of clinical variables (e.g., ANC values across NDBE, DBE, and EAC groups) was not reported alongside p-values↳ Could also: Reporting median with IQR (for right-skewed count data such as ANC) or mean with SD for each group alongside the significance test result is also standard — Dispersion measures and group-level summary statistics allow readers to judge the clinical magnitude of differences that are statistically significant, particularly important in a cohort of n=113 where even small absolute differences may yield low p-values
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36134663
Paper: Ghosh P, Campos VJ, Vo DT, … Sahoo D. AI-assisted discovery of an ethnicity-influenced driver of cell transformation in esophageal and gastroesophageal junction adenocarcinomas. JCI Insight. 2022. PMID 36134663, PMCID PMC9675486, DOI 10.1172/jci.insight.161334.
Code: github.com/sahoo00/BoNE (GPL-3.0). The general tool is in the repo;
this paper's own analysis notebook is shipped at be-eac/be-eac.ipynb (author
Daniella Vo) plus the broader BE notebook BE-Analysis.ipynb, with the cohort
loaders and the published gene clusters embedded as executed-cell outputs.
Method (BoNE): StepMiner-binarized gene expression → Boolean-implication network → clusters of co-regulated genes → a composite score = weighted sum of per-cluster StepMiner-z sums → ROC-AUC for binary sample classification. This is a deterministic bioinformatic pipeline (no training/randomness beyond the network construction, which is published as fixed gene clusters).
In scope (pipeline-derived → attempted)
| # | Result | Where | Pipeline |
|---|---|---|---|
| C1 | NE-vs-BE classification ROC-AUC = 0.88–1.00 across 7 independent validation cohorts (+ training GSE100843, GSE39491) using the SLC44A4-up(24)/SPINK7-down(220) clusters, weights [+1,−1] | Fig 2D | BoNE composite-score → roc_auc_score |
| C2 | Cluster sizes: 220 genes down (SPINK7 cluster), 24 genes up (SLC44A4 cluster) in BE vs NE | Results / Suppl Table 2 | BoNE Boolean network on GSE100843 |
| C3 | EAC transformation clusters: LNX1 (down), IL10RA(30)/LILRB3(31) up; CXCL8/IL8↔neutrophil driver; EAC-from-BE classification | Fig 3 / Suppl Table 3 | BoNE composite-score → ROC-AUC (secondary target) |
Primary reproduction target = C1 + C2 (the cleanest, fully-specified pipeline output: published gene lists applied to public GEO data → ROC-AUC). C3 is a stretch target attempted if C1/C2 land.
Datasets (all public GEO / ArrayExpress; NE vs BE)
- GSE100843 Cummings 2017 (GPL6244) — training, n=76 (36 N / 40 BE)
- GSE39491 Hyland 2014 (GPL571) — NE vs BE
- GSE65013 Yamamoto 2015 #1 (GPL5175); GSE64894 Yamamoto 2015 #2 (GPL570)
- GSE49292 McKeon 2015 (GPL5175); GSE26886 Wang 2013 (GPL570)
- GSE34619 Lao-Sirieix 2012 (GPL6244); GSE13083 Stairs 2008 (GPL96)
- E-MTAB-4054 Maag 2017 (N/BE/EAC) — used for EAC map (C3)
Out of scope (NOT attempted)
- Wet-lab / clinical: IHC, organoid SPT6 knockdown experiments, the Brazil patient cohort histology, survival follow-up collection.
- DARC/ACKR1 SNP / benign-ethnic-neutropenia genotyping (Suppl Table 6) — a case-control genotype association, not a pipeline result reproducible from the shipped expression data.
- Boolean network construction from raw Hegemon databases — the authors'
preprocessed Hegemon
.exprdatabases (hegemon.ucsd.edu) are not redistributed; we instead take the published fixed gene clusters (Suppl Table 2/3, recovered verbatim from the notebook outputs) and re-apply the BoNE score to public GEO data. This is a faithful independent reproduction of the scoring step, not of the upstream network inference. - Survival/prognosis Cox analyses (Fig 5) — depend on clinical metadata not in the expression accessions.
Pinned artifacts
- Repo: github.com/sahoo00/BoNE (cloned on «infra»; commit SHA recorded in manifest).
- Gene clusters:
original/be_clusters.json(verbatim from be-eac.ipynb cell 14). - «infra» work dir:
«path»
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The primary NE-vs-BE BoNE classifier reproduces essentially 1:1: 7 out-of-sample validation cohorts give ROC-AUC 0.876-1.00 vs the reported 0.88-1.00 (the only sub-bound value, GSE39491 0.8763, rounds to 0.88), and cluster sizes (220 down / 24 up) match exactly. All values are derivable from the authors' shipped be-eac.ipynb clusters + public GEO data — no fabrication. Caveats are scope-side, not defects: the GSE100843 training AUC (0.9993) is in-sample because the cluster-generating Boolean network wasn't re-run (Hegemon DBs not redistributed), and the secondary C3 EAC classifier was only partially reproduced. Overall a clean, well-specified reproduction.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.