Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

AI-assisted discovery of an ethnicity-influenced driver of cell transformation in esophageal and gastroesophageal junction adenocarcinomas.

JCI Insight · 2022
L1 84/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
84/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 63% of all assessed papers rank 392 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH -> 1:1 REPRODUCED. The paper (Sahoo lab BoNE) classifies normal esophagus (NE) vs Barretts esophagus (BE) with a composite score over two published gene clusters (SLC44A4 24-up, SPINK7 220-down), reporting ROC-AUC 0.88-1.00 across 7 independent validation cohorts (Fig 2D). The authors ship their own analysis notebook (be-eac/be-eac.ipynb) in github.com/sahoo00/BoNE; the exact gene clusters and weights [+1,-1] were recovered verbatim from its cell outputs. We re-implemented the BoNE scoring (StepMiner-z composite, ported from bone.py getRanks2/mergeRanks + MacUtils.fitstep) independently and applied it to the papers own 8 public GEO cohorts on «our HPC». Result: 7 validation cohorts gave AUC 0.876-1.00 (== reported 0.88-1.00; lower bound rounds to 0.88), the GSE100843 training set reproduced the exact 36N/40BE split at AUC 0.9993, and cluster sizes (220 down / 24 up) match exactly. No fabrication detected; all values derivable from shipped code + public data. CAVEAT: we reproduced the downstream scoring/classification using the published fixed clusters; we did NOT re-run the upstream Boolean-network construction (needs the authors non-redistributed Hegemon databases) - so GSE100843 is in-sample while the 7 validation cohorts are genuine out-of-sample. NOT ATTEMPTED: the EAC classifier ROC-AUC (C3 - LNX1 down list incomplete, E-MTAB-4054 ArrayExpress not parsed), wet-lab IHC/organoid experiments, DARC/ACKR1 SNP case-control genotyping, and survival/Cox analyses (out of scope - see scope.md).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 84
    assessed: 2026-06-16 ⛓ 0fb3c78a3b91
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests what drives cellular transformation from Barrett's esophagus (BE) to esophageal/gastroesophageal junction adenocarcinoma (EAC), and whether some of these drivers are racially/ethnically influenced, explaining the lower EAC incidence in African Americans versus White individuals.

Core claims
  • An AI-guided Boolean network approach (BoNE) models transcriptomic continuum states of normal esophagus, BE, and EAC to derive classifier gene signatures method
  • Loss of the SPT6→TP63 axis drives keratinocyte-to-intestinal transcommitment underlying BE formation mechanism
  • Boolean invariant logic confirms all EACs must originate via the BE metaplastic intermediate finding
  • A CXCL8/IL8-neutrophil immune microenvironment is a driver of cellular transformation in EAC and GEJ adenocarcinoma finding
  • This IL-8/neutrophil-driven immune signature is prominent in White individuals but notably absent in African Americans finding
  • Absolute neutrophil count (ANC) and neutrophil-related gene signatures track and prognosticate risk of BE to EAC progression finding
  • SNPs associated with ethnicity-linked ANC differences (e.g., benign ethnic neutropenia) modify risk of BE to EAC progression finding
  • Low GSTT2 expression combined with high neutrophil-centric inflammation may synergize as risk factors in White individuals finding
Experimental setups
Assay System Perturbation Readout Platform
Boolean network transcriptomic modeling (BoNE) public gene expression datasets (e.g., GSE100843, GSE39491) of normal esophagus, BE, and EAC none gene expression classifier clusters (SPINK7/SLC44A4, LNX1/IL10RA/LILRB3)
RNA-seq/gene expression profiling human keratinocyte-derived BE organoid model SPT6 knockdown (siRNA) differentially expressed genes overlapping network signature
Immunohistochemistry (IHC) human FFPE esophageal biopsy specimens from patients with/without BE none (disease-state comparison) SPT6 and TP63 protein expression
ATAC-seq SPT6-depleted BE organoid model SPT6 knockdown chromatin accessibility at SPINK7-cluster genes
Complete blood cell count human patient cohort (NDBE, DBE, EAC; Brazil) none (disease-stage comparison) absolute neutrophil count, lymphocyte count, platelet count, leukocyte count
Transcriptomic prognostic signature analysis TCGA EAC, ESCC, and gastric adenocarcinoma datasets none prognostic association of EAC/neutrophil/CXCL8 signatures with outcome
Race/sex-stratified transcriptomic analysis GSE77563 esophageal squamous mucosa from White and African American patients none (race/diagnosis comparison) EAC signature, tumor inflammation signature (TIS), neutrophil signatures, GSTT2 expression
Genomic case-control study human BE progression patient cohort none SNPs associated with ethnicity-linked ANC changes (e.g., benign ethnic neutropenia)
Key results
  • BE gene signature classified samples across 7 independent validation cohorts ROC AUC 0.88-1.00
  • Network-derived BE signature overlaps significantly with organoid model differentially expressed genes (up and down) P = 1.37e-4 and 8.65e-63
  • SPT6 and TP63 protein levels are suppressed in esophageal squamous lining of patients with BE versus without P = 0.8e-9 and 0.9e-7
  • CXCL8 high => SLC44A4 high is an invariant Boolean implication relationship across NE, BE, and EAC samples
  • IL-8 and neutrophil process signatures are induced in two waves: NE to NDBE and BE-dysplasia to EAC
  • Protumor N2 tumor-associated neutrophil (TAN) signature is induced in EAC and GEJ-AC
  • ANC is the most significant variable tracking risk of NDBE to DBE to EAC progression in univariate and multivariate analyses
  • EAC signature and TIS are induced in histologically normal squamous lining of White patients with BE but not African American patients with BE
Key statistics
  • correlation r range 0.8-0.99 (TIS versus EAC signature correlation across EAC and GEJ-AC datasets)
  • pvalue P = 1.37 x 10^-4 (overlap of upregulated genes between network BE signature and organoid model)
  • pvalue P = 8.65 x 10^-63 (overlap of downregulated genes between network BE signature and organoid model)
  • pvalue P = 0.8 x 10^-9 (SPT6 protein suppression in BE vs non-BE squamous lining (IHC))
  • pvalue P = 0.9 x 10^-7 (TP63 protein suppression in BE vs non-BE squamous lining (IHC))
  • pvalue P = 1.59 x 10^-10 (overlap of 274 aberrantly methylated genes with SLC44A4 cluster)
  • count n = 932 (samples across public gene expression datasets used for model validation)
  • count n = 113 (cross-sectional cohort of patients with BE and EAC)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper uses an AI-guided Boolean network approach (BoNE) trained on transcriptomic data sets to derive gene signatures for esophageal metaplasia (NE→BE) and neoplastic transformation (BE→EAC), then validates those signatures in up to 7 independent public cohorts (total n=932 samples) using ROC AUC. Key biological predictions are further examined in human organoid models, patient biopsy IHC, a retrospective clinical cohort (n=113), and TCGA prognostic analyses. Supporting statistical approaches include significance tests for gene-set overlaps, pairwise comparisons of signature induction across sequential disease stages, Pearson correlations among signatures, and univariate/multivariate analyses of absolute neutrophil count (ANC) as a predictor of disease progression.

Replicationbiological Sample sizeTraining sets: n=76 (GSE100843, 18 patients with paired matched specimens), n=43 (GSE39491); 7 validation cohorts totaling n=932; retrospective clinical cohort n=113; TCGA EAC cohort and ESCC pooled cohort (n>400) used for prognostic analyses; no formal power calculation stated GroupsNE vs. BE, BE vs. EAC, NDBE vs. DBE vs. EAC, EAC vs. GEJ-AC, EAC vs. ESCC vs. gastric adenocarcinoma, White vs. African American patients, male vs. female Pairingmixed Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Machine-learning classifier evaluated by ROC AUC Classification of NE vs. BE (7 independent validation cohorts) and BE vs. EAC (4 independent validation cohorts) n=932 samples cited across all cohorts; individual cohort sizes not all enumerated in provided text not stated
Overlap significance test (specific method not named; likely hypergeometric or Fisher's exact) Overlap of organoid DEGs with SPINK7/SLC44A4 clusters (P=1.37×10^-4 and 8.65×10^-63); overlap of 274 aberrantly methylated genes with SLC44A4 cluster (P=1.59×10^-10) null not stated
Group comparison, specific test not named (IHC) SPT6 and TP63 protein expression in BE vs. non-BE biopsy specimens (P=0.8×10^-9 and 0.9×10^-7) null not stated
Pairwise comparisons with ROC AUC and P values (specific test not named) Gene signature induction at sequential disease stages: NE vs. NDBE, NDBE vs. DBE, DBE vs. EAC, and EAC vs. GEJ-AC (Figure 4F) null not stated
Pearson correlation TIS vs. EAC signatures across EAC and GEJ-AC data sets (r range 0.8–0.99) null not stated
Univariate and multivariate analysis (specific model not named; likely logistic regression or ordinal regression) ANC as predictor of NDBE→DBE→EAC progression in retrospective Brazilian cohort n=113 (NDBE n=72, DBE n=11, EAC n=30) not stated
Approaches that could also have been used
  • Multiple pairwise comparisons of gene signatures across sequential disease stages (NE→NDBE→DBE→EAC) were performed without a stated multiplicity correction
    Could also: A single omnibus test (e.g., Kruskal-Wallis or one-way ANOVA) on the ordered groups followed by post-hoc pairwise tests with Benjamini-Hochberg FDR or Bonferroni correction would also be standard — Omnibus testing followed by post-hoc correction explicitly controls the family-wise error rate or FDR across the family of comparisons, which is a common expectation when the same outcome is tested at multiple sequential stages
  • Gene-set overlap significance was reported with exact p-values but without naming the specific statistical test or the gene universe used
    Could also: A hypergeometric test or Fisher's exact test with an explicitly defined gene universe (e.g., all genes expressed in the data set or all annotated human genes) is standard practice for overlap analysis — Naming the test and the universe size allows readers to reproduce the calculation and evaluate how sensitive the p-value is to the choice of background universe, which can have a large effect on enrichment results
  • Pearson correlation was used to assess relationships between TIS and EAC signatures (r 0.8–0.99)
    Could also: Spearman rank correlation could also be applied, particularly if score distributions are skewed or individual cohort sizes are small — Spearman correlation requires no assumption about distributional form and is more robust to outliers; it is often preferred when the normality of composite gene-signature scores has not been verified
  • Classifier performance was summarized using point-estimate ROC AUC values across validation cohorts without confidence intervals
    Could also: Bootstrap resampling or the DeLong method for confidence intervals around each AUC, combined with a meta-analytic pooling of AUC across cohorts, would also be standard — Confidence intervals around AUC estimates quantify uncertainty in classifier performance, which is especially informative when individual validation cohorts vary in size and composition
  • The multivariate analysis of ANC and BE→EAC progression is described without naming the specific regression model or listing all covariates included
    Could also: A proportional odds logistic regression (appropriate for the ordered NDBE→DBE→EAC outcome) or a Cox proportional hazards model (if time-to-progression data were available) are standard named approaches, each requiring explicit covariate lists — Naming the model family and listing all covariates allows assessment of model assumptions (e.g., proportional odds assumption) and enables reproducibility; the choice also determines what effect-size metric is reported (odds ratio vs. hazard ratio)
  • Dispersion of clinical variables (e.g., ANC values across NDBE, DBE, and EAC groups) was not reported alongside p-values
    Could also: Reporting median with IQR (for right-skewed count data such as ANC) or mean with SD for each group alongside the significance test result is also standard — Dispersion measures and group-level summary statistics allow readers to judge the clinical magnitude of differences that are statistically significant, particularly important in a cohort of n=113 where even small absolute differences may yield low p-values
Software: BoNE (Boolean Network Explorer) · Reactome (pathway enrichment analysis)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
16
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

rs2814778 RefSNP in Discussion (http://purl.org/orb/Discussion)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36134663

Paper: Ghosh P, Campos VJ, Vo DT, … Sahoo D. AI-assisted discovery of an ethnicity-influenced driver of cell transformation in esophageal and gastroesophageal junction adenocarcinomas. JCI Insight. 2022. PMID 36134663, PMCID PMC9675486, DOI 10.1172/jci.insight.161334.

Code: github.com/sahoo00/BoNE (GPL-3.0). The general tool is in the repo; this paper's own analysis notebook is shipped at be-eac/be-eac.ipynb (author Daniella Vo) plus the broader BE notebook BE-Analysis.ipynb, with the cohort loaders and the published gene clusters embedded as executed-cell outputs.

Method (BoNE): StepMiner-binarized gene expression → Boolean-implication network → clusters of co-regulated genes → a composite score = weighted sum of per-cluster StepMiner-z sums → ROC-AUC for binary sample classification. This is a deterministic bioinformatic pipeline (no training/randomness beyond the network construction, which is published as fixed gene clusters).

In scope (pipeline-derived → attempted)

# Result Where Pipeline
C1 NE-vs-BE classification ROC-AUC = 0.88–1.00 across 7 independent validation cohorts (+ training GSE100843, GSE39491) using the SLC44A4-up(24)/SPINK7-down(220) clusters, weights [+1,−1] Fig 2D BoNE composite-score → roc_auc_score
C2 Cluster sizes: 220 genes down (SPINK7 cluster), 24 genes up (SLC44A4 cluster) in BE vs NE Results / Suppl Table 2 BoNE Boolean network on GSE100843
C3 EAC transformation clusters: LNX1 (down), IL10RA(30)/LILRB3(31) up; CXCL8/IL8↔neutrophil driver; EAC-from-BE classification Fig 3 / Suppl Table 3 BoNE composite-score → ROC-AUC (secondary target)

Primary reproduction target = C1 + C2 (the cleanest, fully-specified pipeline output: published gene lists applied to public GEO data → ROC-AUC). C3 is a stretch target attempted if C1/C2 land.

Datasets (all public GEO / ArrayExpress; NE vs BE)

  • GSE100843 Cummings 2017 (GPL6244) — training, n=76 (36 N / 40 BE)
  • GSE39491 Hyland 2014 (GPL571) — NE vs BE
  • GSE65013 Yamamoto 2015 #1 (GPL5175); GSE64894 Yamamoto 2015 #2 (GPL570)
  • GSE49292 McKeon 2015 (GPL5175); GSE26886 Wang 2013 (GPL570)
  • GSE34619 Lao-Sirieix 2012 (GPL6244); GSE13083 Stairs 2008 (GPL96)
  • E-MTAB-4054 Maag 2017 (N/BE/EAC) — used for EAC map (C3)

Out of scope (NOT attempted)

  • Wet-lab / clinical: IHC, organoid SPT6 knockdown experiments, the Brazil patient cohort histology, survival follow-up collection.
  • DARC/ACKR1 SNP / benign-ethnic-neutropenia genotyping (Suppl Table 6) — a case-control genotype association, not a pipeline result reproducible from the shipped expression data.
  • Boolean network construction from raw Hegemon databases — the authors' preprocessed Hegemon .expr databases (hegemon.ucsd.edu) are not redistributed; we instead take the published fixed gene clusters (Suppl Table 2/3, recovered verbatim from the notebook outputs) and re-apply the BoNE score to public GEO data. This is a faithful independent reproduction of the scoring step, not of the upstream network inference.
  • Survival/prognosis Cox analyses (Fig 5) — depend on clinical metadata not in the expression accessions.

Pinned artifacts

  • Repo: github.com/sahoo00/BoNE (cloned on «infra»; commit SHA recorded in manifest).
  • Gene clusters: original/be_clusters.json (verbatim from be-eac.ipynb cell 14).
  • «infra» work dir: «path»
Figures / tables: Fig 2DTableFig 3
C1
Reported
ROC-AUC 0.88-1.00 across 7 independent validation cohorts (NE vs BE, BoNE SLC44A4-up/SPINK7-down composite)
Reproduced
0.876-1.00 across the same 7 validation cohorts (GSE39491 0.876, GSE65013 1.0, GSE64894 1.0, GSE49292 1.0, GSE26886 0.979, GSE34619 1.0, GSE13083 1.0)
within tolerance
C1a
Reported
GSE100843 training cohort within 0.88-1.00
Reproduced
ROC-AUC 0.9993 (36 N / 40 BE, matches authors notebook split [36,40])
exact
C2
Reported
SPINK7 cluster 220 genes down + SLC44A4 cluster 24 genes up
Reproduced
220 down / 24 up recovered verbatim from authors be-eac.ipynb cell-14
exact
C3
Reported
EAC map IL10RA(30)+LILRB3(31) up, LNX1(471) down; EAC-from-BE classification
Reproduced
up-cluster gene lists captured verbatim; classification AUC not computed (partial)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 84/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

The primary NE-vs-BE BoNE classifier reproduces essentially 1:1: 7 out-of-sample validation cohorts give ROC-AUC 0.876-1.00 vs the reported 0.88-1.00 (the only sub-bound value, GSE39491 0.8763, rounds to 0.88), and cluster sizes (220 down / 24 up) match exactly. All values are derivable from the authors' shipped be-eac.ipynb clusters + public GEO data — no fabrication. Caveats are scope-side, not defects: the GSE100843 training AUC (0.9993) is in-sample because the cluster-generating Boolean network wasn't re-run (Hegemon DBs not redistributed), and the secondary C3 EAC classifier was only partially reproduced. Overall a clean, well-specified reproduction.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

212.9 k
tokens (I/O) · 18.7 M incl. cache
60 min
runtime · 0.03 CPU-h
1.4 GB
peak RAM
5
HPC jobs
hummel
machine