The mouse gastric surface epithelial cell and its response to early Helicobacter pylori infection.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough -> EXACT 1:1 reproduction of the core pipeline result (RNA-seq differential expression). The repo ships the analysis R scripts AND the nf-core/rnaseq salmon quant.sf counts AND the authors' own outputs. I re-implemented the DE portion of script 02 faithfully (tximport -> merge duplicated gene symbols -> DESeq2 ~condition -> results(Infected vs Non_Infected, alpha=0.05) -> IHW) and ran it on «our HPC» on the shipped counts. All 12 pinned numeric claims matched exactly (23678 detected genes, 214 significant DE genes, 33 up / 10 down / 43 |log2FC|>2, and per-gene log2FC for Nkx6-3 -7.02, Dpp4 -4.57, Krt7 -4.40, Zbp1 3.96, Bst2 1.57, Cmpk2 5.17, Gm23935 0.9). Gene-by-gene vs the authors' shipped All_Detected_Genes.tsv: log2FC Pearson r=1.0, max abs diff=0 (median 1.9e-14), adj_p r=1.0 across all 23678 genes — identical to floating-point precision despite using a newer stack (DESeq2 1.50.2 / R 4.5.3 vs authors' 1.42.1 / 4.3.2). No fabrication signal in the DE results. NOT attempted (the hard ~20%, documented in scope.md): SetRank GSEA '119 gene sets' (needs live version-floating Reactome/KEGG/GO/biomaRt/STRING DBs + interactive Cytoscape; README warns it varies between runs), glycomics/glycowork scripts 06-08 (separate mass-spec modality), and histology HAI + immunofluorescence (wet-lab/manual, not pipeline-derived). Raw reads (ENA PRJEB70775) were not re-quantified because the authors' salmon counts are shipped in the repo.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 100assessed: 2026-06-16 ⛓ b869df723551
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThis study aimed to characterize the normal gene expression and mucus glycosylation of gastric surface mucus-producing epithelial cells (SMCs) and to determine how early Helicobacter pylori colonization affects these parameters.
- ★ LCM followed by RNA-Seq is a feasible methodology for characterizing gastric SMCs and their host response to H. pylori infection in vivo method
- ★ SMCs are characterized by high expression of secreted mucus proteins (Tff1, Gkn1, Gkn2, Psca, Muc5ac), mitoribosome RNA, and cytoskeleton proteins finding
- ★ Gastric mucin glycans are large, complex, heavily fucosylated and dense with H-antigen motifs, formed via two main glycosylation pathways corroborated by glycosyltransferase expression finding
- ★ H. pylori infection down-regulates genes required for protein synthesis and oxidative phosphorylation in SMCs finding
- ★ Most up-regulated genes in infected mice are interferon-stimulated genes or genes able to induce interferon production finding
- ★ Depletion of Nkx6-3 in infected mice indicates initiation of a pre-cancerous cascade mechanism
- Mucin glycosylation was consistent between H. pylori-infected and sham control mice finding
- Integrating glycomics with glycosyltransferase transcriptomics identifies key mucin glycan biosynthesis pathways method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RNA-Seq (LCM-RNA-Seq) | C57BL/6Ntac mouse gastric corpus surface mucus epithelial cells (SMCs) | H. pylori SS1 infection vs sham (PBS) control | differential gene expression / transcriptome | Laser capture microdissection (PALM MicroBeam, Carl Zeiss) |
| Mass spectrometry glycomics | Mouse gastric mucins | H. pylori SS1 infection vs sham control | mucin O-glycan structures and motifs (fucosylation, H-antigen) | — |
| Immunofluorescence (IF) | FFPE mouse gastric corpus tissue (SMCs) | H. pylori infection vs sham | protein-level validation; mean signal intensity of RPL19, LGALS3BP, COX6C | Fiji; Alexa Fluor 594 secondary antibody |
| Fluorescence in situ hybridization (FISH) | Mouse gastric tissue (pits) | H. pylori infection | Helicobacter bacterial density (counts per 40x field) | Nikon Eclipse 90i; EUB338-Cy3.5 and Helicobacter-specific Alexa488 probes |
| Histology (H&E, Histological Activity Index) | FFPE mouse gastric corpus tissue | H. pylori infection vs sham | gastritis score (HAI) | — |
- ▼ Genes required for protein synthesis and oxidative phosphorylation were down-regulated in infected mice SMCs
- ▲ Most up-regulated genes were interferon-stimulated genes or inducers of interferon production
- ▼ Nkx6-3 was depleted in infected mice, indicating a pre-cancerous cascade
- – Mucin glycans were large, complex, heavily fucosylated and dense with H-antigen motifs via two main H-antigen pathways
- – Glycosylation was consistent between H. pylori-infected and sham control mice
- ▲ SMCs showed high expression of Tff1, Gkn1, Gkn2, Psca, Muc5ac, mitoribosome RNA and cytoskeleton proteins
- count 3.5 ×10^7 CFU/mouse (two 200 µL doses) (H. pylori SS1 infection dose by oral gavage)
- count 15 ng purified RNA from 12 LCM sections (pilot LCM RNA yield for sequencing)
- count 9 of 10 samples successfully sequenced and passed MultiQC (RNA-Seq cohort sequencing success)
- count 8 infected + 8 sham control mice (glycan cohort group sizes)
- count over 1000 differentially-expressed genes (prior LCM microarray study cited (parietal, chief, mucus cells))
- other OD600 of 1.0 (H. pylori suspension optical density for infection)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study used a two-group experimental design (H. pylori SS1-infected vs PBS sham control male C57BL/6Ntac mice) across an RNA-Seq cohort (n=5 per group; 9 of 10 passed QC) and a glycan cohort (n=4 per group per cohort). Differential gene expression was analysed with DESeq2 with Independent Hypothesis Weighting (IHW) for multiple-testing correction; gene set enrichment was performed with SetRank against Reactome and KEGG; histological and immunofluorescence comparisons used non-parametric tests (Mann-Whitney U, Spearman Rank Correlation) with Benjamini-Hochberg FDR adjustment applied to p-values, all in R. Glycan data were analysed with the Python package Glycowork.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| DESeq2 negative binomial Wald test | Differential gene expression between H. pylori-infected and sham control SMCs (RNA-Seq cohort) | 9 samples (10 collected; 1 failed QC; exact per-group split after failure not stated) | not stated |
| Independent Hypothesis Weighting (IHW) | Multiple-testing correction for RNA-Seq differential expression p-values (applied within DESeq2 workflow) | — | not stated |
| SetRank | Gene set enrichment analysis against Reactome and KEGG pathway databases | — | not stated |
| Mann-Whitney U | Differences in histological assay results (HAI scores, H. pylori FISH density) between infected and sham control groups | not stated per comparison; both RNA-Seq cohort (n=10) and glycan cohorts (n=16 total) were assessed histologically | not stated |
| Spearman Rank Correlation | Correlations between immunofluorescence signal intensity and RNA-Seq gene expression levels for validated targets (RPL19, LGALS3BP, COX6C) | not stated explicitly; likely n=9 RNA-Seq samples | not stated |
-
DESeq2 was used for differential expression with n=9 total samples across two unequal groups after one QC exclusion↳ Could also: edgeR (quasi-likelihood F-test) or limma-voom with quality weights could also be applied to the same count matrix — At very small n, benchmarking studies suggest limma-voom with quality weights and edgeR QL maintain good type-I error control; running a second tool and reporting concordance of top hits is a common practice for increasing confidence in findings from small RNA-Seq experiments
-
IHW was used for multiple-testing correction of RNA-Seq differential expression p-values↳ Could also: Standard Benjamini-Hochberg FDR (already used for other comparisons in the paper) could also be applied directly to DESeq2 p-values — BH-FDR is universally familiar and directly comparable across studies; IHW can gain power by weighting hypotheses by a covariate (typically mean normalised count), but requires explicit reporting of the covariate used to be fully reproducible
-
SetRank was used for gene set enrichment against Reactome and KEGG↳ Could also: Pre-ranked GSEA via fgsea, or camera/ROAST from limma, could also test pathway enrichment using the ranked DESeq2 Wald statistics — fgsea and camera are widely adopted, integrate directly with DESeq2 ranked output, and are supported by extensive benchmarking literature; reporting alongside SetRank would facilitate cross-study comparison
-
Mann-Whitney U was used for histological group comparisons at small n per group↳ Could also: A Welch's unpaired t-test could also be used if the ordinal HAI scores are treated as approximately continuous and roughly normally distributed — Mann-Whitney U is a conservative, distribution-free choice appropriate for ordinal data and small n; the t-test assumes approximate normality but may have marginally greater power when that assumption holds — the choice between them is a common methodological decision point for scored histological data
-
Spearman Rank Correlation was used to relate IF mean signal intensity to RNA-Seq expression across samples↳ Could also: Pearson correlation could also be used if both variables approximate a bivariate normal distribution — Spearman is robust to outliers and does not assume normality, making it well-suited for small n and potentially skewed fluorescence intensity measurements; Pearson is an equally conventional alternative when distributional assumptions can be checked
-
RNA-Seq group size was set at n=5 per group, justified narratively by pilot histological data and a published rule of thumb↳ Could also: A formal a priori power calculation using estimated effect sizes and variance from the pilot data or published gastric RNA-Seq datasets could also have been used — Explicit power analysis makes the assumptions underlying sample size choice transparent and reproducible; for RNA-Seq, power is gene-specific and also influenced by sequencing depth, so simulation-based approaches (e.g. RnaSeqSampleSize or powsimR) can complement empirical justification
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41824631
Paper: Erhardsson et al. 2026, Virulence. "The mouse gastric surface epithelial cell and its response to early Helicobacter pylori infection." DOI 10.1080/21505594.2026.2645859 · PMC13020875
Repo: https://github.com/mattias-erhardsson/lmpc-infection-rnaseq (MIT,
public, default branch main). The repo ships BOTH the analysis code (R scripts
01–08 + one Jupyter notebook) AND the inputs needed for the RNA-seq analysis:
the nf-core/rnaseq STAR+salmon quantifications (P26010/01-RNA-Results/ star_salmon/<sample>/quant.sf), salmon_tx2gene.tsv, and the sample sheet —
plus the authors' own shipped outputs (R_output_files/Tables/All_Detected_Genes.tsv,
etc.). The heavy compute (read alignment / salmon quant) was already done by
the authors; the salmon counts are in the repo. Raw reads: ENA PRJEB70775.
In scope (pipeline-derived, attempted)
The core computational result is script 02 — DESeq2 differential expression of laser-microdissected surface-mucus-cell RNA-seq, Infected vs Non-Infected, with IHW p-value adjustment, run on the shipped salmon quant counts. This is fully specified and deterministic given the counts + package versions. Targets:
| id | reported value | paper location |
|---|---|---|
| detected_genes | 23,678 genes detected in tissue | Results ("Early H. pylori infection…"), Suppl. File 7 |
| sig_genes | 214 significant DE genes (adj p < 0.05) | Results, Fig 11D, Suppl. File 8 |
| up_l2fc_gt2 | 33 up-regulated, log2FC > 2 | Results |
| down_l2fc_lt2 | 10 down-regulated, log2FC < −2 | Results |
| abs_l2fc_gt2 | 43 with | log2FC |
| nkx63_l2fc | Nkx6-3 log2FC = −7.02 (most extreme) | Table 6 |
| dpp4_l2fc | Dpp4 log2FC = −4.57 | Table 6 |
| krt7_l2fc | Krt7 log2FC = −4.40 | Table 6 |
Method of reproduction: a faithful, self-contained re-implementation of the
DE-relevant portion of 02_…-pre-setrank.R (same import → tximport → merge
duplicated gene symbols → DESeqDataSetFromMatrix(condition) → DESeq →
results(contrast=Infected vs Non_Infected, alpha=0.05) → IHW(pvaluebaseMean)),
run on «our HPC» under a conda env pinned to the renv.lock major versions (R 4.3,
DESeq2 1.42, tximport, IHW). Reproduced log2FC / adj-p are also compared
gene-by-gene against the authors' own shipped All_Detected_Genes.tsv.
Out of scope / not attempted (the hard 20%), with reasons
- SetRank GSEA (119 significant gene sets) — scripts 03/04. Requires LIVE, version-floating external databases (Reactome.db, KEGGREST release 107, GO.db, biomaRt Ensembl 109, STRING v12) and an interactive Cytoscape GUI session (script 04 says "launch Cytoscape before running"). The repo README itself warns these results vary between runs because linked databases update. Not reliably reproducible headless → skipped (documented, not failed).
- Glycomics / glycowork (scripts 06–08) — a separate mass-spectrometry data modality (a third-party tool, glycowork, on the glycan data). Distinct effort from the RNA-seq pipeline; not attempted here.
- Histology HAI scoring (script 05) and immunofluorescence validation (Suppl. File 11) — wet-lab / manual blinded scoring, not pipeline-derived.
Honesty note
The README explicitly states re-runs may differ slightly from the article due to un-pinned dependency/database updates between runs. The DESeq2/IHW counts depend only on the (pinned) counts + DESeq2/IHW versions and should be essentially 1:1; small integer drift (±a few genes) across DESeq2 patch versions is expected and is reported honestly rather than hidden.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is an exact 1:1 reproduction of the paper's core pipeline result. Re-implementing the DE portion of script 02 on the authors' shipped salmon counts matched every one of the 12 reported values exactly, and a whole-table cross-check against Suppl. File 7 gave log2FC Pearson r=1.0 with max abs diff=0 across all 23678 genes — even under a newer DESeq2/R stack, ruling out fabrication. The only unreproduced parts (SetRank GSEA, glycomics, histology) are explicitly non-pipeline modalities that the authors themselves flag as run-to-run variable or wet-lab-derived, so they are a scope limitation on our side, not an authors' defect. Deviation: none; severity: negligible.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.