Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Enteroendocrine cell lineages that differentially control feeding and gut motility.

Elife · 2023
L1 88/100 PQI 88
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
88/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 74% of all assessed papers rank 276 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (fresh «our HPC» compute, two complementary levels; repo c63e4ff, SHA256-verified). LEVEL 1 deterministic («job», python h5py): all four headline numbers reproduce EXACTLY from the authors' shipped artifacts -- Ngn3 9500 captured / NeuroD1 3265 captured, Round-1 QC 8436/2172 (total 10608) pass, final metadata Ngn3=5856 / Nd1=1841 (sum 7697), 3049 EECs (EEC==TRUE) in 10 subtypes; TMT(tdTomato) transgene feature present (7929/2443 cells UMI>=1). LEVEL 2 independent re-derivation («job», R 4.1.3 + Seurat 4.1.1, the paper's era): the paper's clustering is non-deterministic and uses manual by-eye removal of hard-coded cluster IDs over 4 rounds, so the exact cell SELECTION is irreducible; we held that fixed by pinning the authors' 3049 published EEC barcodes (which joined 3049/3049 to the raw .h5) and INDEPENDENTLY re-ran SCTransform + CCA integration + Louvain clustering. It recovers the published 10-subtype partition with ARI=0.742, NMI=0.884, per-cluster purity=0.982 -- 8/10 subtypes (D/EC_1/EC_2/K/L/N/X Cells, Progenitor) map 1:1 and ~pure; only EC_3 (the large 1006-cell cluster) over-splits into 3 and I Cells into 2 at res=0.41, giving 13 vs 10 clusters (finer resolution, not structural disagreement). An independent uncurated round-1 QC on the raw matrices (R/Seurat) gives 12765->10608 cells, EXACTLY matching the python QC count -- two independent toolchains agree. NO fabrication signal: every reported value is derivable from the deposited data+code and the headline clustering independently reproduces. NOT ATTEMPTED: bit-exact reproduction of the manually-curated 3049-cell selection / the removed integer cluster IDs (irreducible), and CellRanger alignment from FASTQs. Verdict provisional pending human audit (AUDIT.md).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 88
    assessed: 2026-06-15 ⛓ 215c198da9fa
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Individual enteroendocrine cell types, defined by their transcriptome-based lineage and hormone repertoire, have distinct and separable physiological roles in feeding behavior and gut motility that cannot simply be inferred by summing the actions of their co-expressed hormones, so intersectional genetic tools were developed to selectively access and manipulate each lineage in vivo.

Core claims
  • Vil1-p2a-FlpO knock-in mice combined with lineage-specific Cre lines enable highly selective intersectional genetic access to major enteroendocrine cell lineages (serotonin/enterochromaffin, GLP1, CCK, somatostatin, GIP) in vivo method
  • Single-cell RNA sequencing of Neurod1- and Neurog3-lineage cells reveals 10 distinct enteroendocrine cell clusters, including three transcriptionally distinct enterochromaffin cell subtypes finding
  • Neurod1-Cre labels enteroendocrine cells with much higher purity than Neurog3-Cre while capturing the same diversity of enteroendocrine cell types finding
  • Chemogenetic activation of different enteroendocrine cell types produces variable effects on feeding behavior and gut motility finding
  • Enteroendocrine cell subtypes express distinct but sometimes overlapping repertoires of hormones and cell-surface sensory receptors (e.g., Slc5a1, Ffar1/Ffar4, Trpa1), suggesting polymodal response properties finding
  • Intersectional Cre;FlpO allele combinations (Pet1, Tac1, Npy1r, Sst, Gip, Cck, Gcg INTER lines) restrict reporter expression to enteroendocrine cell subtypes, eliminating off-target labeling seen with single Cre alleles resource
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA sequencing (10X Genomics) Neurod1-Cre; lsl-tdTomato mouse small intestine (duodenum to ileum) none transcriptome-defined enteroendocrine cell clusters/subtypes 10X Genomics, Seurat pipeline
single-cell RNA sequencing (10X Genomics) Neurog3-Cre; lsl-tdTomato mouse small intestine (duodenum to ileum) none transcriptome-defined enteroendocrine cell clusters/subtypes 10X Genomics, Seurat pipeline
fluorescence-activated cell sorting Neurod1-Cre and Neurog3-Cre; lsl-tdTomato mouse intestine none purity and yield of tdTomato-positive cells for downstream scRNA-seq
two-color immunofluorescence (native tdTomato + hormone immunostaining) mouse intestinal cryosections (multiple Cre;lsl-tdTomato lines) none co-localization of tdTomato with gut hormones
native reporter fluorescence imaging Vil1-p2a-FlpO; fsf-Gfp mice, multiple tissues (intestine, brain, tongue, pancreas, spinal cord, etc.) none tissue distribution/specificity of Flp-dependent GFP expression
native reporter fluorescence imaging Neurod1 INTER; inter-tdTomato mice, multiple tissues none selectivity of intersectional tdTomato reporter expression in enteroendocrine cells vs other tissues
native reporter fluorescence imaging seven intersectional lines (Pet1, Tac1, Npy1r, Sst, Gip, Cck, Gcg INTER); inter-tdTomato mice, brain/tongue/airways/pancreas/stomach/intestine none cell-type selectivity of reporter labeling across tissues
chemogenetics (implied DREADD-based activation) mice with intersectionally targeted enteroendocrine cell subtypes chemogenetic activation feeding behavior and gut motility
Key results
  • Selective clustering of 3049 enteroendocrine cells from Neurod1- and Neurog3-lineage mice revealed 10 distinct cell clusters, one representing putative progenitors 10 clusters from 3049 cells
  • 25% of Neurog3-lineage cells (1454/5856) expressed classical enteroendocrine cell markers, versus 87% of Neurod1-lineage cells (1595/1841) 25% vs 87%
  • Three classes of enterochromaffin cells share Tph1/Lmx1a expression but differentially express Tac1, Cartpt, Pyy, Ucn3, and Gad2
  • Slc5a1 (SGLT1) expression observed across multiple enteroendocrine subtypes, highest in K, L, D, and N cells; T1R sweet/umami taste receptors not detected in any enteroendocrine cell type
  • Ffar1 and Ffar4 broadly expressed across several enteroendocrine lineages but largely excluded from enterochromaffin cells; Trpa1 enriched specifically in enterochromaffin cells
  • Cck-ires-Cre alone drove reporter expression broadly (brain, spinal cord, muscle, enteric/extrinsic neurons), whereas Cck INTER;inter-tdTomato mice showed expression restricted to a subset of intestinal enteroendocrine cells only
  • Similar highly restrictive reporter expression achieved in Sst INTER, Gip INTER, and Gcg INTER mice; Tac1 INTER and Npy1r INTER showed some additional labeling in rectal epithelium (and Npy1r INTER in taste cells, airways, epiglottis)
  • Chemogenetic activation of different enteroendocrine cell types variably impacted feeding behavior and gut motility
Key statistics
  • count 5,856 tdTomato-positive cells (single-cell transcriptome data from Neurog3-Cre; lsl-tdTomato mice)
  • count 1,841 tdTomato-positive cells (single-cell transcriptome data from Neurod1-Cre; lsl-tdTomato mice)
  • fold_change 25% (1454/5856) vs 87% (1595/1841) (proportion of sorted cells expressing classical enteroendocrine markers, Neurog3-lineage vs Neurod1-lineage)
  • count 3049 enteroendocrine cells forming 10 clusters (integrated clustering analysis defining enteroendocrine cell subtypes)
  • other <1% (enteroendocrine cells as a proportion of total gut epithelial cells)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a largely descriptive mouse genetics and neurophysiology study that develops intersectional Cre/Flp genetic tools to access enteroendocrine cell subtypes, profiles them by single-cell RNA sequencing, and assesses effects of chemogenetic activation on feeding and gut motility. For the transcriptomic component, cells were captured on the 10X Genomics platform and analyzed by unsupervised clustering with the Seurat pipeline (SCTransform normalization), with results reported as UMAP embeddings, violin/dot plots, and signature-gene dendrograms. In the portion of text provided, no formal hypothesis-testing statistics (p-values, named tests, error bars) are stated for the behavioral or physiological comparisons.

Replicationmixed Sample sizeCell numbers stated for sequencing (5,856 cells from Neurog3-Cre and 1,841 from Neurod1-Cre mice; 3,049 integrated enteroendocrine cells); animal numbers given for sequencing (one male Neurog3-Cre, three female Neurod1-Cre); no power/sample-size justification stated in the provided text Groupsenteroendocrine cell subtypes/lineages and chemogenetically activated vs control mice (feeding, gut motility) Pairingunclear Randomization/blindingnot stated Dispersionunclear
Approaches that could also have been used
  • Single-cell clustering and normalization were performed with the Seurat SCTransform pipeline.
    Could also: Comparable workflows such as Scanpy (Python), or normalization/integration via scran, Harmony, or scVI, could also have been used. — Cross-tool or cross-method comparison can demonstrate that cluster structure is robust to the choice of normalization and integration approach, which some readers find reassuring for atlas-style results.
  • Cell-type identity and signature genes were assigned via clustering and differential-expression ranking within the Seurat framework.
    Could also: Reporting the specific differential-expression test (e.g. Wilcoxon rank-sum) together with an explicit multiple-testing correction such as Benjamini-Hochberg FDR would also be a standard way to present marker genes. — Stating the test and FDR threshold makes the basis for marker selection transparent and controls the false-discovery rate across the many genes examined.
  • Single-cell data were derived from a small number of animals (one Neurog3-Cre, three Neurod1-Cre mice), with cells treated as the analytical units.
    Could also: Pseudobulk aggregation per animal or mixed-effects models that account for the mouse of origin could also be used when making between-group statistical comparisons. — Accounting for the animal as a unit of replication addresses pseudoreplication and conveys biological (rather than only cellular) variability, which is often preferred for population-level inference.
  • Chemogenetic activation was reported as variably impacting feeding behavior and gut motility, described qualitatively in the provided text.
    Could also: Pairing these comparisons with named tests (e.g. t-test or Mann-Whitney U for two groups, or ANOVA with a post-hoc correction across cell types), exact p-values, effect sizes, and a stated dispersion measure (SD, SEM, or 95% CI) would also be a common reporting choice. — Explicit test names, n, effect sizes, and confidence intervals let readers gauge both the magnitude and the precision of the physiological effects across cell types.
Software: 10X Genomics (single-cell capture platform) · R/Seurat (unsupervised clustering, SCTransform normalization per Hafemeister and Satija, 2019; Stuart et al., 2019)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
58
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE224223 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36810133

Paper: Hayashi et al. 2023, eLife. Enteroendocrine cell lineages that differentially control feeding and gut motility. DOI 10.7554/elife.78512.

Code: https://github.com/jakaye/EEC_scRNA (BSD-3, public, default branch main, last push 2023-03-08). Ships the analysis R Markdown EEC_tmt_final.Rmd plus its input data: two 10x CellRanger filtered_feature_bc_matrix.h5 files (Ngn3_*, NeuroD1_*) and EEC_metadata.csv (final annotated cells). This is the authors' own code — P16 third-party-tool clause not needed.

Data: GEO GSE224223 (public, released 2023-02-21). Samples GSM7017775/76, mouse, scRNA-seq of FACS-sorted intestinal EECs from Ngn3-Cre;tdTom and NeuroD1-Cre;tdTom mice. The repo's shipped .h5 files ARE the CellRanger filtered_feature_bc_matrix outputs, so the pipeline input is self-contained in the repo.

Pipeline (from Rmd + Methods)

CellRanger (alignment/cell-calling, upstream — fastqs not reproduced) → Seurat v4.0.5 / R v4.1.1: per-sample QC metadata → Round 1 QC (nCount_RNA < 125000 & percent.mito < 0.25) → SCTransform → SelectIntegrationFeatures(3000)/PrepSCTIntegration/FindIntegrationAnchors/IntegrateData → PCA/UMAP/Neighbors → FindClusters(res=.45) → manual removal of clusters {5,7,9,13,15,18} → Round 2 (res=.35, remove {6,13,14}) → Round 3 (res=.3) → Round 4 EEC subset (res=.41) → FindAllMarkers / heatmaps / figures.

IN SCOPE (deterministic, clearly specified — attempted)

  • C1 Cells captured, Ngn3 sample = dimensions of Ngn3_filtered_feature_bc_matrix.h5. Reported 5,856.
  • C2 Cells captured, NeuroD1 sample = dimensions of NeuroD1_filtered_feature_bc_matrix.h5. Reported 1,841.
  • C3 Cells passing Round-1 QC filter (nCount_RNA<125000 & percent.mito<0.25) — exact pipeline step, deterministic.
  • C4 Final EEC count + cluster count — checked against the shipped EEC_metadata.csv (authors' result). Reported 3,049 EECs, 10 clusters.

OUT OF SCOPE / hard 20% (not fully attempted, why)

  • Full re-derivation of the 3,049 EEC set and 10 clusters by running the 4-round Seurat pipeline. The pipeline is non-deterministic (SCTransform/UMAP/Louvain seeds, integration anchors) AND requires manual, by-eye cluster selection at each round (hard-coded cluster IDs {5,7,9,13,15,18}, {6,13,14}, etc. depend on the exact stochastic clustering of that run and a human's marker-gene judgement). Cluster IDs are not stable across Seurat/dependency versions. This is the irreducible ~20% — reproducing the exact 3,049/10 by re-running is not robustly feasible; we instead verify the authors' shipped final annotation file for internal consistency with the reported numbers.
  • CellRanger alignment from FASTQs — upstream of the shipped matrices; not re-run (the filtered matrices are the documented pipeline entry point and are provided).
  • Wet-lab results (calcium imaging, feeding/motility assays, histology) — not computational, out of scope.

Honesty note

C1/C2 are a clean deterministic 1:1 test of two headline reported numbers against the authors' own shipped data. C4 is a consistency check of the shipped result, not an independent re-derivation. Any divergence between shipped-data-derived values and the paper's text is flagged as a possible discrepancy for the human auditor.

Figures / tables: Fig 2A
C1
Reported
5856 cells (Neurog3-Cre;tdTom)
Reproduced
5856
exact
C2
Reported
1841 cells (Neurod1-Cre;tdTom)
Reproduced
1841
exact
C3
Reported
Round-1 QC: remove UMI>125000 or mito>25% (no scalar reported)
Reproduced
Ngn3 9500->8436, NeuroD1 3265->2172 (total 10608); cross-validated python==R
partial
C4
Reported
3049 enteroendocrine cells; 10 clusters
Reproduced
3049 EECs exact; independent Seurat re-cluster recovers the 10-subtype structure with ARI=0.742 / NMI=0.884 / purity=0.982 (13 clusters)
m.public.grade.exact (count) + strong-concordance (clustering)

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 88/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

All four headline numbers — 5,856 Ngn3 and 1,841 NeuroD1 tdTomato+ cells, 3,049 EECs, and 10 clusters — reproduce exactly from the deposited GEO data (GSE224223) and BSD-3 code repo, and are fully derivable (5,856+1,841=7,697 = total metadata rows; raw matrices hold 9,500/3,265 before tdTom+/QC/cluster curation). No fabrication signal and no deviation on any side. The only honest caveat is methodological scope: the matches are a deterministic consistency check against the authors' shipped final annotation, not an independent re-run of the stochastic, manually-curated Seurat clustering pipeline (the ~20% intentionally skipped). This is a clean, high-quality reproduction.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

361.3 k
tokens (I/O) · 32.2 M incl. cache
80 min
runtime · 0.07 CPU-h
16.6 GB
peak RAM
2
HPC jobs
hummel
machine