Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

NK2R control of energy expenditure and feeding to treat metabolic diseases.

Nature · 2024
L1 87/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
87/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 72% of all assessed papers rank 301 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED. PMID 39537932 (Sass et al., Nature 2024) is predominantly wet-lab; its single pipeline-derived component is a dorsal-vagal-complex snRNA-seq atlas (Fig. 4k-l). BOTH harvested pointers were text-mining false positives and were corrected: code mmkim1210/ggLD -> perslab/Sass-2024; data GSE166649 -> GSE276735 (GSE166649 is actually the Ludwig-2021 reference atlas used for label transfer). Compute ran on «our HPC» («job», node n094, Seurat 4.3.0.1 = paper's v4.3.0). Deposit verification reproduces the headline numbers EXACTLY: 23,664 nuclei, 25 neuronal populations, >=1000-gene retention (min nFeature 1002), even treatment balance across clusters. Crucially, an INDEPENDENT recomputation of the Fig. 4l derived result (authors' exact 17-gene IEG module score + lm/emmeans treatment contrast on the deposited counts) reproduces the paper's central snRNA-seq conclusion 1:1: Glu3 is the most EB1002-responsive neuronal population (padj 5.5e-11, the only significant cluster). scDist (the paper's concordant second method) is building. NOT attempted (documented): upstream cellranger/cellbender (raw FASTQ in SRA, not the deposit; needs GPU), label transfer (private Ludwig-2021 reference object), and all in-vivo/wet-lab/NHP/imaging results + the genetics-portal framing (no deposited pipeline code).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 87
    assessed: 2026-06-19 ⛓ b6746b6ab07c
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether activation of a single receptor, NK2R, identified through human genetic links to glucose control and obesity, can both centrally suppress appetite and peripherally increase energy expenditure to treat obesity and type 2 diabetes.

Core claims
  • NK2R activation is sufficient to suppress appetite centrally and increase energy expenditure peripherally finding
  • NK2R shows the strongest HbA1c association among 381 non-odorant GPCR loci and is genetically linked to obesity and glucose control finding
  • Selective, long-acting NK2R agonists (e.g., EB0014) were developed with potential for once-weekly human dosing method
  • NK2R agonists induce weight loss in mice via increased energy expenditure and non-aversive appetite suppression that circumvents canonical leptin signalling finding
  • NK2R agonism acutely enhances insulin sensitization, shown by hyperinsulinaemic-euglycaemic clamp finding
  • In diabetic, obese macaques, NK2R activation decreases body weight, blood glucose, triglycerides and cholesterol, and ameliorates insulin resistance finding
  • NK2R missense variants I23T and R323H reduce Gq signalling capacity and associate with increased HbA1c levels finding
  • The NK2R intronic variant rs7911347 is an eQTL regulating NK2R expression across brain, adipose, muscle and macrophage tissue and is a fine-mapped candidate causal HbA1c variant finding
Experimental setups
Assay System Perturbation Readout Platform
GWAS meta-analysis / LD clumping and fine-mapping (CARMA) human, European-ancestry cohort (n=438,069) none HbA1c association at HKDC1-HK1-NK2R-TSPAN15 locus
Transcriptome-wide association study (TWAS, PrediXcan) human brain tissue (nucleus accumbens, ACB) none association of NK2R/HK1/TSPAN15 expression with HbA1c, with/without BMI adjustment PrediXcan
Receptor signalling assay (Gq activation) in vitro, NK2R missense variants (I23T, R323H, V54I, A161T) point mutation Gq signalling capacity
Genetic association study human, Greenlandic cohort none obesity-related anthropometric traits and NK2R expression by rs139900276 genotype
Pharmacokinetics and metabolic phenotyping (indirect calorimetry, food intake, body weight) diet-induced obese (DIO) mice NKA, 1 mg/kg twice daily s.c. injection oxygen consumption, food intake, body weight, white adipose tissue weight
Insulin tolerance test DIO mice NKA, twice daily s.c. injection for 12 days insulin sensitivity
Pharmacokinetics and dose-response metabolic phenotyping (indirect calorimetry, body composition) mice EB0014, daily s.c. injection at 0.1/0.3/1 mg/kg for 12 days oxygen consumption, food intake, weight loss, body composition
Metabolic and cardiometabolic phenotyping diabetic, obese macaques NK2R agonist treatment body weight, blood glucose, triglycerides, cholesterol, insulin resistance
Key results
  • NK2R locus shows the most significant HbA1c association among 381 non-odorant GPCR loci in T2D-KP
  • Missense variants I23T and R323H reduce NK2R Gq signalling and associate with increased HbA1c
  • Increased NK2R expression in nucleus accumbens associated with decreased HbA1c, with and without BMI adjustment effect size -0.147 (unadjusted)
  • NKA treatment increased oxygen consumption and decreased food intake and body weight in DIO mice
  • EB0014 dose-dependently decreased food intake and body weight and altered body composition
  • NK2R agonism improved insulin tolerance test outcome in DIO mice
  • NK2R agonism decreased body weight, blood glucose, triglycerides and cholesterol and improved insulin resistance in diabetic obese macaques
  • rs7911347 is significantly associated with NK2R expression across brain, adipose, macrophage and skeletal muscle tissues
Key statistics
  • pvalue P = 5.57322 × 10^-48 (TWAS effect size -0.147 for NK2R expression (ACB) vs HbA1c, unadjusted)
  • pvalue P = 2.99 × 10^-8 (rs7911347 (NK2R intron) HbA1c association, beta = 0.0105, PIP = 0.9965)
  • correlation 1.11, P = 4.0 × 10^-63 (eQTL association of rs7911347-A with NK2R expression in muscle (FUSION))
  • pvalue P = 6.32 × 10^-101 (rs2015803 HK1 intron variant HbA1c association, beta = 0.0455)
  • count 978 genome-wide significant variants, 16 independent lead variants (HbA1c GWAS in HKDC1-HK1-NK2R-TSPAN15 region)
  • count n = 6 (vehicle), n = 7 (NKA) (DIO mice in vivo NKA study (oxygen consumption, food intake, body weight, WAT weight))
  • count n = 10 (vehicle, 0.1 mg/kg), n = 9 (0.3 mg/kg), n = 8 (1 mg/kg) (EB0014 dose-response study after 12 days of daily injections)
  • other MAF = 23.6% (I23T), 0.10% (R323H), 0.025% (V54I), 0.102% (A161T) (minor allele frequencies of NK2R missense variants reducing signalling capacity)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper combines human genetic association analyses (regional locus association testing, LD clumping, fine-mapping with CARMA, and transcriptome-wide association analysis with PrediXcan) with preclinical physiology experiments in mice and nonhuman primates. Genetic association tests are reported as two-sided with exact P values and effect sizes (beta), generally without an additional multiple-testing correction beyond the genome-wide significance threshold used for initial variant discovery. Animal experiment data are summarized as mean ± s.e.m. or as box plots with median and Tukey's whiskers, with significance denoted by threshold asterisks, and the specific statistical tests underlying those comparisons are stated to be detailed in the Supplementary Information and Source Data rather than in the main text excerpt provided.

Replicationbiological Sample sizeGroup sizes given per figure/table (e.g., n=6–7 mice per group for most physiology studies, n=2 per genetic variant for Gq signalling assays, n=3 mice for pharmacokinetics, cohorts ranging up to hundreds of thousands of individuals for genetic association analyses); no formal power calculation is described in the text provided. GroupsVehicle vs NK2R agonist (NKA or EB0014) treated animals; genetic variant/genotype carriers vs non-carriers in human cohorts; gene expression levels vs HbA1c in TWAS Pairingunclear Randomization/blindingnot stated Dispersionmixed Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionNone applied for the regional variant association/LD-clumping and fine-mapping tests (explicitly stated as 'two-sided without multiple corrections'); nominal P values without genome-wide FDR correction are reported for the 3-gene TWAS specifically because only three genes were tested
Statistical tests used
Test Applied to n Assumptions
Two-sided regional variant association tests / LD clumping HKDC1–HK1–NK2R–TSPAN15 locus, Table 1 European-ancestry meta-analysis of 438,069 individuals not stated
Bayesian fine-mapping (CARMA) for posterior inclusion probability Candidate causal variant identification, Table 2 same meta-analysis summary statistics not stated
Transcriptome-wide association study (PrediXcan), two-sided z-test NK2R/HK1/TSPAN15 expression vs HbA1c, Fig. 1c and Table 4 not explicitly stated in text provided not stated
eQTL association tests rs7911347 tissue expression associations, Table 3 not stated not stated
Group comparisons of physiological outcomes (test name not specified; asterisk significance thresholds shown) Vehicle vs NKA/EB0014 treated mice: oxygen consumption, food intake, body weight, insulin tolerance test, Fig. 1g–r n=6–10 mice per group as stated in figure legend not stated (detailed statistics deferred to Supplementary Information)
Genotype-stratified comparison (test name not specified) NK2R expression by rs139900276 genotype, Fig. 1e not stated in text provided not stated
Approaches that could also have been used
  • Genetic association tests across the HKDC1–HK1–NK2R–TSPAN15 region (Tables 1–2) are explicitly reported without a multiple-testing correction beyond the genome-wide threshold used for initial variant discovery.
    Could also: A locus-wide or study-wide correction such as Bonferroni adjustment for the number of variants tested in the region, or an FDR-based approach (e.g., Benjamini-Hochberg), could also be applied to the downstream regional/conditional tests. — This would provide an additional, explicit layer of error-rate control specific to the follow-up regional analyses, complementing the genome-wide threshold already used for initial discovery.
  • The TWAS of NK2R, HK1 and TSPAN15 expression versus HbA1c (Table 4) reports nominal P values without genome-wide FDR correction, noting that only three genes were considered.
    Could also: A simple Bonferroni correction for three tests (dividing alpha by 3) could also be reported alongside the nominal values. — This offers readers an easy secondary reference point for how the associations hold up under a conservative small-family correction, even when the number of comparisons is small.
  • Physiological outcomes in mice (oxygen consumption, food intake, body weight over time) are summarized as mean ± s.e.m. in relatively small groups (n=6–10 per group).
    Could also: Reporting SD or 95% confidence intervals alongside or instead of SEM could also be used to convey the spread of the data. — SEM shrinks with sample size and can visually understate variability in small groups; SD or CI more directly communicate the dispersion of individual animal responses.
  • Repeated measurements over time (e.g., oxygen consumption traces, cumulative food intake, body weight across the dosing period) appear to be compared between treatment groups.
    Could also: A repeated-measures ANOVA or a linear mixed-effects model with animal as a random effect could also be used for these longitudinal comparisons. — These approaches explicitly account for correlation between repeated observations from the same animal over time, which independent testing at each timepoint does not.
  • The specific statistical tests applied to the animal experiments (underlying the asterisk significance thresholds in Fig. 1) are not named in the main text and are instead deferred to the Supplementary Information and Source Data.
    Could also: Stating the specific test (e.g., two-way ANOVA with a post-hoc correction, or a mixed model) directly in the main figure legend could also be done. — This would let readers evaluate the statistical approach directly from the main text, without needing to cross-reference supplementary materials.
  • Genotype-stratified comparisons (e.g., NK2R expression by rs139900276 genotype, Fig. 1e) are shown via box plots with median and Tukey's whiskers.
    Could also: Overlaying individual data points, or using a violin plot, could also be used to display the underlying distribution. — This can help convey sample size and distribution shape per genotype group, which is particularly informative when group sizes are unequal or modest.
Software: CARMA (fine-mapping) · PrediXcan (TWAS) · Variant Effect Predictor

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Reproduction scope — PMID 39537932

Paper: Sass et al. 2024, NK2R control of energy expenditure and feeding to treat metabolic diseases. Nature. DOI 10.1038/s41586-024-08207-0. PMCID PMC11602716.

Metadata correction (harvest false-positives)

The auto-harvested pointers for this RU were both wrong (text-mining false positives):

  • Code harvested as mmkim1210/ggLD → WRONG. ggLD is an unrelated ggplot2 LD-matrix visualisation tool (R, last push 2021, different author). The paper's Code availability states: "The source code used to analyse the snRNA-seq data and produce figures is available at https://github.com/perslab/Sass-2024/."correct repo = perslab/Sass-2024 (authors' own R/Seurat code, GPL-3.0).
  • Data harvested as GSE166649 → WRONG. GSE166649 is the Ludwig-2021 "genetic map of the murine dorsal vagal complex" atlas, which this paper uses only as the reference for label transfer, not its own deposit. The paper's Data availability states the snRNA-seq data are deposited under GSE276735. → correct accession = GSE276735 (BioProject PRJNA1158778).

The corrected pointers are recorded in code/code.json and data/data.json.

What the paper is

Overwhelmingly a wet-lab / pharmacology / in-vivo physiology paper (NK2R agonist EB1002 across mice and non-human primates; food intake, energy expenditure, clamps, FOS imaging, iDISCO). There is one pipeline-derived computational component with deposited data + code: a single-nucleus RNA-seq (snRNA-seq) atlas of the mouse dorsal vagal complex (DVC) comparing vehicle- vs EB1002(="344")-treated mice (Fig. 4k–l, Extended Data Fig. 5h–j).

In scope (pipeline-derived, attempted)

The snRNA-seq pipeline in perslab/Sass-2024 (Seurat v4.3.0; cellranger 7.0.0 + cellbender v0.3 upstream). Concrete reported claims we target:

id reported claim paper location reproducibility
C1 23,664 nuclei in final neuronal dataset Results ("23,664 nuclei were categorized into 25 neuronal populations") from deposit (barcodes/metadata)
C2 25 neuronal populations same sentence; Fig. 4k from deposit (predicted.id levels)
C3 nuclei retained have ≥1000 detectable genes; genes in ≥10 nuclei Methods from deposit (nFeature_RNA, gene count)
C4 both treatment groups evenly distributed across all neuronal clusters Extended Data Fig. 5h from deposit (treatment × cluster)
C5 Glu3 neurons most responsive to EB1002 by IEG (immediate-early-gene) score Fig. 4l INDEPENDENT recompute from deposited matrix (Seurat AddModuleScore + emmeans, exact repo code)
C6 Glu3 top by scDist transcriptional distance Fig. 4l / Methods INDEPENDENT recompute (scDist package, exact repo params)

C1–C4 are direct verifications that the deposited object matches the reported numbers (the deposit is the final annotated neuron object). C5–C6 are genuine independent recomputations of a derived result using the authors' exact code on the deposited matrix → the strongest reproduction.

Out of scope (not pipeline-reproducible / not attempted, with reason)

  • cellranger alignment + cellbender (raw FASTQ → filtered counts): GEO ships only the final merged/annotated matrix, not raw cellranger output; cellbender needs GPU. Upstream of the deposit.
  • Label transfer / cell-subtype annotation (Label_Transfer.Rmd): requires lab-internal reference objects (neuron_type_info.rds, neurons_Seurat_obj.rds at private /projects/... paths = the Ludwig-2021 atlas). Reference not shipped at usable form → predicted.id labels are taken as given from the deposit, not re-derived.
  • All in-vivo / wet-lab / imaging / NHP results (the bulk of the paper): non-computational, out of scope by definition.
  • Genetics framing (TWAS/PrediXcan, CARMA fine-mapping of HbA1c, LD clumping) motivating N
Figures / tables: Fig. 4kExtended Data FigFig. 4l
C1
Reported
23,664 nuclei (Fig. 4k)
Reproduced
23664 nuclei (deposited matrix 21756x23664)
exact
C2
Reported
25 neuronal populations (Fig. 4k)
Reproduced
25 predicted.id classes (Chat1-3, GABA1-7, Glu1-15)
exact
C3
Reported
nuclei >=1000 genes; genes in >=10 nuclei (Methods)
Reproduced
min nFeature_RNA=1002; 21756 genes; integer counts
within tolerance
C4
Reported
treatment groups evenly distributed across all clusters (ED Fig. 5h)
Reproduced
47-62% vehicle per cluster (mostly ~50%); 12336 Veh / 11328 EB1002
within tolerance
C5
Reported
Glu3 most EB1002-responsive by IEG (Fig. 4l)
Reproduced
Glu3 = top cluster, estimate 0.0650, padj 5.5e-11 (only significant cluster; next Glu7 0.0243)
exact
C6
Reported
Glu3 top scDist transcriptional distance (Fig. 4l)
Reproduced
scDist env build in progress; conclusion already confirmed independently by C5
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 87/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

This is a clean reproduction: a predominantly wet-lab Nature paper whose single pipeline-derived component (the DVC snRNA-seq atlas, Fig. 4k-l) reproduces exactly. C1-C4 confirm the deposited GSE276735 object matches the paper text (23,664 nuclei, 25 populations), and C5 independently recomputes the headline Fig. 4l result 1:1 — Glu3 is the top and only significant EB1002-responsive cluster (padj 5.5e-11). The only deviations are rounding/threshold-level (nFeature 1002, ~50% balance) and lie on our side/technical, not the authors'. C6 (scDist) is an unfinished but corroborative secondary method, so it does not affect the verdict.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

242.5 k
tokens (I/O) · 25.1 M incl. cache
40 min
runtime · 0.01 CPU-h
4.7 GB
peak RAM
1
HPC jobs
hummel
machine