Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Distinct sympathetic projections to brown fat regulate thermogenesis and glucose tolerance.

Nat Metab · 2026
L1 85/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (1:1 within tolerance) on «our HPC» for both in-scope sequencing pipelines using the authors' own code on the deposited GEO data. scRNA (GSE310280): 4 clusters exact; 91.6% iBAT(A555) inputs in SG1+SG2 (>85%); NPY 7.3-54x in SG1 (7-50x); independent Seurat re-clustering recovers the deposited 4-cluster partition at 100% purity (merged ARI=1.0); N=878 vs reported 879 (off-by-one). Bulk (GSE310279): DESeq2+sva LRT gives 88 DEGs / 58 up / 30 down vs reported 87/56/31 (within version-drift tolerance, order_match=TRUE); all 6 paper-named marker genes (Jun,Egr3,Adrb2,Irf4,Irs2,Thbs1) present; Ucp1 correctly NOT a DEG. The one miss: top GO BP term ('cell activation' not recovered as #1; adipocyte/hormone terms instead) - attributed to org.Mm.eg.db/GO.db annotation drift over 4 years. Methods-vs-code discrepancies flagged (bulk used kallisto not STAR/featureCounts as text says; scRNA used res 1.9 + manual relabel, Seurat v4 not v3, not res 0.6) - documentation inaccuracies, not result fabrication. NOT attempted (out of scope): anesthetized recordings, metabolic cages, iBAT imaging (wet-lab/physiology/imaging, no pipeline/accession).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 85
    assessed: 2026-06-19 ⛓ 28c4e6e1581c
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Given that thermogenic and glucoregulatory effects of brown adipose tissue (BAT) activation can be dissociated in certain physiological states, the paper tests whether distinct sympathetic neuron subpopulations projecting to intrascapular BAT (iBAT) parenchyma versus vasculature independently control thermogenesis/blood flow and systemic glucose tolerance, respectively.

Core claims
  • Distinct sympathetic neuron subpopulations in the stellate ganglion (SG) innervating iBAT parenchyma versus its vasculature mediate separable functions of the depot. finding
  • Stimulation of parenchymal SG projections to iBAT increases local blood flow and thermogenesis without altering circulating glucose. finding
  • Stimulation of vascular SG projections to iBAT improves glucose tolerance without altering iBAT blood flow or thermogenesis. finding
  • The authors developed a Cre-dependent MaCPNS1 viral toolkit enabling targeted retrograde labeling and chemogenetic manipulation of SG projections to iBAT while sparing other organs and sensory (DRG) inputs. method
  • scRNA-seq coupled with retrograde tracing identifies four molecular SG neuron clusters (SG1-4); SG1 and SG2 provide >85% of sympathetic input to iBAT. finding
  • Chemogenetic stimulation of iBAT-projecting SNS neurons (SG_BAT) alters the iBAT transcriptome and plasma/iBAT lipid profiles in a sex-dependent manner. finding
  • The MaCPNS1 viral serotype achieves substantially more efficient retrograde labeling of the SG from iBAT than conventional AAVrg. method
  • Chemogenetic effects on thermogenesis and blood flow require adrenergic signaling, as shown by reversal with propranolol and partial reduction of blood flow with phentolamine. mechanism
Experimental setups
Assay System Perturbation Readout Platform
Retrograde viral/tracer labeling (AAVrg-GFP, CTB, MaCPNS1) WT and Th-Cre mice, stellate ganglion and DRG intra-iBAT injection of tracer/virus efficiency and specificity of neuronal labeling in SG vs DRG
Chemogenetic activation (DREADD hM3Dq) with iBAT thermography and laser Doppler blood flow Th-Cre mice injected with Cre-dependent MaCPNS1-hM3Dq in iBAT, under anaesthesia CNO (10 mg/kg, i.p.) vs saline; propranolol or phentolamine blockade iBAT-lower back temperature difference and iBAT blood flow
Indirect calorimetry (metabolic cages) Awake, freely moving Th-Cre mice expressing Cre-dependent hM3Dq in iBAT DCZ (200 µg/kg, i.p.) vs saline, random crossover energy expenditure and respiratory exchange ratio (RER)
Oral glucose tolerance test (OGTT) Awake Th-Cre mice expressing Cre-dependent hM3Dq (excitatory) or hM4Di (inhibitory) DREADD in iBAT, plus WT controls DCZ vs saline before oral glucose gavage (2 g/kg) circulating blood glucose over time
Bulk RNA sequencing iBAT from anaesthetized Th-Cre mice expressing Cre-dependent hM3Dq-MaCPNS1 DCZ vs saline, 30 min post-treatment differentially expressed genes, GSEA/GO enrichment, controlling for sex
Lipidomics Plasma and iBAT from the same DCZ/SAL-treated Th-Cre hM3Dq mice, separated by sex DCZ vs saline lipid species and lipid class concentrations
Single-cell RNA-sequencing with CTB retrograde tracing Stellate ganglion neurons from 8-week-old mice, FACS-sorted after CTB-A555 (iBAT) and CTB-A488 (forelimb) injection none (tracing only) transcriptomic clustering (SG1-4), projection target identity, marker gene expression FACS + scRNA-seq
Single-molecule FISH (smFISH) and immunofluorescence SG tissue from WT, Th-Cre::tdTomato, Rxfp1-Cre, Vmat1-Cre and Rxfp1-Flp::tdTOM;Vmat1-Cre mice intra-iBAT injection of tracer/MaCPNS1 reporter viruses validation of cluster marker gene expression (Npy, Fst, Sctr, Rxfp1, Slc18a1, Colq, Aqp1) and co-localization with NPY/mCherry/EGFP
Key results
  • CNO stimulation of SG_BAT hM3Dq neurons increased iBAT temperature by ~0.8°C and blood flow by ~25% within 5 min, with no effect in control-virus mice ~0.8°C; ~25%
  • DCZ stimulation of excitatory DREADD lowered blood glucose after oral glucose gavage at 15 min P = 0.0016 at 15 min
  • Chemogenetic inhibition (hM4Di) of SG_BAT neurons impaired glucose tolerance at 30 min after gavage P = 0.04271 at 15 min
  • DCZ treatment produced a significantly greater increase in energy expenditure than saline in DREADD-expressing mice of both sexes P < 0.001
  • Average RER in the 30 min post-injection decreased with DCZ but not SAL P < 0.001
  • Bulk RNA-seq identified differentially expressed genes in iBAT with DCZ treatment; 'cell activation' was the top enriched GSEA term; Ucp1 was not increased at this timepoint 87 DEGs (56 up, 31 down)
  • Chemogenetic stimulation reduced circulating triglyceride (TG) species in plasma to a similar degree in females and males 43.1% (females), 46.4% (males)
  • Most sympathetic input to iBAT arises from SG1 and SG2 neuronal clusters identified by scRNA-seq >85%
Key statistics
  • pvalue P < 0.001 (change in iBAT temperature and blood flow following SAL vs CNO)
  • other ~0.8 °C temperature increase; ~25% blood flow increase (iBAT thermogenesis and blood flow within 5 min of CNO)
  • mean 60 ± 24.3% reduction in blood flow (effect of phentolamine (α-adrenergic antagonist) on iBAT blood flow)
  • pvalue P = 0.0016 (DCZ vs SAL effect on blood glucose at 15 min during OGTT (excitatory DREADD))
  • pvalue P = 0.04271 (DCZ vs SAL effect on blood glucose at 15 min after glucose gavage (inhibitory DREADD, hM4Di))
  • count 87 differentially expressed genes (56 upregulated, 31 downregulated) (bulk RNA-seq of iBAT, DCZ vs SAL, controlling for sex)
  • fold_change 43.1% (female) and 46.4% (male) reduction in plasma TG (lipidomics comparison of DCZ vs SAL plasma triglycerides)
  • count 1,039 single cells (433 SG_BAT, 220 SG_FL, 375 non-labelled) from n = 12 SG (scRNA-seq with retrograde tracing from iBAT and forelimb)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper uses a mouse model with chemogenetic (DREADD) activation or inhibition of specific sympathetic neuron populations projecting to brown adipose tissue (iBAT), comparing physiological readouts (iBAT thermogenesis, blood flow, energy expenditure, RER, oral glucose tolerance) between vehicle (SAL) and agonist (CNO/DCZ) conditions, largely using within-subject, randomized crossover designs analyzed with two-tailed paired t-tests without correction for multiple comparisons. Bulk RNA-sequencing and lipidomics were used to characterize transcriptional (87 differentially expressed genes, GSEA/GO-ORA enrichment) and lipid changes in iBAT and plasma following stimulation, and single-cell RNA-sequencing coupled with retrograde tracing was used to define sympathetic neuron subtypes. Results are reported throughout as mean ± s.e.m., with exact p-values given for most statistically compared endpoints.

Replicationbiological Sample sizeSample sizes are stated per figure/comparison (e.g., n = 16, 19, 5, 3-5 mice), but no formal power analysis or sample-size justification is described. GroupsVehicle (SAL) vs DREADD agonist (CNO or DCZ) within the same animals; male vs female comparisons for transcriptomic/lipidomic data; control vs hM3Dq/hM4Di virus Pairingmixed Randomization/blindingstated DispersionSEM Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionnone stated (explicitly 'no adjustment')
Statistical tests used
Test Applied to n Assumptions
Paired t-test, two-tailed, no adjustment iBAT thermogenesis and blood flow, SAL vs CNO (Fig. 1f,g) n = 16 mice, tested in both conditions not stated
Paired t-test, two-tailed, no adjustment Energy expenditure and RER, SAL vs DCZ (Fig. 1j,k) n = 16 mice each, tested in both conditions not stated
Paired t-test, two-tailed, no adjustment Blood glucose during OGTT, SAL vs DCZ, excitatory DREADD (Fig. 1l,m) n = 19 mice, tested in both conditions not stated
Paired t-test, two-tailed, no adjustment Blood glucose after glucose gavage, SAL vs DCZ, inhibitory DREADD (Fig. 1o,p) n = 5 mice not stated
Paired t-test, two-tailed, no adjustment Plasma and iBAT lipid class/species concentrations, SAL vs DCZ, by sex (Fig. 2f-m) n = 3-5 SAL-treated and n = 4 DCZ-treated mice per sex not stated
Differential gene expression analysis (method not named) plus GSEA/GO over-representation analysis iBAT bulk RNA-seq, DCZ vs SAL controlling for sex, and male vs female comparisons (Fig. 2a-e) separate cohort of female and male mice, exact n not stated in text not stated
Approaches that could also have been used
  • Many related physiological endpoints (thermogenesis, blood flow, EE, RER, glucose, lipid species) were each compared with separate paired t-tests, explicitly without adjustment for multiple comparisons.
    Could also: A repeated-measures ANOVA or mixed-effects model spanning the related endpoints, combined with a multiple-comparison correction (e.g., Holm-Šidák or Benjamini-Hochberg FDR) — This would jointly model related outcomes and help control the family-wise error rate across the many paired comparisons performed.
  • Differentially expressed genes (87 total) were identified from bulk RNA-seq comparing DCZ vs SAL, controlling for sex, without the specific test/correction method stated in the text.
    Could also: A named RNA-seq differential expression pipeline with explicit FDR control (e.g., DESeq2 Wald test with Benjamini-Hochberg correction, or edgeR/limma-voom) — Explicitly stating and applying an FDR-based correction across the transcriptome-wide set of tested genes helps readers gauge the expected false discovery rate given the large number of genes tested simultaneously.
  • Lipid species and lipid classes in plasma and iBAT were compared between SAL- and DCZ-treated mice across many individual lipid species (volcano plots) using paired t-tests.
    Could also: Multiple-testing correction across lipid species (e.g., Benjamini-Hochberg FDR, as is common in lipidomics volcano plot analyses) — This would help distinguish lipid species with robust changes from those expected by chance given the large number of simultaneous species-level comparisons.
  • Dispersion is reported as mean ± s.e.m. throughout, including for small cohorts (e.g., n = 3-5 per group in lipidomics/RNA-seq).
    Could also: Reporting SD or a 95% confidence interval alongside or instead of s.e.m. — SD directly conveys the spread of individual data points, and a CI conveys the precision of the estimated group difference, which can be especially informative when group sizes are small.
  • Sex differences and treatment effects were examined through separate contrasts (treatment effect within each sex; male vs female independent of treatment) rather than a single combined model.
    Could also: A two-way (factorial) model explicitly testing a treatment × sex interaction term — This would formally quantify whether the treatment effect differs by sex, rather than inferring this from separate within-sex analyses.
  • Effect magnitudes are described mainly as percentage changes or raw mean differences (e.g., ~0.8 °C increase, ~25% increase in blood flow).
    Could also: Standardized effect size measures (e.g., Cohen's d) alongside the raw differences — Standardized effect sizes can facilitate comparison of effect magnitude across different studies, measures, or labs.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41559445

Paper: Neri et al., "Distinct sympathetic projections to brown fat regulate thermogenesis and glucose tolerance." Nat Metab (2026). DOI 10.1038/s42255-025-01429-0. Repo: https://github.com/DanieleN90/Sympathetic-control-of-brown-fat @ f021fa29ff0576e7a9cbfb20f39ba23e5b629392 (authors' own code, P16 satisfied).

IN SCOPE (pipeline-derived sequencing results)

  1. Bulk RNA-seq DEGs (GSE310279, iBAT chemogenetics, ThCre). Pipeline: counts (kallisto est_counts, deposited all_samples_raw.txt) -> DESeq2 LRT with sva surrogate-variable correction (3 SVs), design ~SV1+SV2+SV3+Sex*Injection; injection main effect (Wald contrast, 0.5 weight on interaction), padj<0.05. Code: DESeq2 analysis.R. Target: 87 DEGs (56 up / 31 down); top GO BP term "cell activation"; Ucp1 not increased.
  2. scRNA-seq clustering (GSE310280, stellate ganglion, plate-based). Pipeline: Seurat QC (nFeature>1000, percent.mt<10) -> LogNormalize -> 6000 HVG -> ScaleData (regress percent.mt,nCount_RNA,set) -> PCA(30) -> FindNeighbors(1:30) -> FindClusters(res 1.9) -> manual relabel 12 louvain -> 4 (SG1-SG4) -> t-SNE. Code: scSeq getting from raw data to annotated dataset.R. Target: 879 cells post-QC; 4 clusters; >85% iBAT(A555) inputs in SG1+SG2; NPY 7-50x higher in SG1.

OUT OF SCOPE (wet-lab / physiology / imaging — not bioinformatic pipelines)

  • anesthetized recordings/ (temperature/Doppler/sympathetic-nerve recordings, R plotting of physiology traces)
  • metabolic cages analysis/ (CLAMS food intake/heat/RER analysis)
  • whole iBAT imaging analysis/ (ImageJ macro nerve/vessel measurement; the imaging-analysis folder itself was deleted in the last commit) These are device/wet-lab measurements visualized in R, not reproducible sequencing pipelines, and have no deposited sequencing accession.

Notes / methods-vs-code discrepancies (for AUDIT)

  • Methods text states bulk alignment "STAR v2.5.3a + featureCounts v1.5.3 + UMI-tools"; the deposited counts + repo script actually use kallisto pseudoalignment (est_counts_genes_kallisto.txt / all_samples_raw.txt; GEO summary.csv reports "Pseudoaligned_Reads"). DESeq2 step is faithful; upstream aligner differs from the text.
  • Methods text states scRNA used "Seurat v3", "42 PCs of which 30 retained", "resolution 0.6"; the repo script uses Seurat v4.3.0, 30 PCs, resolution 1.9 then manual 12->4 relabel. The deposited object stores all resolutions 0.1-2.0; final_idents = 4 clusters.
scrna_ncells
Reported
879 cells post-QC
Reproduced
878 cells
within tolerance
scrna_nclusters
Reported
4 clusters (SG1-SG4)
Reproduced
4 (388/335/101/54)
exact
scrna_iBAT_SG12
Reported
>85% iBAT inputs in SG1+SG2
Reproduced
91.6% of A555 cells
exact
scrna_NPY_SG1
Reported
NPY 7-50x higher in SG1
Reproduced
7.3 / 9.0 / 54.0x (linear-space mean)
within tolerance
scrna_reclustering
Reported
Seurat pipeline -> 4 clusters
Reproduced
re-run nests into deposited 4 at 100% purity, merged ARI=1.0
exact
bulk_DEG_total
Reported
87 DEGs
Reproduced
88
within tolerance
bulk_DEG_up
Reported
56 up
Reproduced
58
within tolerance
bulk_DEG_down
Reported
31 down
Reproduced
30
within tolerance
bulk_markers
Reported
Jun,Egr3,Adrb2,Irf4,Irs2,Thbs1 among DEGs
Reproduced
all 6 present (positive log2FC)
exact
bulk_Ucp1
Reported
Ucp1 not increased
Reproduced
Ucp1 not in DEGs
exact
bulk_GO_top
Reported
'cell activation' top GO BP term (up genes)
Reproduced
top terms 'fat cell differentiation'/'pancreatic cell proliferation'; 'cell activation' not top-8
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1

Both in-scope sequencing pipelines reproduce 1:1 within tolerance on the deposited GEO data using the authors' own code: scRNA gives exactly 4 clusters with >85% iBAT inputs in SG1+SG2 and NPY enrichment in SG1, and bulk gives 88/58/30 DEGs (vs 87/56/31) with all named markers up and Ucp1 not a DEG. The deviations are minor and technical — ±2 DEGs from version drift, an off-by-one cell count (878 vs 879), and a GO top-term label change attributable to GO.db annotation drift. The notable caveats are authors'-side documentation discrepancies (methods text cites STAR/featureCounts and Seurat v3/res 0.6, but kallisto and Seurat v4/res 1.9 were actually used) which do not alter the results. Overall a strong reproduction with explainable, non-substantive deviations.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

276.6 k
tokens (I/O) · 29.7 M incl. cache
37 min
runtime · 0.03 CPU-h
2.8 GB
peak RAM
2
HPC jobs
hummel
machine