Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Intact innervation is essential for diet-induced recruitment of brown adipose tissue

· 2018
PubMed 30576247 ↗ pmid-30576247
L1 No computation 2/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Total score +7
✓ What held up
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
Reproduction agent’s raw note

DROP (non_pipeline). 'Intact innervation is essential for diet-induced recruitment of brown adipose tissue' (Fischer AW, Schlein C, Cannon B, Heeren J, Nedergaard J; Am J Physiol Endocrinol Metab 2019; PMID 30576247; PMCID PMC6459298; DOI 10.1152/ajpendo.00443.2018) is a pure wet-lab / in-vivo mouse physiology study. Methods are surgical denervation of interscapular BAT, Western blot (Li-Cor Image Studio Lite densitometry), qRT-PCR (TaqMan), indirect calorimetry (TSE PhenoMaster), immunofluorescence/IHC, and H&E histology; statistics by Student's t-test in GraphPad Prism 6 / Excel. There is NO bioinformatic or computational pipeline, NO RNA-seq/microarray/sequencing, NO deposited public dataset (GEO/SRA/ENA/ArrayExpress/figshare/zenodo/dbGaP/EGA/PRIDE all checked: zero accessions), NO Data/Code Availability statement, and NO analysis-code repository. Therefore there is nothing to reproduce computationally and no third-party tool applies (no deposited data to run one on). No «our HPC» compute was spent (correctly). This is an honest, valid drop, not a reproduction failure: the experimental science is sound but is out of scope for computational-pipeline reproduction. Headline reported claims (c1-c8) are documented in original/claims.tsv for the human audit trail only and were NOT attempted. Secondary applicable reasons: no_code, no_data_accession.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-18 ⛓ dab66d36b725
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Whether the recruitment of brown adipose tissue (BAT) during diet-induced thermogenesis (DIT) requires intact sympathetic innervation of the tissue, or whether circulating/humoral factors alone can drive this recruitment under obesogenic conditions.

Core claims
  • Intact innervation of BAT is essential for diet-induced thermogenesis/recruitment. finding
  • Circulating factors cannot by themselves initiate recruitment of brown adipose tissue under obesogenic conditions. finding
  • Denervation totally abolished the diet-induced increase in total UCP1 protein levels seen in intact mice, while basal UCP1 expression was innervation-independent. finding
  • IBAT denervation did not enhance obesity but increased metabolic efficiency, so equal obesity was reached with lower food intake. finding
  • Denervation altered the feeding pattern of the mice. finding
  • Denervation of interscapular BAT did not detectably hyper-recruit other BAT depots. finding
  • No UCP1 protein was detected in the browning-competent inguinal white adipose tissue depot under any tested condition. finding
  • Surgical denervation method for interscapular BAT (cutting nerve bundles at two locations, removing intermediate piece) as established by Vaughan et al. method
Experimental setups
Assay System Perturbation Readout Platform
surgical denervation with body weight/food intake monitoring interscapular BAT (IBAT), male C57BL/6J mice surgical denervation vs. sham, chow vs. high-fat diet, thermoneutrality (30°C) body weight gain, food/energy intake, metabolic efficiency
indirect calorimetry whole-body, male C57BL/6J mice IBAT denervation vs. sham, chow vs. HFD O2 consumption, CO2 production, energy expenditure, respiratory quotient, food/water intake TSE PhenoMaster System (TSE Systems)
Western blotting interscapular BAT and other adipose depots IBAT denervation vs. sham, chow vs. HFD UCP1, tyrosine hydroxylase, and γ-tubulin protein levels Amersham Imager 600 (GE); Li-Cor Image Studio Lite
quantitative real-time PCR (TaqMan) adipose tissue (IBAT and other depots) IBAT denervation vs. sham, chow vs. HFD mRNA levels of Ucp1, Adrb3, Dio2 (normalized to Tbp) 7900HT Fast Real-Time PCR System (Applied Biosystems)
histology (H&E staining) IBAT sections IBAT denervation vs. sham, chow vs. HFD tissue morphology Leica Microtome/Autotechnikon
immunofluorescence/confocal imaging IBAT sections IBAT denervation vs. sham tyrosine hydroxylase and perilipin1 staining (innervation confirmation, adipocyte marker) Nikon A1 Ti confocal laser scanning microscope
Key results
  • Body weight of chow-fed mice was stable and unaffected by IBAT denervation.
  • HFD feeding caused significant weight gain, not affected by IBAT denervation.
  • HFD mice had significantly higher energy intake than chow, statistically similar (slightly lower trend) in denervated mice.
  • Metabolic efficiency (body weight gained per food energy) was significantly higher in IBAT-denervated mice than sham.
  • Denervation abolished the diet-induced increase in total UCP1 protein observed in intact (sham) mice.
  • No UCP1 protein detected in inguinal WAT under any diet or denervation condition.
Key statistics
  • pvalue P ≤ 0.001 between diets; P ≤ 0.05 between surgical intervention groups (Figure 1 body weight/metabolic efficiency comparisons)
  • other IBAT estimated to contribute ~70% (or ~25% by another method) of total BAT thermogenic potential (relative contribution of IBAT depot to whole-body BAT thermogenesis)
  • other ~45% of total UCP1 mRNA in IBAT, ~50% in other BAT depots, ~5% in brite/beige depots at 30°C (distribution of UCP1 mRNA across adipose depots in mice at thermoneutrality)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a 2×2 between-subjects design (diet: chow vs. high-fat × surgery: sham vs. interscapular BAT denervation) with male C57BL/6J mice (n = 5 per group) housed at thermoneutrality (30°C). Student's t-test was the sole stated inferential test, applied separately to detect diet effects and denervation effects across all outcomes (body weight, energy intake, metabolic efficiency, UCP1 protein/mRNA, calorimetry). Results were expressed as means ± SE with threshold-based p-value notation (asterisks for diet effect, hash symbols for denervation effect) rather than exact p-values.

Replicationbiological Sample sizen = 5 per group stated in Fig. 1 legend; no formal power calculation or sample size justification described GroupsFour groups: chow-sham, chow-denervated, HFD-sham, HFD-denervated (2 × 2 factorial) Pairingunpaired Randomization/blindingnot stated DispersionSEM Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Student's t-test (direction/tails not stated) Diet effect (chow vs. HFD): body weight gain, energy intake, metabolic efficiency (Fig. 1C–F) n = 5 per group not stated
Student's t-test (direction/tails not stated) Denervation effect (sham vs. DNV): body weight gain, energy intake, metabolic efficiency (Fig. 1C–F) n = 5 per group not stated
Student's t-test (direction/tails not stated) UCP1 protein levels and mRNA (Western blot and qRT-PCR; diet and denervation effects; figures not fully visible in provided text) n = 5 per group (inferred from stated n; exact n per blot not stated in provided text) not stated
Student's t-test (direction/tails not stated) Indirect calorimetry outcomes (energy expenditure, RQ, food intake during calorimetry; Fig. 3) n = 5 per group (inferred; exact n per calorimetry measurement not stated in provided text) not stated
Approaches that could also have been used
  • The 2×2 factorial design (diet × surgery) was analyzed by applying Student's t-tests separately for each factor across each outcome
    Could also: A two-way ANOVA followed by a post-hoc test (e.g., Tukey HSD or Sidak) for each outcome variable could also have been applied — Two-way ANOVA directly tests the diet × denervation interaction term, which is central to the paper's question of whether denervation modifies the diet effect; applying separate t-tests does not formally test this interaction and can inflate type I error across the factorial comparisons
  • Multiple Student's t-tests were conducted across numerous outcome variables (body weight, energy intake, metabolic efficiency, UCP1 protein, multiple mRNA targets, calorimetry variables) without a stated correction for multiple comparisons
    Could also: A multiplicity correction such as Benjamini-Hochberg FDR adjustment or Holm-Bonferroni could also have been applied across the family of tests — With many simultaneous tests, the probability of at least one false-positive result increases; a correction procedure makes the accepted type I error rate explicit and consistent across the study
  • Dispersion was reported as SEM (means ± SE) with n = 5 per group
    Could also: SD or 95% confidence intervals could also have been used to describe spread — With n = 5, SEM is approximately half the SD, which can make variability appear smaller than it is; SD directly describes the observed biological variability in the sample, and 95% CIs additionally convey uncertainty about the group mean estimate, both of which are often preferred for small-n experiments
  • P-values were reported as threshold bands (*, **, ***) rather than exact values
    Could also: Exact p-values (e.g., P = 0.023) could also have been reported — Exact p-values allow readers to apply their own significance thresholds, facilitate meta-analysis, and are recommended by many journals and statistical reporting guidelines (e.g., APA, Nature reporting standards)
  • No effect size metrics were reported alongside significance tests
    Could also: Cohen's d (for t-tests) or partial η² (for ANOVA) could also have been reported — With n = 5, studies are underpowered to detect small effects, so reporting effect sizes helps readers evaluate biological magnitude independently of statistical significance and supports future power calculations
  • UCP1 Western blot signals were quantified as continuous outcomes and compared with t-tests; the provided text does not indicate whether normality was assessed for the small groups
    Could also: A nonparametric alternative such as the Mann-Whitney U test could also have been applied for pairwise comparisons with n = 5 — With only 5 observations per group, normality assumptions are difficult to verify empirically; nonparametric tests make no distributional assumption and are a common choice for small-n Western blot quantification data
Software: Microsoft Excel 2016 · GraphPad Prism 6

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

NOTHING — DROP (non_pipeline). Pure wet-lab/in-vivo mouse physiology study (surgical iBAT denervation; Western blot, qRT-PCR, indirect calorimetry, histology/IF). No bioinformatic pipeline, no deposited data, no code. See reproduction/scope.md and reproduction/ROOM_RESULT.json.

Figures / tables: Fig 6DFig 6CFig 3ETableFig 2Fig 6A

No individual results have been recorded for this entry yet.

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 56/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Total score +7

This is a valid, honest non_pipeline drop, not a reproduction failure. The paper (Fischer/Schlein/Nedergaard, Am J Physiol Endocrinol Metab 2019, PMC6459298) is pure wet-lab/in-vivo mouse physiology — denervation, Western blot, qRT-PCR, TSE PhenoMaster calorimetry, histology — with no computational pipeline, no code, and no deposited data (all major repositories checked, zero accessions). Nothing could be compared 1:1 (q1/q2 red) because nothing was shared, but per the rubric this is a data-availability/scope limitation on the input side, not an authors' defect or fabrication signal (no too-perfect pattern; q5 yellow only because non-derivable, not suspect). The core claim is simply untestable in this framework, so q7/q8 are yellow rather than red.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

36.6 k
tokens (I/O) · 2.2 M incl. cache
5 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.