Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Purinergic adipocyte-macrophage crosstalk promotes degeneration of thermogenic brown adipose tissue

· 2025
PubMed 41261284 ↗ pmid-41261284
L1 78/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Total score +3
✓ What held up
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 51% of all assessed papers rank 533 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the ONE fully-public, pipeline-derived result: Fig 3O, a gene-based PheWAS of the P2X receptor family (P2RX1-7) vs UK Biobank metabolic phenotypes built with the third-party ExPheWas tool. Re-querying the ExPheWas v1 API for all 7 genes x 11 phenotypes reproduces the paper's own source-data values (3O.xlsx) essentially 1:1: Pearson r=1.0000, Spearman rho=0.9989 over 28 non-null cells, with the only difference a uniform +3.21 -log10(P) offset (~1620x stronger p) consistent with ExPheWas now using a larger UK Biobank freeze. P2X4 and P2X7 dominate; CRP/GGT/IGF-1 are driven only by them; P2X1/2/5/6 are null. Honest nuance: the prose claims 'no other P2X member' correlates, but P2X3 carries a secondary adiposity signal visible in the authors' own figure and our re-query (no CRP) - a wording overstatement, not a reproduction failure. No fabrication found. NOT attempted: (a) the two snRNA-seq re-analyses (Fig EV3B Behrens-2025/GSE218710; Fig EV3G Sun-2020) because the paper deposits no Expanded-View source data, so there is no reported number to compare a re-run against - datasets were profiled instead, no compute forced; (b) the bulk RNA-seq DE/GO pipeline because no raw reads or DEG table are deposited ('This study includes no data deposited in external repositories') = data_unavailable. No «our HPC» compute was needed (the reproducible result is a control-plane API requery). The rest of the paper is wet-lab/proprietary-lipidomics and out of scope.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 78
    assessed: 2026-06-18 ⛓ 8abc1f0d11e8
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether an imbalance between sympathetic activation and mitochondrial energy handling in brown adipose tissue (BAT) drives its inflammatory degeneration, and specifically whether ATP released by stressed brown adipocytes acts as a paracrine signal that activates purinergic receptors (P2X4/P2X7) on BAT-resident macrophages to initiate this process.

Core claims
  • Imbalanced sympathetic/thermogenic activation (via UCP1 deficiency or pharmacological futile thermogenesis) causes BAT inflammation, fibrosis and degeneration finding
  • Brown adipocytes secrete ATP in response to imbalanced thermogenic activation finding
  • Secreted ATP activates P2X4 and P2X7 purinergic receptors on BAT-resident macrophages mechanism
  • Myeloid-specific loss of P2X4/P2X7 activity protects mice against BAT inflammation, thermogenic dysfunction and systemic metabolic disturbances finding
  • Combined etomoxir + CL316,243 treatment is a pharmacological in vivo model that mimics futile thermogenic activation seen in UCP1-deficient adipocytes method
  • BAT degeneration is preceded by early infiltration of pro-inflammatory myeloid cells (neutrophils, macrophages), with T and B cell infiltration occurring later finding
  • Surgical BAT denervation prevents the inflammatory and fibrotic phenotype of Ucp1-/- mice finding
  • RNA-seq of degenerating BAT shows enrichment of purine nucleotide and ATP metabolism pathways alongside immune regulation pathways finding
Experimental setups
Assay System Perturbation Readout Platform
Histology (HE, Sirius Red) and MAC2 immunostaining BAT of Ucp1-/- and WT mice UCP1 knockout, housing at 22°C or 6°C lipid accumulation, fibrosis, macrophage abundance
qRT-PCR gene expression BAT of Ucp1-/- and WT mice, denervated (DNV) vs sham surgical BAT denervation expression of Ucp1, inflammatory, immune cell and fibrosis marker genes
Norepinephrine quantification BAT of DNV/sham Ucp1-/- and WT mice surgical BAT denervation tissue norepinephrine levels
Pharmacological in vivo model (etomoxir + CL316,243) with qPCR and histology BAT of wild-type mice etomoxir (FAO inhibitor) and/or CL316,243 (β3-adrenergic agonist), 3-day daily injection lipid accumulation, fibrosis, MAC2+ macrophages, Ucp1 and inflammatory gene expression
Indirect calorimetry wild-type mice mock/CL vs Eto/CL treatment energy expenditure, respiratory quotient, body core temperature
Flow cytometry (FACS) immune cell phenotyping BAT of wild-type mice Eto/CL treatment, kinetic time course 0-48 h quantification of neutrophils, macrophage subsets, NK, T and B cells
Bulk RNA sequencing with DESeq2 and GO enrichment analysis BAT of wild-type mice Eto/CL vs mock/mock treatment differentially expressed genes and enriched biological pathways Novogene analysis pipeline (DESeq2)
Supernatant analysis (ATP, FFA, cytokine measurement) primary brown adipocytes (in vitro culture) etomoxir, CL316,243, or combination secreted CCL2, IL1β, free fatty acids, extracellular ATP
Key results
  • Ucp1-/- mice at 22°C show higher BAT weight, increased pro-inflammatory/fibrosis gene expression, more macrophages and fibrosis versus WT
  • Combined Eto/CL treatment (but not single compounds) caused lipid accumulation, fibrosis, MAC2+ macrophage infiltration and reduced Ucp1 expression
  • Eto/CL-treated mice showed reduced energy expenditure and lower body core temperature after CL injection compared to CL alone
  • BAT denervation prevented the MAC2 elevation and inflammatory/fibrotic phenotype otherwise seen in Ucp1-/- mice
  • Kinetic time course showed Ccl2 induction at 4 h, Tnf/Il1b/Emr1 induction at 8 h, and myeloid cell (neutrophil, macrophage) infiltration preceding T and B cell infiltration which appeared only at 48 h
  • RNA-seq comparing Eto/CL vs mock showed regulation of nearly 10,000 genes, with GO enrichment highlighting immune regulation and purine nucleotide/ATP metabolism pathways ~10,000 genes
  • In primary brown adipocytes, combined etomoxir+CL (but not single treatments) increased extracellular ATP accumulation; FFA increased with CL treatment; CCL2 and IL1β secretion were largely unaltered
  • Denervation abolished sympathetic innervation marker tyrosine hydroxylase and norepinephrine levels in BAT of both genotypes
Key statistics
  • pvalue P<0.0001 (Ucp1 expression, denervated vs sham mice (Fig 1B))
  • pvalue P=0.0212 / P=0.0303 / P=0.0006 (Norepinephrine levels, DNV vs sham comparisons (Fig 1D))
  • pvalue P<0.0001 (Energy expenditure 5h pre/post first CL injection, mock/CL vs Eto/CL (Fig 1K))
  • pvalue P=0.0026 (Change in body core temperature after CL injection (Fig 1L))
  • count ~10,000 differentially expressed genes (RNA-seq of BAT, Eto/CL vs mock/mock (Fig 2M))
  • pvalue P=2.65E-05 (Ccl2 gene expression 4h after Eto/CL injection (Fig 2A))
  • pvalue P=1.73E-06 (Ccl2 gene expression 48h after Eto/CL injection (Fig 2D))
  • other adjusted P value < 0.01 (Threshold used to define differentially expressed gene list in Dataset EV1)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper used two-way ANOVA, unspecified one-way ANOVA, and Student's t-tests on in vivo mouse experiments with 4–6 biological replicates per group to compare gene expression, immune cell counts, protein levels, and metabolic parameters across genotypes, surgical interventions, and pharmacological treatment groups. Bulk RNA-seq differential expression was analyzed via DESeq2 with an adjusted P-value threshold of 0.01. All continuous outcomes were reported as mean ± SEM with exact P-values; no effect sizes or confidence intervals were reported. Post-hoc correction methods for ANOVA comparisons and multiplicity handling outside of RNA-seq were not explicitly stated.

Replicationbiological Sample sizeN values stated per figure to indicate biological replicates; no a priori power calculation or sample size justification described GroupsUcp1−/− vs wild-type mice; sham-operated vs BAT-denervated; four pharmacological groups (mock/mock, mock/CL, Eto/mock, Eto/CL); kinetic timepoints (0, 4, 8, 24, 48 h post-injection) Pairingunpaired Randomization/blindingnot stated DispersionSEM Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR via DESeq2 for RNA-seq (adjusted P < 0.01 threshold stated); post-hoc correction method following ANOVA not stated; no correction method stated for the panel of gene-by-gene t-tests
Statistical tests used
Test Applied to n Assumptions
Two-way ANOVA Genotype (WT vs Ucp1−/−) factor for Ucp1 expression, norepinephrine levels, and inflammatory gene expression in BAT (Fig 1B, D, E) n=4 per group not stated
Two-way ANOVA Surgical intervention (sham vs denervated) factor for Ucp1 expression, norepinephrine levels, and inflammatory gene expression in BAT (Fig 1B, D, E) n=4 per group not stated
ANOVA (type not specified) Pharmacological treatment group comparisons (mock/mock, mock/CL, Eto/mock, Eto/CL) for BAT and WAT gene expression (Fig 1G–I) n=4–5 per group not stated
ANOVA (type not specified) Energy expenditure comparison across pharmacological treatment groups (Fig 1K) n=6 per group not stated
Student's t-test Body core temperature response to CL injection comparing mock/CL vs Eto/CL (Fig 1L) n=6 per group not stated
Student's t-test on log-transformed values Kinetic gene expression comparisons (Ucp1, Ccl2, Tnf, Il1b, Emr1, Cd4, Cd8b1, Ifng) in BAT at 4 h, 8 h, 24 h, 48 h after Eto/CL vs mock/mock (Fig 2A–D) n=4 per group not stated
DESeq2 Wald test (negative binomial generalized linear model) Bulk RNA-seq differential expression of Eto/CL vs mock/mock BAT (~10,000 genes regulated; Fig 2M) n=4 per group not stated
Gene ontology enrichment analysis Pathway enrichment on DESeq2-derived differentially expressed genes (Fig 2N); specific enrichment algorithm not named not stated
Approaches that could also have been used
  • Dispersion was reported as SEM throughout, including for groups with n=4
    Could also: SD or 95% confidence intervals could also be used to represent variability — With n as small as 4, SD conveys actual biological variability rather than precision of the mean estimate; 95% CI additionally supports visual inference about group differences and is increasingly recommended by journals for small preclinical samples
  • Many individual gene markers were each tested with separate t-tests or ANOVAs across treatment conditions without a stated family-wise or FDR correction
    Could also: A Benjamini-Hochberg FDR or Bonferroni correction applied across the panel of simultaneously tested markers could also be reported — Specifying a multiplicity-control procedure for multi-marker panels is standard when several outcomes are tested in parallel; it makes explicit how the type-I error rate is handled across the gene-expression comparisons
  • Post-hoc test following ANOVA comparisons is not named
    Could also: A named post-hoc procedure such as Tukey HSD (for all-pairs comparisons) or Dunnett's test (for comparisons vs a single control) could also be specified — Naming the post-hoc method clarifies which pairwise contrasts are controlled and at what error rate, directly supporting reproducibility of the reported P-values
  • Gene expression at kinetic timepoints was log-transformed before Student's t-tests
    Could also: A non-parametric Mann-Whitney U test on the original scale could also have been applied — With n=4 per group normality is difficult to verify empirically; non-parametric tests require no distributional assumption and are a common alternative for small-sample RT-qPCR comparisons
  • Randomization of animals to treatment groups and blinding of outcome assessors are not described
    Could also: Explicit reporting of randomization procedure and blinding status (or a statement that blinding was not feasible) could also be included — ARRIVE 2.0 and most preclinical reporting guidelines recommend describing these elements; their inclusion allows readers to assess potential sources of performance and detection bias
  • RNA-seq differential expression was analyzed by a single pipeline (DESeq2, outsourced to Novogene) without a reported sensitivity or cross-validation analysis
    Could also: edgeR or limma-voom could also be applied as complementary pipelines, with findings reported as the intersection or union of significant genes — Different count-model frameworks can differ in their low-count behavior and dispersion estimation; cross-validating with a second pipeline is sometimes used to identify the most robust subset of differentially expressed genes when thousands of genes are regulated
Software: DESeq2

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41261284

Paper: Jaeckstein et al. (2025) Purinergic adipocyte-macrophage crosstalk promotes degeneration of thermogenic brown adipose tissue. EMBO reports. DOI 10.1038/s44319-025-00642-y · PMCID PMC12715258.

Data availability statement (verbatim): "This study includes no data deposited in external repositories." Only an EMBO Source Data record exists: BioStudies S-SCDT-10_1038-S44319-025-00642-y (per-figure .xlsx of plotted values).

In scope (pipeline-derived computational results)

id result pipeline / tool data reproducible?
3O Fig 3O — metabolic phenotypes associated with the P2X-receptor family (P2RX1–7) in the UK Biobank ExPheWas gene-based PheWAS browser (Legault 2022), a third-party tool on public UK Biobank gene-based association summary statistics ExPheWas v1 API + paper's own source data 3O.xlsx YES — reproduced (primary result; no heavy compute)
EV3B Fig EV3B — P2rx4/P2rx7 double-positive nuclei per cell type re-analysis of published murine snRNA-seq Behrens 2025 (GEO GSE218710, open; code github.com/AdlungLab/ChREBP) partial — no pinnable target (no EV source data deposited); dataset profiled only
EV3G Fig EV3G — P2RX4/P2RX7 expression across human BAT clusters re-analysis of published human snRNA-seq Sun et al. Nature 2020 (PMID 33116305) partial — no pinnable target (no EV source data); dataset profiled only
(methods) bulk mRNA-seq: mapping → DEG (padj<0.01, log2FC >0) → GO enrichment Novogene standard RNA-seq pipeline (NovaSeq 6000 PE150)

Out of scope (wet-lab / manual / proprietary — not attempted)

qPCR (TaqMan/SYBR), Lipidyzer lipidomics (proprietary SCIEX MRM + Lipidomics Workflow Manager; raw data not deposited), histochemistry (HE/Sirius Red/MAC2), indirect calorimetry, flow cytometry, Western blot, ATP/FFA/cytokine assays, denervation surgery, animal phenotyping. All wet-lab measurements; GraphPad Prism t-test/ANOVA statistics on small-n biological replicates — not a bioinformatic pipeline.

Primary reproduction target

Fig 3O is the one clearly-specified, fully public, pipeline-derived result: a third-party tool (ExPheWas) applied to public data (UK Biobank), with the exact plotted values shipped in 3O.xlsx. This is the honest 1:1 comparison.

Figures / tables: Fig 3O
3O-pattern
Reported
Fig 3O heatmap: only P2X4 & P2X7 strongly associate with metabolic phenotypes across UK Biobank (P2X3 secondary adiposity; P2X1/2/5/6 null)
Reproduced
ExPheWas re-query reproduces the identical pattern; Pearson r=1.0000, Spearman rho=0.9989 over 28 non-null cells
within tolerance
3O-CRP-P2X4
Reported
-log10(P)=119.24 (P2RX4 vs C-reactive protein)
Reproduced
p=3.55e-123, -log10(P)=122.45
within tolerance
3O-CRP-P2X7
Reported
-log10(P)=208.51 (P2RX7 vs C-reactive protein)
Reproduced
p=1.90e-212, -log10(P)=211.72
within tolerance
3O-BMI-P2X4
Reported
-log10(P)=8.21 (P2RX4 vs body mass index)
Reproduced
p=3.73e-12, -log10(P)=11.43
within tolerance
3O-text-exclusivity
Reported
text: variants in P2RX4 and P2RX7 'but no other members of the P2X family' correlate with metabolic phenotypes
Reproduced
P2RX4/P2RX7 dominate, but P2RX3 also reaches FDR-significant adiposity associations (present in the paper's own Fig 3O) — text mildly overstates exclusivity
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Total score +3

The single fully-public, pipeline-derived result (Fig 3O ExPheWas P2X-family PheWAS) reproduces essentially 1:1 against the authors' own source data (Pearson r=1.0000, Spearman rho=0.9989 over 28 cells), with the only deviation a uniform +3.21 -log10(P) offset explained by a later, larger UK Biobank freeze — a technical/version effect, not an authors' defect. The central claim that P2X4/P2X7 dominate metabolic associations while P2X1/2/5/6 are null holds exactly; the only blemish is a mild text-vs-figure overstatement (P2X3 secondary adiposity signal contradicts 'no other P2X member'), visible in the authors' own data. Overall judged yellow: a clean, explainable reproduction of the reproducible part, but most of the paper (snRNA-seq re-analyses, bulk RNA-seq DEG/GO) is non-reproducible because no data were deposited.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

158.4 k
tokens (I/O) · 9.2 M incl. cache
16 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.