Purinergic adipocyte-macrophage crosstalk promotes degeneration of thermogenic brown adipose tissue
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the ONE fully-public, pipeline-derived result: Fig 3O, a gene-based PheWAS of the P2X receptor family (P2RX1-7) vs UK Biobank metabolic phenotypes built with the third-party ExPheWas tool. Re-querying the ExPheWas v1 API for all 7 genes x 11 phenotypes reproduces the paper's own source-data values (3O.xlsx) essentially 1:1: Pearson r=1.0000, Spearman rho=0.9989 over 28 non-null cells, with the only difference a uniform +3.21 -log10(P) offset (~1620x stronger p) consistent with ExPheWas now using a larger UK Biobank freeze. P2X4 and P2X7 dominate; CRP/GGT/IGF-1 are driven only by them; P2X1/2/5/6 are null. Honest nuance: the prose claims 'no other P2X member' correlates, but P2X3 carries a secondary adiposity signal visible in the authors' own figure and our re-query (no CRP) - a wording overstatement, not a reproduction failure. No fabrication found. NOT attempted: (a) the two snRNA-seq re-analyses (Fig EV3B Behrens-2025/GSE218710; Fig EV3G Sun-2020) because the paper deposits no Expanded-View source data, so there is no reported number to compare a re-run against - datasets were profiled instead, no compute forced; (b) the bulk RNA-seq DE/GO pipeline because no raw reads or DEG table are deposited ('This study includes no data deposited in external repositories') = data_unavailable. No «our HPC» compute was needed (the reproducible result is a control-plane API requery). The rest of the paper is wet-lab/proprietary-lipidomics and out of scope.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 78assessed: 2026-06-18 ⛓ 8abc1f0d11e8
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether an imbalance between sympathetic activation and mitochondrial energy handling in brown adipose tissue (BAT) drives its inflammatory degeneration, and specifically whether ATP released by stressed brown adipocytes acts as a paracrine signal that activates purinergic receptors (P2X4/P2X7) on BAT-resident macrophages to initiate this process.
- ★ Imbalanced sympathetic/thermogenic activation (via UCP1 deficiency or pharmacological futile thermogenesis) causes BAT inflammation, fibrosis and degeneration finding
- ★ Brown adipocytes secrete ATP in response to imbalanced thermogenic activation finding
- ★ Secreted ATP activates P2X4 and P2X7 purinergic receptors on BAT-resident macrophages mechanism
- ★ Myeloid-specific loss of P2X4/P2X7 activity protects mice against BAT inflammation, thermogenic dysfunction and systemic metabolic disturbances finding
- ★ Combined etomoxir + CL316,243 treatment is a pharmacological in vivo model that mimics futile thermogenic activation seen in UCP1-deficient adipocytes method
- ★ BAT degeneration is preceded by early infiltration of pro-inflammatory myeloid cells (neutrophils, macrophages), with T and B cell infiltration occurring later finding
- ★ Surgical BAT denervation prevents the inflammatory and fibrotic phenotype of Ucp1-/- mice finding
- RNA-seq of degenerating BAT shows enrichment of purine nucleotide and ATP metabolism pathways alongside immune regulation pathways finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Histology (HE, Sirius Red) and MAC2 immunostaining | BAT of Ucp1-/- and WT mice | UCP1 knockout, housing at 22°C or 6°C | lipid accumulation, fibrosis, macrophage abundance | — |
| qRT-PCR gene expression | BAT of Ucp1-/- and WT mice, denervated (DNV) vs sham | surgical BAT denervation | expression of Ucp1, inflammatory, immune cell and fibrosis marker genes | — |
| Norepinephrine quantification | BAT of DNV/sham Ucp1-/- and WT mice | surgical BAT denervation | tissue norepinephrine levels | — |
| Pharmacological in vivo model (etomoxir + CL316,243) with qPCR and histology | BAT of wild-type mice | etomoxir (FAO inhibitor) and/or CL316,243 (β3-adrenergic agonist), 3-day daily injection | lipid accumulation, fibrosis, MAC2+ macrophages, Ucp1 and inflammatory gene expression | — |
| Indirect calorimetry | wild-type mice | mock/CL vs Eto/CL treatment | energy expenditure, respiratory quotient, body core temperature | — |
| Flow cytometry (FACS) immune cell phenotyping | BAT of wild-type mice | Eto/CL treatment, kinetic time course 0-48 h | quantification of neutrophils, macrophage subsets, NK, T and B cells | — |
| Bulk RNA sequencing with DESeq2 and GO enrichment analysis | BAT of wild-type mice | Eto/CL vs mock/mock treatment | differentially expressed genes and enriched biological pathways | Novogene analysis pipeline (DESeq2) |
| Supernatant analysis (ATP, FFA, cytokine measurement) | primary brown adipocytes (in vitro culture) | etomoxir, CL316,243, or combination | secreted CCL2, IL1β, free fatty acids, extracellular ATP | — |
- ▲ Ucp1-/- mice at 22°C show higher BAT weight, increased pro-inflammatory/fibrosis gene expression, more macrophages and fibrosis versus WT
- ▲ Combined Eto/CL treatment (but not single compounds) caused lipid accumulation, fibrosis, MAC2+ macrophage infiltration and reduced Ucp1 expression
- ▼ Eto/CL-treated mice showed reduced energy expenditure and lower body core temperature after CL injection compared to CL alone
- ▼ BAT denervation prevented the MAC2 elevation and inflammatory/fibrotic phenotype otherwise seen in Ucp1-/- mice
- ▲ Kinetic time course showed Ccl2 induction at 4 h, Tnf/Il1b/Emr1 induction at 8 h, and myeloid cell (neutrophil, macrophage) infiltration preceding T and B cell infiltration which appeared only at 48 h
- – RNA-seq comparing Eto/CL vs mock showed regulation of nearly 10,000 genes, with GO enrichment highlighting immune regulation and purine nucleotide/ATP metabolism pathways ~10,000 genes
- ▲ In primary brown adipocytes, combined etomoxir+CL (but not single treatments) increased extracellular ATP accumulation; FFA increased with CL treatment; CCL2 and IL1β secretion were largely unaltered
- ▼ Denervation abolished sympathetic innervation marker tyrosine hydroxylase and norepinephrine levels in BAT of both genotypes
- pvalue P<0.0001 (Ucp1 expression, denervated vs sham mice (Fig 1B))
- pvalue P=0.0212 / P=0.0303 / P=0.0006 (Norepinephrine levels, DNV vs sham comparisons (Fig 1D))
- pvalue P<0.0001 (Energy expenditure 5h pre/post first CL injection, mock/CL vs Eto/CL (Fig 1K))
- pvalue P=0.0026 (Change in body core temperature after CL injection (Fig 1L))
- count ~10,000 differentially expressed genes (RNA-seq of BAT, Eto/CL vs mock/mock (Fig 2M))
- pvalue P=2.65E-05 (Ccl2 gene expression 4h after Eto/CL injection (Fig 2A))
- pvalue P=1.73E-06 (Ccl2 gene expression 48h after Eto/CL injection (Fig 2D))
- other adjusted P value < 0.01 (Threshold used to define differentially expressed gene list in Dataset EV1)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper used two-way ANOVA, unspecified one-way ANOVA, and Student's t-tests on in vivo mouse experiments with 4–6 biological replicates per group to compare gene expression, immune cell counts, protein levels, and metabolic parameters across genotypes, surgical interventions, and pharmacological treatment groups. Bulk RNA-seq differential expression was analyzed via DESeq2 with an adjusted P-value threshold of 0.01. All continuous outcomes were reported as mean ± SEM with exact P-values; no effect sizes or confidence intervals were reported. Post-hoc correction methods for ANOVA comparisons and multiplicity handling outside of RNA-seq were not explicitly stated.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Two-way ANOVA | Genotype (WT vs Ucp1−/−) factor for Ucp1 expression, norepinephrine levels, and inflammatory gene expression in BAT (Fig 1B, D, E) | n=4 per group | not stated |
| Two-way ANOVA | Surgical intervention (sham vs denervated) factor for Ucp1 expression, norepinephrine levels, and inflammatory gene expression in BAT (Fig 1B, D, E) | n=4 per group | not stated |
| ANOVA (type not specified) | Pharmacological treatment group comparisons (mock/mock, mock/CL, Eto/mock, Eto/CL) for BAT and WAT gene expression (Fig 1G–I) | n=4–5 per group | not stated |
| ANOVA (type not specified) | Energy expenditure comparison across pharmacological treatment groups (Fig 1K) | n=6 per group | not stated |
| Student's t-test | Body core temperature response to CL injection comparing mock/CL vs Eto/CL (Fig 1L) | n=6 per group | not stated |
| Student's t-test on log-transformed values | Kinetic gene expression comparisons (Ucp1, Ccl2, Tnf, Il1b, Emr1, Cd4, Cd8b1, Ifng) in BAT at 4 h, 8 h, 24 h, 48 h after Eto/CL vs mock/mock (Fig 2A–D) | n=4 per group | not stated |
| DESeq2 Wald test (negative binomial generalized linear model) | Bulk RNA-seq differential expression of Eto/CL vs mock/mock BAT (~10,000 genes regulated; Fig 2M) | n=4 per group | not stated |
| Gene ontology enrichment analysis | Pathway enrichment on DESeq2-derived differentially expressed genes (Fig 2N); specific enrichment algorithm not named | — | not stated |
-
Dispersion was reported as SEM throughout, including for groups with n=4↳ Could also: SD or 95% confidence intervals could also be used to represent variability — With n as small as 4, SD conveys actual biological variability rather than precision of the mean estimate; 95% CI additionally supports visual inference about group differences and is increasingly recommended by journals for small preclinical samples
-
Many individual gene markers were each tested with separate t-tests or ANOVAs across treatment conditions without a stated family-wise or FDR correction↳ Could also: A Benjamini-Hochberg FDR or Bonferroni correction applied across the panel of simultaneously tested markers could also be reported — Specifying a multiplicity-control procedure for multi-marker panels is standard when several outcomes are tested in parallel; it makes explicit how the type-I error rate is handled across the gene-expression comparisons
-
Post-hoc test following ANOVA comparisons is not named↳ Could also: A named post-hoc procedure such as Tukey HSD (for all-pairs comparisons) or Dunnett's test (for comparisons vs a single control) could also be specified — Naming the post-hoc method clarifies which pairwise contrasts are controlled and at what error rate, directly supporting reproducibility of the reported P-values
-
Gene expression at kinetic timepoints was log-transformed before Student's t-tests↳ Could also: A non-parametric Mann-Whitney U test on the original scale could also have been applied — With n=4 per group normality is difficult to verify empirically; non-parametric tests require no distributional assumption and are a common alternative for small-sample RT-qPCR comparisons
-
Randomization of animals to treatment groups and blinding of outcome assessors are not described↳ Could also: Explicit reporting of randomization procedure and blinding status (or a statement that blinding was not feasible) could also be included — ARRIVE 2.0 and most preclinical reporting guidelines recommend describing these elements; their inclusion allows readers to assess potential sources of performance and detection bias
-
RNA-seq differential expression was analyzed by a single pipeline (DESeq2, outsourced to Novogene) without a reported sensitivity or cross-validation analysis↳ Could also: edgeR or limma-voom could also be applied as complementary pipelines, with findings reported as the intersection or union of significant genes — Different count-model frameworks can differ in their low-count behavior and dispersion estimation; cross-validating with a second pipeline is sometimes used to identify the most robust subset of differentially expressed genes when thousands of genes are regulated
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41261284
Paper: Jaeckstein et al. (2025) Purinergic adipocyte-macrophage crosstalk promotes degeneration of thermogenic brown adipose tissue. EMBO reports. DOI 10.1038/s44319-025-00642-y · PMCID PMC12715258.
Data availability statement (verbatim): "This study includes no data
deposited in external repositories." Only an EMBO Source Data record exists:
BioStudies S-SCDT-10_1038-S44319-025-00642-y (per-figure .xlsx of plotted values).
In scope (pipeline-derived computational results)
| id | result | pipeline / tool | data | reproducible? |
|---|---|---|---|---|
| 3O | Fig 3O — metabolic phenotypes associated with the P2X-receptor family (P2RX1–7) in the UK Biobank | ExPheWas gene-based PheWAS browser (Legault 2022), a third-party tool on public UK Biobank gene-based association summary statistics | ExPheWas v1 API + paper's own source data 3O.xlsx |
YES — reproduced (primary result; no heavy compute) |
| EV3B | Fig EV3B — P2rx4/P2rx7 double-positive nuclei per cell type | re-analysis of published murine snRNA-seq | Behrens 2025 (GEO GSE218710, open; code github.com/AdlungLab/ChREBP) | partial — no pinnable target (no EV source data deposited); dataset profiled only |
| EV3G | Fig EV3G — P2RX4/P2RX7 expression across human BAT clusters | re-analysis of published human snRNA-seq | Sun et al. Nature 2020 (PMID 33116305) | partial — no pinnable target (no EV source data); dataset profiled only |
| (methods) | bulk mRNA-seq: mapping → DEG (padj<0.01, | log2FC | >0) → GO enrichment | Novogene standard RNA-seq pipeline (NovaSeq 6000 PE150) |
Out of scope (wet-lab / manual / proprietary — not attempted)
qPCR (TaqMan/SYBR), Lipidyzer lipidomics (proprietary SCIEX MRM + Lipidomics Workflow Manager; raw data not deposited), histochemistry (HE/Sirius Red/MAC2), indirect calorimetry, flow cytometry, Western blot, ATP/FFA/cytokine assays, denervation surgery, animal phenotyping. All wet-lab measurements; GraphPad Prism t-test/ANOVA statistics on small-n biological replicates — not a bioinformatic pipeline.
Primary reproduction target
Fig 3O is the one clearly-specified, fully public, pipeline-derived result:
a third-party tool (ExPheWas) applied to public data (UK Biobank), with the
exact plotted values shipped in 3O.xlsx. This is the honest 1:1 comparison.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The single fully-public, pipeline-derived result (Fig 3O ExPheWas P2X-family PheWAS) reproduces essentially 1:1 against the authors' own source data (Pearson r=1.0000, Spearman rho=0.9989 over 28 cells), with the only deviation a uniform +3.21 -log10(P) offset explained by a later, larger UK Biobank freeze — a technical/version effect, not an authors' defect. The central claim that P2X4/P2X7 dominate metabolic associations while P2X1/2/5/6 are null holds exactly; the only blemish is a mild text-vs-figure overstatement (P2X3 secondary adiposity signal contradicts 'no other P2X member'), visible in the authors' own data. Overall judged yellow: a clean, explainable reproduction of the reproducible part, but most of the paper (snRNA-seq re-analyses, bulk RNA-seq DEG/GO) is non-reproducible because no data were deposited.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.