Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Apoptotic brown adipocytes enhance energy expenditure via extracellular inosine

· 2022
PubMed 35790189 ↗ pmid-35790189
L1 57/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Total score +8
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🔴A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
57/100
Reproducibility score
1.0 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 17% of all assessed papers rank 965 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH FOR THE IDENTIFICATION COUNT, PARTIALLY FOR THE STATISTICS. The single pipeline-derived dataset (phosphoproteomics, PRIDE PXD032153) was re-fetched to «infra» and analysed on «our HPC» («job»). C1 (38,451 phosphopeptides) reproduces to within 0.8% directly from the shipped MaxQuant Phospho (STY)Sites.txt (38,157) — clean. C2-C4 (regulated-site counts 7,875/8,613/2,535) reproduce only PARTIALLY: a faithful EasyPhos/Perseus two-sample permutation-FDR test, run at BOTH the collapsed-site and the canonical multiplicity-expanded levels across a full valid-value x S0 x FDR-method sweep, yields ~36-58% of the reported counts in EVERY configuration (best: FORSK 3,769 / inosine 5,018 / both 919). The qualitative pattern IS reproduced (inosine regulates more sites than FORSK; large overlap). The Methods do not publish the Perseus parameters and none are deposited, so the reported counts are not derivable from the public data as-is — flagged for human review (agreement.json possible_fabrication_notes), not asserted as fabrication. C5 (PKA over-representation) reproduces for the forskolin/cAMP/PKA arm (OR 1.8-2.5, p<=1e-17) but not for the inosine arm. NOT ATTEMPTED (out of scope): untargeted metabolomics (Compound Discoverer/MetaboAnalyst, no public accession), RNA-seq of Slc29a1/a2 (no resolvable accession), and all wet-lab/in-vivo/human-cohort results (energy expenditure, KO phenotypes, ENT1 Ile216Thr-BMI association).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 59
    assessed: 2026-06-19 ⛓ 34c11c5c1c15
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether apoptotic brown adipocytes release signaling metabolites (particularly purines) that promote replacement/activation of the thermogenic program in surrounding brown/beige adipocytes, and whether the transporter ENT1 (SLC29A1) regulates this process by controlling extracellular inosine levels to influence energy expenditure and obesity.

Core claims
  • Apoptotic brown adipocytes release a specific secretome enriched in purine metabolites (including inosine, AMP, hypoxanthine, ATP) finding
  • The apoptotic secretome enhances thermogenic and adipogenic gene expression in healthy brown adipocytes finding
  • Inosine activates the cAMP-PKA signaling pathway (via p38, Sik2, Crtc3, Creb) to drive the thermogenic program mechanism
  • Inosine treatment increases BAT-dependent energy expenditure in vivo, induces browning of white adipose tissue, and counteracts diet-induced obesity finding
  • ENT1 (SLC29A1) regulates extracellular inosine levels in BAT; ENT1 deficiency increases extracellular inosine and enhances thermogenic adipocyte differentiation mechanism
  • Pharmacological inhibition or genetic ablation of ENT1 in mice enhances BAT activity and counteracts diet-induced obesity finding
  • In human brown adipocytes, ENT1 knockdown/blockade increases extracellular inosine and thermogenic capacity; high ENT1 expression correlates with lower UCP1 in human adipose tissue finding
  • The ENT1 Ile216Thr loss-of-function variant is associated with significantly lower BMI and 59% lower odds of obesity in humans finding
Experimental setups
Assay System Perturbation Readout Platform
untargeted metabolomics murine brown adipocytes (differentiated SVF) nutlin-3 (MDM2 inhibitor, p53-mediated apoptosis) secreted metabolite levels
untargeted metabolomics murine brown adipocytes UV irradiation (caspase-dependent apoptosis) secreted metabolite levels
targeted purine metabolite profiling murine brown adipocytes, endothelial cells, fibroblasts UV irradiation / nutlin-3 extracellular ATP, ADP, AMP, adenosine, inosine, hypoxanthine concentrations
cAMP assay murine brown adipocytes inosine, AMP, hypoxanthine, or forskolin treatment intracellular cAMP levels
phosphoproteomics (high-sensitivity) murine brown adipocytes inosine or forskolin treatment phospho-site regulation across phosphoproteome
indirect calorimetry / metabolic cages mice (WT, A2A-KO, A2B-KO; DIO/HFD models) inosine injection or micro-osmotic pump infusion oxygen consumption, energy expenditure, body weight, body composition
3H-inosine uptake assay murine brown adipocytes (WT vs ENT1-KO) genetic ENT1 knockout radiolabeled inosine uptake and extracellular inosine accumulation
human genetic association study human population (ENT1/SLC29A1 gene) Ile216Thr loss-of-function variant (natural genetic variation) BMI and odds of obesity
Key results
  • Nutlin-3-induced apoptosis significantly enriched 84 metabolites and reduced 13 metabolites, with purine metabolism the most significantly altered pathway 84 up / 13 down
  • UV-induced apoptosis significantly altered the secretome with 50 metabolites up and 19 down, purine pathway again significantly involved 50 up / 19 down
  • UV-induced apoptosis increased extracellular ATP, AMP, inosine and hypoxanthine, with AMP, inosine and hypoxanthine reaching the highest concentrations
  • Supernatant from apoptotic brown adipocytes significantly increased Ucp1, Ppargc1a, Pparg and Fabp4 expression in healthy brown adipocytes
  • Inosine, but not AMP or hypoxanthine, significantly increased intracellular cAMP in brown adipocytes
  • Inosine injection increased oxygen consumption in WT mice; this effect was suppressed in A2A-KO and A2B-KO mice
  • Chronic inosine delivery (micro-osmotic pump) during HFD reduced body weight gain, increased oxygen consumption and cold-induced thermogenesis, and increased UCP1 expression in BAT and WATi
  • ENT1-KO brown adipocytes showed significantly reduced 3H-inosine uptake and correspondingly increased extracellular inosine accumulation
Key statistics
  • count 84 metabolites significantly increased, 13 significantly reduced (nutlin-3-induced apoptosis metabolomics in brown adipocytes)
  • count 50 metabolites up, 19 down (UV-induced apoptosis metabolomics in brown adipocytes)
  • count 330 compounds detected in supernatant (targeted purinergic metabolite screen)
  • count 38,451 phosphopeptides identified; 7,875 (FORSK) and 8,613 (inosine) phospho-sites regulated (FDR<0.05); 2,535 overlapping (phosphoproteomic analysis of inosine vs forskolin treatment)
  • pvalue P < 0.05 (increased TUNEL-positive apoptotic cells in BAT after thermoneutrality (30°C, 3 and 7 days))
  • pvalue P = 0.0592 (trend for reduced fasting blood glucose after 25 days of inosine injection during HFD)
  • other 59% lower odds of obesity (human carriers of ENT1 Ile216Thr Thr variant vs Ile)
  • other 100 µg kg-1 inosine dose (acute inosine injection used to increase oxygen consumption in mice)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper uses primarily murine and cell-culture experiments with small group sizes (n typically 3–16), comparing treatment vs. control (e.g., inosine vs. vehicle, knockout vs. wild-type) using two-tailed Student's t-tests for most pairwise comparisons and one-way ANOVA with Tukey's post-hoc test for comparisons among more than two groups; an ANCOVA was used for one energy-expenditure/body-weight covariate analysis, and FDR-based thresholds were applied to phosphoproteomic site-level statistics. Results are reported as mean ± s.e.m. with significance stars and a note that exact P values are available in source data.

Replicationbiological Sample sizeSample sizes are given per panel in figure legends (e.g., n=3–16 animals, cells, or explants); no formal power calculation is described in the excerpted text Groupstreatment vs. vehicle/control, genotype (WT vs. knockout), diet (control diet vs. HFD), and time points Pairingunclear Randomization/blindingnot stated DispersionSEM Exact p-valuesyes Multiplicity correctionTukey's post hoc test (for one-way ANOVA panels) and an FDR<0.05 threshold (for phosphoproteomic site-level analysis); no correction method is stated for the many separate two-tailed t-tests applied across different panels/comparisons
Statistical tests used
Test Applied to n Assumptions
Two-tailed Student's t-test Fig. 1 panels d, e, h–m, o–q (e.g., extracellular purine levels, Ucp1 expression, oxygen consumption, lipolysis, body weight, thermogenic gene expression) varies by panel, stated as n=3 to n=16 not stated
One-way ANOVA with Tukey's post hoc test Fig. 1 panels a–c, f (TUNEL quantification, metabolomics volcano/enrichment comparisons, intracellular cAMP across treatments) varies by panel (e.g., n=3, n=6) not stated
ANCOVA (non-linear fit) on area under the curve Fig. 1n, oxygen consumption/body weight AUC at 23 °C n=5 not stated
Multiple t-tests Fig. 2f, 14C-triolein lipid uptake across organs in WT vs ENT1-KO mice n=5 not stated
FDR thresholding (FDR<0.05) on regulated phosphopeptides Phosphoproteomic comparison of inosine vs forskolin treatment (Extended Data Fig. 2a) 38,451 phosphopeptides detected; group n not stated in this excerpt not stated
Untargeted metabolomics differential abundance testing (volcano plot) Fig. 1b, nutlin-3-induced apoptosis metabolite changes n=6 not stated
Approaches that could also have been used
  • Many independent panels each use a two-tailed t-test for a single pairwise comparison (e.g., treatment vs. vehicle) across dozens of separate figures/panels.
    Could also: A two-way or repeated-measures ANOVA (or linear mixed model) with a post-hoc correction (e.g., Sidak or Tukey) across related comparisons — Jointly modelling related comparisons (e.g., across time points or genotypes) can account for shared variance structure and also provide a built-in adjustment for multiple comparisons within a related family of tests.
  • Variability is summarized as mean ± s.e.m. throughout, including for panels with small group sizes (e.g., n=3).
    Could also: Reporting standard deviation (SD) or a 95% confidence interval alongside or instead of SEM — SD or CIs directly convey the spread of the underlying data rather than the precision of the mean estimate, which can be informative when group sizes are small.
  • Fig. 2f uses 'multiple t-tests' to compare lipid uptake across several organs between WT and ENT1-KO mice.
    Could also: A two-way ANOVA (organ × genotype) with a post-hoc multiple-comparison correction (e.g., Holm-Sidak or Benjamini-Hochberg) — This approach would test for an organ-by-genotype interaction directly and control the family-wise error rate across the multiple organ comparisons in one framework.
  • Body weight and body composition are compared at discrete time points (e.g., days −1 and 25, or from day 7 onward) using t-tests.
    Could also: A repeated-measures or mixed-effects model that treats the full longitudinal trajectory within the same animals over time — This would use the correlation between repeated measurements on the same animals and could capture the trajectory of weight change over the full study period, not just isolated time points.
  • An explicit multiple-testing correction (FDR) is described for the phosphoproteomics dataset but is not explicitly stated for the metabolomics volcano-plot comparisons (Fig. 1b) despite testing hundreds of metabolites.
    Could also: Applying an explicit false-discovery-rate procedure (e.g., Benjamini-Hochberg) to the metabolomics comparisons as well — Given the large number of metabolites tested simultaneously, an FDR-based approach applied consistently across both omics datasets would extend the same false-positive control already used for the phosphoproteomics data.
  • The human ENT1 variant association with BMI and obesity odds is reported as a significant association without further statistical detail in this excerpt.
    Could also: A covariate-adjusted regression model (e.g., logistic regression adjusted for age, sex, and population structure) — Adjusting for common confounders and stratification variables is a standard complementary approach that can help characterize the robustness of a genetic association across subgroups.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35790189

Paper: Niemann et al. 2022, Nature 609:361–368. "Apoptotic brown adipocytes enhance energy expenditure via extracellular inosine." DOI 10.1038/s41586-022-05041-0, PMCID PMC9452294.

How this paper is structured

Predominantly a wet-lab / in-vivo physiology paper (mouse BAT thermogenesis, energy expenditure, ENT1/SLC29A1 pharmacology & genetics, human cohort genotype– phenotype association). The single substantial pipeline-derived computational dataset with a public accession is the phosphoproteomics (PRIDE PXD032153).

In scope (pipeline-derived, attempted)

result pipeline data status
C1 — 38,451 phosphopeptides identified MaxQuant v2.0.1.0 (EasyPhos, label-free, MBR) PXD032153 MaxQuant output tables attempt from shipped tables
C2 — 7,875 phospho-sites regulated on FORSK (FDR<0.05) Perseus two-sample test on MaxQuant Phospho (STY)Sites PXD032153 attempt (Perseus-equivalent in Python)
C3 — 8,613 phospho-sites regulated on inosine (FDR<0.05) Perseus two-sample test PXD032153 attempt
C4 — 2,535 sites regulated in BOTH intersection of C2 & C3 PXD032153 attempt
C5 — PKA target sites over-represented motif/kinase enrichment (qualitative) PXD032153 attempt (qualitative)

Because PRIDE ships the MaxQuant result tables (Phospho (STY)Sites.txt, proteinGroups.txt, peptides.txt, …), the 80/20 quick win is to recompute the identification count and the differential phospho-site counts directly from those tables — no MaxQuant re-search required (re-search of the 44 GB raw is the optional harder extension).

Out of scope (not attempted; reason)

  • Untargeted metabolomics (Compound Discoverer v3.2, MetaboAnalyst v5.0): no public accession (MetaboLights/Metabolomics Workbench) found in the text or via search → no_data_accession for that modality. Not attempted.
  • RNA-seq of Slc29a1/Slc29a2 expression: mentioned in text; no resolvable GEO accession found → not attempted.
  • All wet-lab / in-vivo / human-cohort results (energy expenditure, qPCR, histology, mouse KO phenotypes, ENT1 Ile216Thr–BMI association): manual/ experimental, not a bioinformatic pipeline → out of scope by brief rule 2.

Key reproducibility caveat

The paper's main-text Methods (as available via PMC/EuropePMC) do not specify the Perseus parameters that drive the regulated-site counts: valid-value filter, imputation width/downshift, t-test S0, and permutation-FDR randomization count. These choices materially affect C2–C4. The identification count (C1) is deterministic from the MaxQuant output and is the cleanest target.

Third-party-tool note (brief rule P16)

The "code" here is the standard MaxQuant + Perseus (EasyPhos) stack, not a bespoke authors' repo. Per the brief this is an equally valid reproduction: run the described tools on the paper's own deposited data with the described parameters.

Figures / tables: Fig 3
C1
Reported
38,451 phosphopeptides identified
Reproduced
38,157 multiplicity-expanded phosphopeptide species (FORSK/Ino/O, rev+cont filtered)
within tolerance
C2
Reported
7,875 phospho-sites regulated on FORSK (FDR<0.05)
Reproduced
3,769 (multiplicity-expanded, class I, full-valid, S0=0, perm-FDR); 3,264 collapsed — ~48% of reported, robust to full parameter sweep
partial
C3
Reported
8,613 phospho-sites regulated on inosine (FDR<0.05)
Reproduced
5,018 (multiplicity-expanded); 3,952 collapsed — ~58% of reported; ordering inosine>FORSK reproduced
partial
C4
Reported
2,535 phospho-sites regulated in BOTH
Reproduced
919 (multiplicity-expanded); 870 collapsed — ~36% of reported
partial
C5
Reported
PKA target sites over-represented among regulated sites
Reproduced
FORSK arm: enriched (basophilic R-x-x-S OR=1.82 p=2e-35; strict R-R-x-S OR=2.54 p=1e-17). Inosine arm: not enriched (OR=0.74 p=1.0)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 57/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🔴3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Total score +8

Identification-level claim C1 reproduced cleanly (38,157 vs 38,451, 0.8% below) from the public, complete PXD032153 raw data, so input identity and comparability are sound. The deviation lives entirely in the downstream statistics: regulated-site counts (C2–C4) come out ~2.5x lower with EasyPhos/Perseus defaults because the paper never publishes the Perseus permutation-FDR parameters, and the reproduction is still preliminary (sweep running). This is best read as a methodology/underspecification gap on the analysis side rather than a fabrication concern — the qualitative conclusion of broad PKA-linked phospho-regulation likely survives, but exact counts remain unconfirmed pending the sweep.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

419.8 k
tokens (I/O) · 30.7 M incl. cache
113 min
runtime · 0.01 CPU-h
0.8 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine