Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

MACanalyzeR scRNAseq analysis tool reveals PPARγHIGH/GDF15HIGH lipid-associated macrophages facilitate thermogenic

Nat Commun · 2025
L1 70/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
70/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 37% of all assessed papers rank 732 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH? Yes -- authors' own public R tool (MACanalyzeR, pinned d5f21f94), data + pre-trained models ship in-repo (brief P16 valid). Compute RAN on «our HPC» (SLURM «job», partition std/n114; conda R 4.4.3 env built on «infra»; prior session's VPN-2FA blocker is gone). OUTCOME = partial. 1:1 vs DIFFERENT: C2 pipeline reproduces 1:1 QUALITATIVELY -- CreateMacObj->FoamSpotteR->MacPolarizeR runs end-to-end on the shipped mono_macs_human.rds (24794x2002) and emits the described categorical outputs (fMAC- 1946/fMAC+ 56; Inflammatory 695/Transitional 855/Healing 452). C1 is a MISMATCH and a possible OVERSTATEMENT: the paper's headline classifier accuracy K=0.9879 is NOT recoverable from the shipped models -- their honest out-of-bag Cohen's kappa is 0.8742 (mouse) / 0.8893 (human), and no statistic stored in the randomForest objects yields 0.9879. It is consistent with a training-set/resubstitution or caret-CV kappa whose artifacts (training matrix / caret train object) are NOT shipped, so 0.9879 is not independently verifiable from what ships -- flagged per brief rule 5. The model is genuinely strong (95%+ OOB accuracy); only the specific reported kappa is in question. WHAT I DID NOT ATTEMPT: the mouse-BAT scRNAseq figures (57% LAMs foamy/27% proliferating; 6 SVF + 4 MonoMac clusters; Pparg/Gdf15 DE; PathAnalyzeR OXPHOS/FA) need raw primary-BAT FASTQ -> CellRanger -> Seurat integration/clustering (heavy multi-stage); and all wet-lab results (flow, qPCR, phenotyping) are non-pipeline. No reproduced value fabricated; every value is recomputable from reproduction/outputs/repro_values.json.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-23
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Macrophage-mediated mechanisms underlying brown adipose tissue (BAT) dysfunction in obesity are unclear; the paper tests whether a distinct PpargHIGH lipid-associated macrophage (LAM) subset, identified using a new scRNAseq tool (MACanalyzeR), actively supports BAT thermogenic identity via GDF15 secretion rather than merely reflecting tissue degeneration.

Core claims
  • MACanalyzeR is a novel computational scRNAseq framework (with FoamSpotteR, MacPolarizeR, and PathAnalyzeR modules) for comprehensive monocyte/macrophage metabolic and phenotypic profiling method
  • Lipid-associated macrophages (LAMs) significantly accumulate in BAT of obese mouse models (db/db and HFD) compared to WT finding
  • Unlike db/db BAT LAMs, HFD BAT LAMs correlate with thermogenic gene expression and PPAR signaling pathway activation, showing a healing phenotype finding
  • FoamSpotteR, a machine-learning classifier trained on aortic fMAC+/fMAC- macrophages, identifies foamy-like macrophages (fMACs) in obese BAT, predominantly within LAM and proliferating subclusters method
  • MacPolarizeR, a machine-learning module trained on M1/M2 bulk RNAseq data, classifies macrophages into Inflammatory, Transitional, and Healing polarization states and was validated on in vitro polarized macrophages and a sepsis WAT dataset method
  • A distinct PpargHIGH LAM subcluster progressively accumulates in thermogenically active BAT finding
  • Macrophage-specific Pparg depletion disrupts BAT thermogenesis, inducing a white-like phenotype and metabolic dysfunction finding
  • PpargHIGH LAMs secrete GDF15, which supports maintenance of brown fat identity and lipogenic characteristics under high-energy demand mechanism
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA sequencing BAT of db/db and HFD mice diet/genetic obesity model expression of canonical brown/white fat marker genes
proteomics BAT of db/db and HFD mice diet/genetic obesity model levels of thermogenic proteins including UCP1
scRNAseq stromal vascular fraction (SVF) from BAT of WT, db/db, and HFD mice genetic (db/db) and dietary (HFD) obesity cell clustering, subpopulation abundance, gene expression
multiplex immunofluorescence paraffin-embedded BAT tissue sections from LFD and HFD mice HFD Adiponectin, UCP1, TREM2, and CD9 staining/co-localization
scRNAseq (public dataset, GSE97310/GSE116240) aortic CD45+ (Ptprc) cells from healthy and atherosclerotic mice atherosclerosis model fMAC+ vs fMAC- macrophage classification for FoamSpotteR training
bulk RNA sequencing (public dataset, GSE116239) foamy vs non-foamy macrophages isolated from atherosclerotic aortas via lipid probe-based flow sorting none differentially expressed genes between foamy and non-foamy macrophages flow cytometry (lipid probe sorting)
scRNAseq (public dataset, GSE117176) bone marrow-derived macrophages (in vitro) IFN-γ+LPS (M1 induction) or IL-4+IL-13 (M2 induction) polarization classification validation via MacPolarizeR
scRNAseq (public dataset, PRJNA626597) WAT MonoMacs from mice post-sepsis induction sepsis (acute, 1 day; chronic, 1 month) macrophage polarization state (M1 vs M2) via MacPolarizeR
Key results
  • db/db BAT shows reduced brown adipocyte identity and a more white-like phenotype, while HFD BAT preserves brown fat identity and shows increased thermogenic proteins including UCP1
  • Linear model analysis shows an inverse relationship between macrophage activation and thermogenic function scores in db/db BAT, but a positive correlation in HFD BAT
  • scRNAseq of BAT SVF (n=14,715 cells) identified 6 clusters; MonoMacs was the largest immune cluster and massively increased in both db/db and HFD BAT vs WT
  • MonoMac subclustering (n=5578) identified Monocytes, PVM, LAM, and Proliferating subclusters; LAMs significantly accumulated in db/db and HFD BAT
  • TREM2+ macrophages significantly increased in BAT of HFD-fed mice by immunofluorescence, validating scRNAseq findings
  • FoamSpotteR-predicted fMACs predominantly localize to LAM and proliferating macrophage subclusters 57% (LAM) and 27% (proliferating)
  • Comparison of DEGs between bulk RNAseq fMACs and scRNAseq-predicted fMAC+ LAMs found shared upregulation of lipid metabolism/localization pathways, while BAT-unique fMAC pathways involved metabolite precursor generation, energy production, and vesicle organization (vs. cholesterol metabolism uniquely in atherosclerotic fMACs) 593 vs 193 genes, 73 shared
  • MacPolarizeR correctly identifies inflammatory (M1) shift at 1 day post-sepsis and healing (M2) shift at 1 month post-sepsis in WAT macrophages
Key statistics
  • count n = 14,715 (total BAT SVF cells analyzed by scRNAseq across WT, db/db, HFD)
  • count n = 5578 (MonoMacs subclustered from BAT SVF scRNAseq)
  • count ~4900 cells sampled per condition (approximate cells per condition (WT, db/db, HFD) after quality filtering)
  • count 593 upregulated genes (bulk fMAC), 193 (fMAC+ LAM scRNAseq), 73 shared (DEG overlap between bulk RNAseq foam macrophages and scRNAseq-predicted fMAC+ LAMs)
  • other 57% and 27% (proportion of fMACs occurring within LAM and proliferating macrophage subclusters, respectively)
  • pvalue p-value < 0.05; log2FC > 1 (threshold used for GSEA of GO Biological Process terms on fMAC DEGs)
  • pvalue two-side ANOVA test (statistical test for linear model correlating macrophage activation and thermogenesis scores)
  • other 75%/25% train/test split (FoamSpotteR model training and validation split of aortic macrophage scRNAseq data)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper combines bulk RNAseq/proteomics comparisons, linear-model correlation analysis of pathway scores, and single-cell RNAseq clustering/sub-clustering with a custom computational tool (MACanalyzeR) that includes machine-learning classifier modules (FoamSpotteR, MacPolarizeR) trained on public reference datasets. Statistical comparisons reported so far include a two-sided ANOVA for a fitted linear model relationship between pathway scores, and Benjamini-Hochberg (BH)-corrected p-values for gene set enrichment analysis (GSEA) of differentially expressed genes. Classifier modules were validated using a 75%/25% train/test split and cross-dataset comparisons rather than classical hypothesis tests.

Replicationunclear Sample sizescRNAseq: 14,715 total SVF cells (~4900 cells per condition: WT, db/db, HFD); MonoMacs sub-cluster n = 5578 cells; number of biological animals/replicates per condition not stated in this excerpt GroupsWT vs db/db vs HFD mouse BAT (and LFD vs HFD in imaging); fMAC+ vs fMAC- macrophages; M1 vs M2 vs M0 polarization states Pairingunclear Randomization/blindingnot stated Dispersionunclear Multiplicity correctionBenjamini-Hochberg (BH)
Statistical tests used
Test Applied to n Assumptions
two-sided ANOVA Fig. 1C, linear models relating Macrophage Activation and Adaptive Thermogenesis GO pathway scores in db/db and HFD BAT not stated
Benjamini-Hochberg (BH) correction on GSEA p-values Fig. 2E–G, GO Biological Process gene set enrichment of DEGs shared/unique between bulk RNAseq foam macrophages and FoamSpotteR-predicted fMAC+ cells DEGs filtered at p-value < 0.05 and log2FC > 1 not stated
Machine-learning classifier (train/test split) for cell-state prediction FoamSpotteR module trained on aortic CD45+ macrophage scRNAseq data (75% train / 25% test) and applied to BAT macrophages trained on aortic macrophage scRNAseq dataset; test performance on held-out 25% not stated
Machine-learning-based polarization scoring/clustering MacPolarizeR module, classifying BAT macrophages into Inflammatory/Transitional/Healing states using genes from a public M1/M2 bulk RNAseq dataset (GSE129253), validated against GSE117176 and PRJNA626597 not stated
Approaches that could also have been used
  • The relationship between macrophage-activation and thermogenesis pathway scores was assessed by fitting a linear model and testing it with a two-sided ANOVA.
    Could also: A correlation coefficient (Pearson or Spearman) with a reported confidence interval could also be used to summarize the association — Reporting a correlation coefficient alongside its CI conveys both the strength and precision of the association, complementing the significance test from the ANOVA on the fitted model
  • Differentially expressed gene sets for GSEA were filtered using fixed thresholds (p-value < 0.05, log2FC > 1) prior to Benjamini-Hochberg correction.
    Could also: A pre-ranked GSEA approach using the full ranked gene list (e.g., by test statistic) without a hard DEG cutoff could also be used — Pre-ranked GSEA retains information from genes with subtler but consistent changes across the whole transcriptome, which can complement threshold-based enrichment analysis
  • The FoamSpotteR and MacPolarizeR classifiers were trained and evaluated with a single 75%/25% train/test split.
    Could also: k-fold cross-validation could also be used to estimate classifier performance — Cross-validation uses all available labeled cells for both training and testing across folds, which can provide a more stable estimate of classifier accuracy, especially with modestly sized reference datasets
  • Cell-count changes across conditions (e.g., LAM or fMAC+ proportions) are presented as bar plots/proportions across WT, db/db, and HFD.
    Could also: A compositional data analysis method (e.g., scCODA or a beta-binomial/Dirichlet-multinomial model) could also be applied to formally test shifts in cell-type proportions — Proportion-based compositional methods explicitly account for the non-independence of cell-type fractions summing to one, which can complement descriptive bar-plot comparisons of cluster abundance
  • The scRNAseq design pools cells per condition (WT, db/db, HFD) for clustering and downstream comparisons.
    Could also: A pseudobulk aggregation across multiple biological replicates per condition, analyzed with tools such as DESeq2 or edgeR, could also be used for differential expression testing — Pseudobulk approaches that incorporate biological replicate variability are often used alongside single-cell clustering to formally test for condition differences while accounting for animal-to-animal variation
  • Dispersion measures and exact p-values for quantitative comparisons are not detailed in the portion of text reviewed.
    Could also: Reporting exact p-values together with a chosen dispersion measure (SD, SEM, or 95% CI) for each comparison could also be included — Exact values and a stated dispersion measure allow readers to gauge effect magnitude and variability directly, complementing significance labels alone
Software: MACanalyzeR (custom scRNAseq analysis framework developed in this study, including FoamSpotteR and MacPolarizeR modules)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40450001 (MACanalyzeR)

  • Title: MACanalyzeR scRNAseq analysis tool reveals PPARγ^HIGH/GDF15^HIGH lipid-associated macrophages facilitate thermogenic… (Nat Commun 2025)
  • DOI: 10.1038/s41467-025-60295-2 · PMCID PMC12126529 · PMID 40450001
  • Code: https://github.com/andreeedna/MACanalyzeR — public, NOT archived, no SPDX license file ("License: tSNEland" joke in DESCRIPTION), R package v1.0.0/1.0.1, last push 2025-05-23, commit d5f21f94f19adf1179afb746a6f5b1f45f057033.
  • Authors' OWN tool (Andrea Ninni = andreeedna = first author). Per BRIEF rule P16, applying this tool to its data is fully valid.

What ships in the repo (zero-compute inspection via GitHub API)

  • R/ source for all modules (FoamSpotteR, MacPolarizeR/macSpectrum, PathwayAnalysis, MitoScanneR).
  • Pre-trained models / references in data/:
    • mouse_LAMprey.rda (5.0 MB), human_LAMprey.rda (12.7 MB) — the FoamSpotteR randomForest classifiers (object has $importance, used via predict(FoamModel, x, type="prob")).
    • M0/M1/M2_mean_new.rda, con/foam_mean_new.rda — MacPolarizeR reference centroids.
    • mono_macs_human.rds (53.7 MB) — shipped example Seurat object (human mono/macrophages).
    • GSEA/GO/KEGG/Reactome gene sets for PathAnalyzeR.
  • Vignette vignette/MACanalyzeR_Vignette.md (refers to a vignette/mac.rds that is NOT in the tree → vignette example file missing; the shipped data/mono_macs_human.rds is the usable example).

Pipeline tools (Methods)

CellRanger v7.1.0 → Seurat v5.0.3 (R 4.4.0): normalize, 2000 HVG, CCA integration, PCA, clustering, UMAP/tSNE, Wilcoxon DE. Then MACanalyzeR modules (FoamSpotteR RF, MacPolarizeR, PathAnalyzeR, MitoScanneR).

In scope (attempted) — pipeline-derived, clearly specified, data ships

# Reported result Paper location Pipeline Why in scope
C1 FoamSpotteR final RF accuracy K = 0.9879 (Cohen's kappa) Results, verbatim "The final RF showed a high level of accuracy, i.e., K = 0.9879." randomForest/caret model shipped as mouse_LAMprey Model ships pre-trained; kappa recoverable from its confusion/CV. Cleanest 1:1 numeric claim.
C2 MACanalyzeR pipeline runs and yields fMAC± (FoamSpotteR) + M0/M1/M2 (MacPolarizeR) classifications as described Methods/Results + vignette MACanalyzeR on shipped mono_macs_human.rds Demonstrates the tool reproduces its described categorical outputs on its own shipped data.

Out of scope (NOT attempted) — the hard ~20%, per 80/20 rule

  • 57% of LAMs foamy / 27% proliferating (Fig — FoamSpotteR on mouse BAT): requires the primary BAT scRNAseq (~14,715 cells) → CellRanger from FASTQ → Seurat integration → clustering, an expensive full re-run; primary BAT accession not the brief's GSE116239 (that is the bulk atherosclerosis training set). Skipped: heavy, multi-stage, raw-FASTQ.
  • 6 SVF clusters / 4 MonoMac subclusters, Pparg/Gdf15 DE values, PathAnalyzeR OXPHOS/FA scores on mouse BAT — all depend on the same full BAT re-clustering. Skipped for same reason.
  • All wet-lab results (flow cytometry, qPCR, mouse phenotyping) — out of scope (non-pipeline).

Accessions named in paper

GSE116239 (bulk foamy/non-foamy aorta macs — FoamSpotteR training), GSE97310, GSE116240, GSE129253, GSE117176, PRJNA626597, GSE182233, GSE207706, GSE111105. Primary BAT scRNAseq accession in (truncated) Data-availability — not needed for C1/C2.

Eligibility: eligible

Resolvable public code (own tool) + data ships inside the repo + an identifiable exact expected value (K=0.9879). Compute on «our HPC».

C1
Reported
FoamSpotteR final RF accuracy K = 0.9879
Reproduced
OOB Cohen's kappa 0.8742 (mouse_LAMprey, n=4976) / 0.8893 (human_LAMprey, n=10765); OOB accuracy 0.9584 / 0.9473
did not match
C2a
Reported
FoamSpotteR labels cells fMAC+/fMAC- on shipped example data
Reproduced
ran end-to-end: fMAC- 1946 / fMAC+ 56 on mono_macs_human.rds (24794 genes x 2002 cells)
exact
C2b
Reported
MacPolarizeR classifies cells into Inflammatory/Transitional/Healing
Reproduced
Inflammatory 695 / Transitional 855 / Healing 452
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 70/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴

The reproduction is split: the tool pipeline (C2) reproduces 1:1 qualitatively on its shipped example data, emitting the described fMAC±/polarization categories. The headline classifier accuracy (C1, K=0.9879) is not derivable from the shipped pre-trained models — their honest out-of-bag Cohen's kappa is 0.8742/0.8893 and no stored statistic yields 0.9879. The gap is on the authors' side / data-availability: the training matrix or caret object needed to verify the reported kappa was never deposited, and the reported value sits materially above the model's own honest estimate, making it a possible overstatement. Severity is moderate, not critical — the classifier is genuinely strong (95%+ OOB accuracy) and the qualitative tool behaviour holds — so overall solid-with-explainable-deviation, with the C1 number clearly flagged.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

265 k
tokens (I/O) · 17.8 M incl. cache
60 min
runtime · 0.03 CPU-h
2.1 GB
peak RAM
1
HPC jobs
hummel
machine