MACanalyzeR scRNAseq analysis tool reveals PPARγHIGH/GDF15HIGH lipid-associated macrophages facilitate thermogenic
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH? Yes -- authors' own public R tool (MACanalyzeR, pinned d5f21f94), data + pre-trained models ship in-repo (brief P16 valid). Compute RAN on «our HPC» (SLURM «job», partition std/n114; conda R 4.4.3 env built on «infra»; prior session's VPN-2FA blocker is gone). OUTCOME = partial. 1:1 vs DIFFERENT: C2 pipeline reproduces 1:1 QUALITATIVELY -- CreateMacObj->FoamSpotteR->MacPolarizeR runs end-to-end on the shipped mono_macs_human.rds (24794x2002) and emits the described categorical outputs (fMAC- 1946/fMAC+ 56; Inflammatory 695/Transitional 855/Healing 452). C1 is a MISMATCH and a possible OVERSTATEMENT: the paper's headline classifier accuracy K=0.9879 is NOT recoverable from the shipped models -- their honest out-of-bag Cohen's kappa is 0.8742 (mouse) / 0.8893 (human), and no statistic stored in the randomForest objects yields 0.9879. It is consistent with a training-set/resubstitution or caret-CV kappa whose artifacts (training matrix / caret train object) are NOT shipped, so 0.9879 is not independently verifiable from what ships -- flagged per brief rule 5. The model is genuinely strong (95%+ OOB accuracy); only the specific reported kappa is in question. WHAT I DID NOT ATTEMPT: the mouse-BAT scRNAseq figures (57% LAMs foamy/27% proliferating; 6 SVF + 4 MonoMac clusters; Pparg/Gdf15 DE; PathAnalyzeR OXPHOS/FA) need raw primary-BAT FASTQ -> CellRanger -> Seurat integration/clustering (heavy multi-stage); and all wet-lab results (flow, qPCR, phenotyping) are non-pipeline. No reproduced value fabricated; every value is recomputable from reproduction/outputs/repro_values.json.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-23no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetMacrophage-mediated mechanisms underlying brown adipose tissue (BAT) dysfunction in obesity are unclear; the paper tests whether a distinct PpargHIGH lipid-associated macrophage (LAM) subset, identified using a new scRNAseq tool (MACanalyzeR), actively supports BAT thermogenic identity via GDF15 secretion rather than merely reflecting tissue degeneration.
- ★ MACanalyzeR is a novel computational scRNAseq framework (with FoamSpotteR, MacPolarizeR, and PathAnalyzeR modules) for comprehensive monocyte/macrophage metabolic and phenotypic profiling method
- ★ Lipid-associated macrophages (LAMs) significantly accumulate in BAT of obese mouse models (db/db and HFD) compared to WT finding
- ★ Unlike db/db BAT LAMs, HFD BAT LAMs correlate with thermogenic gene expression and PPAR signaling pathway activation, showing a healing phenotype finding
- ★ FoamSpotteR, a machine-learning classifier trained on aortic fMAC+/fMAC- macrophages, identifies foamy-like macrophages (fMACs) in obese BAT, predominantly within LAM and proliferating subclusters method
- ★ MacPolarizeR, a machine-learning module trained on M1/M2 bulk RNAseq data, classifies macrophages into Inflammatory, Transitional, and Healing polarization states and was validated on in vitro polarized macrophages and a sepsis WAT dataset method
- ★ A distinct PpargHIGH LAM subcluster progressively accumulates in thermogenically active BAT finding
- ★ Macrophage-specific Pparg depletion disrupts BAT thermogenesis, inducing a white-like phenotype and metabolic dysfunction finding
- ★ PpargHIGH LAMs secrete GDF15, which supports maintenance of brown fat identity and lipogenic characteristics under high-energy demand mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA sequencing | BAT of db/db and HFD mice | diet/genetic obesity model | expression of canonical brown/white fat marker genes | — |
| proteomics | BAT of db/db and HFD mice | diet/genetic obesity model | levels of thermogenic proteins including UCP1 | — |
| scRNAseq | stromal vascular fraction (SVF) from BAT of WT, db/db, and HFD mice | genetic (db/db) and dietary (HFD) obesity | cell clustering, subpopulation abundance, gene expression | — |
| multiplex immunofluorescence | paraffin-embedded BAT tissue sections from LFD and HFD mice | HFD | Adiponectin, UCP1, TREM2, and CD9 staining/co-localization | — |
| scRNAseq (public dataset, GSE97310/GSE116240) | aortic CD45+ (Ptprc) cells from healthy and atherosclerotic mice | atherosclerosis model | fMAC+ vs fMAC- macrophage classification for FoamSpotteR training | — |
| bulk RNA sequencing (public dataset, GSE116239) | foamy vs non-foamy macrophages isolated from atherosclerotic aortas via lipid probe-based flow sorting | none | differentially expressed genes between foamy and non-foamy macrophages | flow cytometry (lipid probe sorting) |
| scRNAseq (public dataset, GSE117176) | bone marrow-derived macrophages (in vitro) | IFN-γ+LPS (M1 induction) or IL-4+IL-13 (M2 induction) | polarization classification validation via MacPolarizeR | — |
| scRNAseq (public dataset, PRJNA626597) | WAT MonoMacs from mice post-sepsis induction | sepsis (acute, 1 day; chronic, 1 month) | macrophage polarization state (M1 vs M2) via MacPolarizeR | — |
- – db/db BAT shows reduced brown adipocyte identity and a more white-like phenotype, while HFD BAT preserves brown fat identity and shows increased thermogenic proteins including UCP1
- – Linear model analysis shows an inverse relationship between macrophage activation and thermogenic function scores in db/db BAT, but a positive correlation in HFD BAT
- ▲ scRNAseq of BAT SVF (n=14,715 cells) identified 6 clusters; MonoMacs was the largest immune cluster and massively increased in both db/db and HFD BAT vs WT
- ▲ MonoMac subclustering (n=5578) identified Monocytes, PVM, LAM, and Proliferating subclusters; LAMs significantly accumulated in db/db and HFD BAT
- ▲ TREM2+ macrophages significantly increased in BAT of HFD-fed mice by immunofluorescence, validating scRNAseq findings
- – FoamSpotteR-predicted fMACs predominantly localize to LAM and proliferating macrophage subclusters 57% (LAM) and 27% (proliferating)
- – Comparison of DEGs between bulk RNAseq fMACs and scRNAseq-predicted fMAC+ LAMs found shared upregulation of lipid metabolism/localization pathways, while BAT-unique fMAC pathways involved metabolite precursor generation, energy production, and vesicle organization (vs. cholesterol metabolism uniquely in atherosclerotic fMACs) 593 vs 193 genes, 73 shared
- – MacPolarizeR correctly identifies inflammatory (M1) shift at 1 day post-sepsis and healing (M2) shift at 1 month post-sepsis in WAT macrophages
- count n = 14,715 (total BAT SVF cells analyzed by scRNAseq across WT, db/db, HFD)
- count n = 5578 (MonoMacs subclustered from BAT SVF scRNAseq)
- count ~4900 cells sampled per condition (approximate cells per condition (WT, db/db, HFD) after quality filtering)
- count 593 upregulated genes (bulk fMAC), 193 (fMAC+ LAM scRNAseq), 73 shared (DEG overlap between bulk RNAseq foam macrophages and scRNAseq-predicted fMAC+ LAMs)
- other 57% and 27% (proportion of fMACs occurring within LAM and proliferating macrophage subclusters, respectively)
- pvalue p-value < 0.05; log2FC > 1 (threshold used for GSEA of GO Biological Process terms on fMAC DEGs)
- pvalue two-side ANOVA test (statistical test for linear model correlating macrophage activation and thermogenesis scores)
- other 75%/25% train/test split (FoamSpotteR model training and validation split of aortic macrophage scRNAseq data)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper combines bulk RNAseq/proteomics comparisons, linear-model correlation analysis of pathway scores, and single-cell RNAseq clustering/sub-clustering with a custom computational tool (MACanalyzeR) that includes machine-learning classifier modules (FoamSpotteR, MacPolarizeR) trained on public reference datasets. Statistical comparisons reported so far include a two-sided ANOVA for a fitted linear model relationship between pathway scores, and Benjamini-Hochberg (BH)-corrected p-values for gene set enrichment analysis (GSEA) of differentially expressed genes. Classifier modules were validated using a 75%/25% train/test split and cross-dataset comparisons rather than classical hypothesis tests.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| two-sided ANOVA | Fig. 1C, linear models relating Macrophage Activation and Adaptive Thermogenesis GO pathway scores in db/db and HFD BAT | — | not stated |
| Benjamini-Hochberg (BH) correction on GSEA p-values | Fig. 2E–G, GO Biological Process gene set enrichment of DEGs shared/unique between bulk RNAseq foam macrophages and FoamSpotteR-predicted fMAC+ cells | DEGs filtered at p-value < 0.05 and log2FC > 1 | not stated |
| Machine-learning classifier (train/test split) for cell-state prediction | FoamSpotteR module trained on aortic CD45+ macrophage scRNAseq data (75% train / 25% test) and applied to BAT macrophages | trained on aortic macrophage scRNAseq dataset; test performance on held-out 25% | not stated |
| Machine-learning-based polarization scoring/clustering | MacPolarizeR module, classifying BAT macrophages into Inflammatory/Transitional/Healing states using genes from a public M1/M2 bulk RNAseq dataset (GSE129253), validated against GSE117176 and PRJNA626597 | — | not stated |
-
The relationship between macrophage-activation and thermogenesis pathway scores was assessed by fitting a linear model and testing it with a two-sided ANOVA.↳ Could also: A correlation coefficient (Pearson or Spearman) with a reported confidence interval could also be used to summarize the association — Reporting a correlation coefficient alongside its CI conveys both the strength and precision of the association, complementing the significance test from the ANOVA on the fitted model
-
Differentially expressed gene sets for GSEA were filtered using fixed thresholds (p-value < 0.05, log2FC > 1) prior to Benjamini-Hochberg correction.↳ Could also: A pre-ranked GSEA approach using the full ranked gene list (e.g., by test statistic) without a hard DEG cutoff could also be used — Pre-ranked GSEA retains information from genes with subtler but consistent changes across the whole transcriptome, which can complement threshold-based enrichment analysis
-
The FoamSpotteR and MacPolarizeR classifiers were trained and evaluated with a single 75%/25% train/test split.↳ Could also: k-fold cross-validation could also be used to estimate classifier performance — Cross-validation uses all available labeled cells for both training and testing across folds, which can provide a more stable estimate of classifier accuracy, especially with modestly sized reference datasets
-
Cell-count changes across conditions (e.g., LAM or fMAC+ proportions) are presented as bar plots/proportions across WT, db/db, and HFD.↳ Could also: A compositional data analysis method (e.g., scCODA or a beta-binomial/Dirichlet-multinomial model) could also be applied to formally test shifts in cell-type proportions — Proportion-based compositional methods explicitly account for the non-independence of cell-type fractions summing to one, which can complement descriptive bar-plot comparisons of cluster abundance
-
The scRNAseq design pools cells per condition (WT, db/db, HFD) for clustering and downstream comparisons.↳ Could also: A pseudobulk aggregation across multiple biological replicates per condition, analyzed with tools such as DESeq2 or edgeR, could also be used for differential expression testing — Pseudobulk approaches that incorporate biological replicate variability are often used alongside single-cell clustering to formally test for condition differences while accounting for animal-to-animal variation
-
Dispersion measures and exact p-values for quantitative comparisons are not detailed in the portion of text reviewed.↳ Could also: Reporting exact p-values together with a chosen dispersion measure (SD, SEM, or 95% CI) for each comparison could also be included — Exact values and a stated dispersion measure allow readers to gauge effect magnitude and variability directly, complementing significance labels alone
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-40450001 (MACanalyzeR)
- Title: MACanalyzeR scRNAseq analysis tool reveals PPARγ^HIGH/GDF15^HIGH lipid-associated macrophages facilitate thermogenic… (Nat Commun 2025)
- DOI: 10.1038/s41467-025-60295-2 · PMCID PMC12126529 · PMID 40450001
- Code: https://github.com/andreeedna/MACanalyzeR — public, NOT archived, no SPDX license file ("License: tSNEland" joke in DESCRIPTION), R package v1.0.0/1.0.1, last push 2025-05-23, commit
d5f21f94f19adf1179afb746a6f5b1f45f057033. - Authors' OWN tool (Andrea Ninni = andreeedna = first author). Per BRIEF rule P16, applying this tool to its data is fully valid.
What ships in the repo (zero-compute inspection via GitHub API)
- R/ source for all modules (FoamSpotteR, MacPolarizeR/macSpectrum, PathwayAnalysis, MitoScanneR).
- Pre-trained models / references in
data/:mouse_LAMprey.rda(5.0 MB),human_LAMprey.rda(12.7 MB) — the FoamSpotteR randomForest classifiers (object has$importance, used viapredict(FoamModel, x, type="prob")).M0/M1/M2_mean_new.rda,con/foam_mean_new.rda— MacPolarizeR reference centroids.mono_macs_human.rds(53.7 MB) — shipped example Seurat object (human mono/macrophages).- GSEA/GO/KEGG/Reactome gene sets for PathAnalyzeR.
- Vignette
vignette/MACanalyzeR_Vignette.md(refers to avignette/mac.rdsthat is NOT in the tree → vignette example file missing; the shippeddata/mono_macs_human.rdsis the usable example).
Pipeline tools (Methods)
CellRanger v7.1.0 → Seurat v5.0.3 (R 4.4.0): normalize, 2000 HVG, CCA integration, PCA, clustering, UMAP/tSNE, Wilcoxon DE. Then MACanalyzeR modules (FoamSpotteR RF, MacPolarizeR, PathAnalyzeR, MitoScanneR).
In scope (attempted) — pipeline-derived, clearly specified, data ships
| # | Reported result | Paper location | Pipeline | Why in scope |
|---|---|---|---|---|
| C1 | FoamSpotteR final RF accuracy K = 0.9879 (Cohen's kappa) | Results, verbatim "The final RF showed a high level of accuracy, i.e., K = 0.9879." | randomForest/caret model shipped as mouse_LAMprey |
Model ships pre-trained; kappa recoverable from its confusion/CV. Cleanest 1:1 numeric claim. |
| C2 | MACanalyzeR pipeline runs and yields fMAC± (FoamSpotteR) + M0/M1/M2 (MacPolarizeR) classifications as described | Methods/Results + vignette | MACanalyzeR on shipped mono_macs_human.rds |
Demonstrates the tool reproduces its described categorical outputs on its own shipped data. |
Out of scope (NOT attempted) — the hard ~20%, per 80/20 rule
- 57% of LAMs foamy / 27% proliferating (Fig — FoamSpotteR on mouse BAT): requires the primary BAT scRNAseq (~14,715 cells) → CellRanger from FASTQ → Seurat integration → clustering, an expensive full re-run; primary BAT accession not the brief's GSE116239 (that is the bulk atherosclerosis training set). Skipped: heavy, multi-stage, raw-FASTQ.
- 6 SVF clusters / 4 MonoMac subclusters, Pparg/Gdf15 DE values, PathAnalyzeR OXPHOS/FA scores on mouse BAT — all depend on the same full BAT re-clustering. Skipped for same reason.
- All wet-lab results (flow cytometry, qPCR, mouse phenotyping) — out of scope (non-pipeline).
Accessions named in paper
GSE116239 (bulk foamy/non-foamy aorta macs — FoamSpotteR training), GSE97310, GSE116240, GSE129253, GSE117176, PRJNA626597, GSE182233, GSE207706, GSE111105. Primary BAT scRNAseq accession in (truncated) Data-availability — not needed for C1/C2.
Eligibility: eligible
Resolvable public code (own tool) + data ships inside the repo + an identifiable exact expected value (K=0.9879). Compute on «our HPC».
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The reproduction is split: the tool pipeline (C2) reproduces 1:1 qualitatively on its shipped example data, emitting the described fMAC±/polarization categories. The headline classifier accuracy (C1, K=0.9879) is not derivable from the shipped pre-trained models — their honest out-of-bag Cohen's kappa is 0.8742/0.8893 and no stored statistic yields 0.9879. The gap is on the authors' side / data-availability: the training matrix or caret object needed to verify the reported kappa was never deposited, and the reported value sits materially above the model's own honest estimate, making it a possible overstatement. Severity is moderate, not critical — the classifier is genuinely strong (95%+ OOB accuracy) and the qualitative tool behaviour holds — so overall solid-with-explainable-deviation, with the C1 number clearly flagged.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.