Macrophages on the run: Exercise balances macrophage polarization for improved health.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡A deviation arose in the data or preprocessing
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce 1:1. The paper's central computational result — SupplementaryTable1's per-dataset ROC-AUC + Welch p-value of the BoNE M1/M2 macrophage-polarization score across ~75 public exercise datasets (the numbers behind Figures 4 & 5) — was reproduced by reusing the authors' OWN scoring code (bone.py getRanks2/mergeRanks/getMetrics/getScores) and their exact dataset-grouping definitions UNCHANGED, swapping only the data backend from the lab's non-public local Hegemon files to the lab's PUBLIC Hegemon web service (the same endpoint the repo's Download_Data.ipynb uses). Of 68 gradeable dataset/grouping constructions: 42 reproduce the reported AUC to the exact 2nd decimal, 12 more within 0.02, 11 within 0.05; median |dAUC|=0.000; p-values match to ~2-3 sig figs where AUC is exact (e.g. GSE221210 P=0.0107 and 0.00387 identical). This is a faithful 1:1 reproduction with no fabrication signal — the reported numbers are re-derivable from the shipped code + the lab's public data. NOT attempted (the hard ~20%): re-deriving the M1/M2 signature gene clusters themselves (out of scope; shipped gene lists consumed); 9 datasets errored on the public data path (7 gene→probe mapping gaps on RNA-seq platforms, 2 server HTTP 500) — coverage gaps, not disagreements; 3 small mismatches (incl. a tiny n=3/3 series). Compute on «our HPC» SLURM «job» (COMPLETED 5:37). Repo pinned at 4cf2f16.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 81assessed: 2026-06-14 ⛓ 55d70774b29b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetExercise appears to trigger both pro-inflammatory (M1) and reparative (M2) macrophage activation in the literature; this paper tests whether a temporally-resolved Boolean gene-expression model of macrophage polarization can reconcile this apparent paradox across a large, diverse meta-analysis of exercise and immobilization datasets.
- ★ Immediate/acute exercise triggers an M1 (pro-inflammatory) macrophage polarization surge. finding
- ★ Long-term exercise training leads to sustained M2 (reparative) macrophage activation. finding
- ★ Immobilization has the opposite effect of exercise, triggering immediate M2 activation. finding
- ★ These temporal M1/M2 patterns are consistent across species (human vs mouse), sampling method (blood vs muscle biopsy), and exercise type (resistance vs endurance). finding
- ★ Gender, exercise intensity, and age modulate the degree/strength of macrophage polarization but not the overall directional pattern. finding
- ★ A focused 25-gene signature (3 M1, 22 M2) can predict pre- vs post-exercise and trained vs control status specifically in muscle tissue. resource
- Boolean implication/relationship analysis increases signal-to-noise across heterogeneous datasets and reveals gene relationships missed by standard differential expression analysis. method
- ★ Compiled a database of 75 exercise/immobilization gene expression datasets (7000+ samples) spanning microarray, RNA-seq, and scRNA-seq. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Meta-analysis of microarray/RNA-seq/scRNA-seq (Boolean macrophage polarization scoring) | human, mouse, rat (blood and muscle biopsy samples) | acute exercise, long-term training (resistance/endurance), immobilization | macrophage polarization ('macrophage') score (M1 vs M2) | — |
| StepMiner algorithm / Boolean implication (BooleanNet) analysis | gene expression matrices from 75 compiled datasets | none (computational method applied to existing data) | binarized (high/low) gene expression thresholds and Boolean relationships between gene pairs | — |
| Gene expression normalization (RMA for microarray, CPM/RPKM for RNA-seq) | Affymetrix microarray and RNA-Seq platforms from GEO | none | log2(CPM+1) or log2(RPKM+1) normalized expression values | Affymetrix (RMA); RNA-seq (CPM/RPKM) |
| Gene signature validation (ROC-AUC classification) | pooled macrophage dataset (GSE134312, 170 samples) | M1 vs M2 macrophage state | ROC-AUC for M1/M2 classification using new gene set model groups | — |
| Exercise-prediction validation of muscle-resident macrophage signature | 20 muscle tissue datasets from the 75-dataset database | exercise vs control/pre-exercise | ROC-AUC for pre- vs post-exercise / trained vs control classification | — |
| Tissue-comparison of macrophage gene groups | spleen vs heart/skeletal muscle (GSE142068, GSE230102, GSE59927) and exercise dataset GSE224146 | tissue type / exercise | macrophage gene group scores compared across tissues | — |
| Disease-context macrophage scoring (M1-only genes) | GSE27536 and GSE14798 (disease-status samples) | disease state | macrophage score using only M1-associated genes | — |
- – Modeling revealed a temporal dynamic: exercise triggers an immediate M1 surge, while long-term training transitions to sustained M2 activation.
- – Patterns were consistent across species, sampling methods, and exercise type, and were routinely statistically significant.
- ▲ Immobilization triggered an immediate M2 activation, opposite to exercise.
- – Gender, exercise intensity, and age affected the degree of polarization without changing overall patterns.
- – Identified a 25-gene muscle-resident macrophage polarization signature (3 M1, 22 M2) predictive of exercise status. ROC-AUC |0.7|
- – Boolean macrophage polarization model path of 338 genes accurately separated M1 vs M2 macrophage samples in training and validation datasets. 338 genes (M1=48, M2=290)
- count 75 datasets (compiled exercise and immobilization gene expression datasets)
- count 7000+ samples (total samples across the 75-dataset meta-analysis)
- count 338 genes (M1=48, M2=290) (Boolean pathway gene set separating M1 vs M2 macrophages)
- count 25 genes (3 M1, 22 M2) (muscle-resident macrophage polarization signature)
- other ROC-AUC |0.7| (accuracy threshold for including genes in the new muscle signature model)
- count 170 pooled macrophage samples (GSE134312 dataset used to validate M1/M2 prediction)
- other 2-fold change noise margin (±0.5 around StepMiner threshold) (StepMiner/Boolean binarization noise margin)
- count 20 muscle tissue datasets (subset of the 75 datasets used to validate exercise prediction)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper performs a computational meta-analysis of 75 published gene expression datasets (7000+ samples, spanning microarray, RNA-seq, and scRNA-seq) by applying a Boolean logic framework to binarize gene expression and a pre-built Boolean macrophage polarization network (338 genes) to derive a continuous 'macrophage score' per sample. Standard t-tests (Python) are used to compare scores between pre- and post-exercise conditions within each dataset, and ROC-AUC is used both for gene selection in the muscle-specific signature and for classification validation. Results are reported across subgroups stratified by species, tissue, and exercise modality.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| StepMiner F-statistic (adaptive regression) | Per-gene expression binarization threshold detection across all datasets | varies by dataset; 7000+ samples total across all 75 datasets | not stated |
| BooleanNet sparsity statistics (S_ij, p_ij) | Identification of Boolean implication relationships between gene pairs | varies by dataset | not stated |
| Standard two-sample t-test | Comparison of macrophage polarization scores between pre- and post-exercise sample groups within each dataset | per-dataset n; aggregate n stated as 7000+ samples across 75 datasets | not stated |
| ROC-AUC (Receiver Operating Characteristic Area Under the Curve) | Gene selection for muscle-resident macrophage signature (|AUC| ≥ 0.7 threshold) and validation of the gene set model on a pooled macrophage dataset (GSE134312, n=170) and 20 muscle tissue datasets | 170 pooled macrophage samples for M1/M2 validation; 20 muscle datasets for exercise validation | na |
-
Gene expression was binarized (high/low) via StepMiner thresholds and analyzed using Boolean implication relationships to aggregate signal across heterogeneous datasets↳ Could also: Gene Set Enrichment Analysis (GSEA) or single-sample GSEA (ssGSEA) applied to M1/M2 gene sets could also aggregate polarization signal continuously across diverse datasets without binarization — Continuous enrichment scoring preserves quantitative variation within the high/low bins and has established null distributions and FDR procedures; comparing it to Boolean scores would clarify how much information the binarization step retains or discards
-
Per-dataset standard t-tests were used to compare pre- vs post-exercise macrophage scores, with results described as 'routinely showing statistically significant results' across 75 datasets↳ Could also: A formal random-effects meta-analysis (e.g., DerSimonian–Laird or restricted maximum likelihood) pooling standardized effect sizes (Cohen's d or Hedges' g) across datasets could also synthesize evidence — Random-effects meta-analysis explicitly models between-study heterogeneity, yields a single pooled effect estimate with a confidence interval, and produces I² to quantify consistency — all of which complement the Boolean consistency counts the paper reports
-
A binary M1/M2 macrophage polarization model with a single composite score was applied; macrophage state is described in the text as a spectrum rather than discrete states↳ Could also: Computational deconvolution methods (e.g., CIBERSORT, MuSiC) or multi-state transcriptional scoring (e.g., a gradient score along a continuum) could also quantify macrophage activation spectra from bulk RNA data — Spectrum-based or deconvolution approaches can capture intermediate or mixed polarization states without forcing samples into two categories, which may be informative given the paper's own acknowledgment that M1/M2 is a simplification
-
A ROC-AUC threshold of |0.7| on a single training dataset (one muscle exercise dataset) was used to select genes for the muscle-specific macrophage signature↳ Could also: Cross-validated feature selection (e.g., leave-one-out or k-fold cross-validation combined with LASSO or elastic-net regularization) across multiple muscle datasets could also identify a stable gene signature — Cross-validated selection reduces the risk that chosen genes are optimized to a specific dataset; reporting AUC confidence intervals across held-out folds would also convey signature stability
-
Multiple t-tests are applied across 75 independent dataset comparisons without a stated procedure to adjust for the number of tests↳ Could also: Applying a Benjamini-Hochberg FDR correction across all per-dataset p-values, or reporting the proportion of datasets exceeding a significance threshold, could also characterize consistency of the exercise effect — With 75 simultaneous tests, a small fraction of significant results could arise by chance at an uncorrected α; FDR adjustment or a sign-test on the direction of effects would make the consistency claim more formally quantified
-
Macrophage polarization scores are described as higher or lower between conditions, but the available text does not report a measure of spread (SD, SEM, or CI) around the score distributions↳ Could also: Reporting 95% confidence intervals or standard deviations alongside point estimates of the macrophage score difference would also convey precision, especially for subgroups with few datasets — For small subgroup comparisons (e.g., a single species or modality), effect estimates without spread measures are difficult to interpret; confidence intervals are especially informative when subgroup n is low
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
A 338-gene Boolean model separates M1 (48 genes) from M2 (290 genes) macrophage polarization states and forms the basis for the exercise polarization signature.microarray human macrophage 2024×1papers★ This paper is the founder (earliest)
-
Exercise-induced macrophage polarization patterns are reproducible across human and mouse, blood and muscle biopsy, and resistance and endurance exercise modalities.RNA-seq human-mouse muscle none 2024×1papers★ This paper is the founder (earliest)
-
M1 macrophage polarization is acutely elevated in muscle and blood immediately following exercise across multiple datasets.RNA-seq human muscle up 2024×1papers★ This paper is the founder (earliest)
-
M2 macrophage polarization is sustainedly elevated in muscle following long-term exercise training.RNA-seq human muscle up 2024×1papers★ This paper is the founder (earliest)
-
Gender, age, and exercise intensity modulate the magnitude of macrophage polarization shifts without altering the direction of the exercise-induced pattern.RNA-seq human muscle mixed 2024×1papers★ This paper is the founder (earliest)
-
A 25-gene muscle macrophage signature (3 M1, 22 M2 genes) accurately predicts pre- vs post-exercise and trained vs untrained state in muscle tissue.RNA-seq human muscle 2024×1papers★ This paper is the founder (earliest)
-
M2 macrophage polarization is increased in muscle during immobilization/inactivity, opposite to the acute exercise response.RNA-seq mouse muscle up 2024×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-39476967
Paper: Voskoboynik Y, McCulloch AD, Sahoo D. Macrophages on the run: Exercise
balances macrophage polarization for improved health. Mol Metab. 2024;90:102058.
Repo: https://github.com/YoyoVosko/MacrophagesOntheRun @ 4cf2f16 (authors' own code)
Method (what the paper computes)
A Boolean-network (BoNE / StepMiner / Hegemon) M1↔M2 macrophage-polarization
gene signature (gene clusters 13, 14, 3 with weights −1/+1/+2) is used to score
each sample of ~75 public exercise transcriptomics datasets. Per dataset a
composite score separates two sample groups (pre vs post exercise = "Immediate",
or sedentary vs trained = "Training"); separation is quantified by ROC-AUC and
a Welch t-test p-value. These per-dataset AUC/Pval are tabulated in the repo's
SupplementaryTable1_…ROCAUC_Pvalues.csv and drive Figure 4 / Figure 5.
In scope (pipeline-derived → attempted)
- SupplementaryTable1: per-dataset ROC-AUC + Pval of the M1/M2 BoNE score
across all exercise datasets. Pipeline:
bone.getRanks2/mergeRanks(StepMiner rank-composite score) →getMetrics(sklearnroc_curve/auc) →getPval(scipy.stats.ttest_ind, Welch). This is the central reproducible computational result and the one we reproduced. Figures 4/5 are deterministic dot-plots of this table (geometry not separately reproduced — the underlying numbers are).
Out of scope (not pipeline / not attempted, per 80/20)
- Signature construction itself (Figs 1–3): building clusters 13/14/3 from the
M1/M2 reference compendium (
getSigMacGenes,mut.MacAnalysis,import Datasetsmodule not shipped). We instead consume the shipped signature gene lists — i.e. we reproduce the application of the signature, not its derivation. - Wet-lab / manual interpretation, pathway narratives, SupplementaryTable2 (gender/age/intensity metadata join) — non-pipeline.
- The hard ~20%: datasets whose gene-symbol→probe mapping fails on the public
server (RNA-seq platforms with differing annotation), or whose Hegemon id
returns HTTP 500 / empty — recorded as
error, not chased.
Reproduction approach (1:1, authors' own code)
The authors' scoring code is reused unchanged. The only substitution: the data
backend is moved from the lab's local Hegemon files (/booleanfs2/…, not
public) to the lab's public Hegemon web service (hegemon.ucsd.edu/.../explore.php)
— the exact endpoint the repo's own Download_Data.ipynb uses, serving the same
processed expression + StepMiner thresholds + clinical annotation the paper scored.
All compute ran in a «our HPC» SLURM job (2175577); data stayed on the server/«infra».
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a strong 1:1 reproduction: the authors' own scoring code and dataset definitions were reused unchanged, with only the data backend swapped from the lab's non-public local Hegemon files to its public web service. Of 68 gradeable per-dataset ROC-AUCs, 42 are exact and the median |dAUC| is 0.000, with p-values matching to 2-3 sig figs — the SupplementaryTable1 values (behind Figs 4 & 5) are clearly derivable from shared code+data. Deviations are on our/data-path side, not the authors': 9 errors are public-server coverage gaps and the 3 mismatches are tiny-sample (n=3/3) sensitivity, not substantive disagreements. The central claim (M1/M2 polarization score discriminates exercise states) is fully confirmed.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.