Exosomal microRNA miR-92a concentration in serum reflects human brown fat activity
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH AND REPRODUCES 1:1. Two independent pipeline-derived tracks both reproduced. RA-1 (Fig 2 / Suppl Table 1): from the deposited GEO series GSE79440 (DataAssist global-normalized signal) we recomputed group-mean fold-changes (CL/cold vs wt serum exosomes, cAMP vs control BA exosomes) with the paper's >2-fold criterion on «our HPC» (SLURM «job», «infra») - all 7 candidate miRNA fold-changes (incl. star-strand miR-34c*/miR-93* under '#' assays) match Supplementary Table 1 to 3 decimals, and the Fig 2a Venn (41 altered in both in-vivo groups, 12 up) reproduces exactly (downregulated 24 vs 29 within-tol due to undetected-assay flooring). RA-2 (Fig 3 human correlations): recomputed the headline Pearson correlations with scipy from the authors' own published per-subject data (Suppl Tables 3/4/5) - 5 of 6 Fig-3 panels reproduce exactly/within-tol (Fig 3d r=-0.528=-0.528; Fig 3e r=-0.506 vs -0.505; Fig 3f R2=0.286 vs 0.29; Fig 3h cohort-2 r=-0.631 vs -0.630; Fig 3c within-tol). The cohort-1 N=22 set was reconstructed (12 thermoneutral + 10 cold-acclimation baseline) and cross-validated against independent cells of Suppl Table 6. NOT ATTEMPTED (honestly out of scope): Fig 3b/3g group-wise t-tests (undisclosed BAT-detection threshold; only 1/22 subjects has SUVmean=0 vs reported n=6 no-BAT), cohort-1 miR-133a per-subject values (not shipped), and all Fig 1 wet-lab exosome assays (EM/WB/NTA, non-pipeline). NO fabrication indicators: 12 of 14 pinned numeric claims match to 3 decimals, the remainder within-tol with documented mechanistic reasons. Note: paper co-authored by C. Schlein (UKE).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 97assessed: 2026-06-18 ⛓ 8cdc74bfd34a
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether brown/beige adipocytes release exosomes containing miRNAs that change with thermogenic activation, and whether one such exosomal miRNA (miR-92a) in serum can serve as a non-invasive biomarker reflecting human BAT activity, as an alternative to radiation-based 18F-FDG PET/CT imaging.
- ★ Brown and beige adipocytes release exosomes, and thermogenic activation increases exosome release both in vitro and in vivo. finding
- ★ cAMP signalling selectively increases exosome release from brown and beige adipocytes but not from white adipocytes (3T3-L1), muscle cells, or hepatocytes. finding
- ★ Cold exposure increases exosome release from BAT and inguinal WAT in vivo, with BAT accounting for ~87% of total exosome output, while gonadal WAT, muscle, liver and brain show no change. finding
- ★ miRNA profiling identifies miR-92a, miR-133a and miR-34c* as showing coherent expression changes across in vitro (cAMP) and in vivo (cold, CL-316,243) BAT activation models; miR-92a was prioritized because miR-34c* was undetectable in human serum. finding
- ★ Exosomal miR-92a is downregulated upon BAT/beige activation (cAMP treatment, cold-exposure) and upregulated upon BAT 'whitening' induced by high-fat diet. finding
- ★ Serum exosomal miR-92a levels inversely correlate with human BAT activity (SUV and glucose uptake rate measured by 18F-FDG PET/CT) in two independent cohorts. finding
- ★ Exosomal miR-92a represents a potential serum biomarker for BAT activity in mice and humans. resource
- Comparative miRNA profiling of in vitro brown-adipocyte exosomes and in vivo mouse serum exosomes was used to nominate BAT-activation-associated miRNA candidates. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Electron microscopy and Western blotting (CD63-GFP) | murine brown adipocytes / BAT | none | exosome detection/release | — |
| Exosome release quantification | murine brown adipocytes, beige adipocytes, 3T3-L1 white adipocytes, C2C12 muscle cells, HepG2 hepatocytes | cAMP treatment | fold-change in exosome release | — |
| Exosome release quantification (CD63/Hsp70 markers) | mouse tissues: BAT, WATi, WATg, muscle, liver, brain | cold-exposure | tissue exosome release | — |
| miRNA profiling (array covering 757 murine miRNAs) | mouse serum exosomes | cold-exposure or CL-316,243 (β3-adrenoreceptor agonist) | miRNA expression changes vs control | — |
| miRNA profiling | exosomes from cultured brown adipocytes | cAMP treatment | miRNA expression changes | — |
| qPCR | exosomes from mouse BAT, WATi, WATg, muscle, liver, brain | cold-exposure | miR-92a abundance per tissue | — |
| qPCR + histology (lipid droplet size, triglyceride content) | mouse BAT | high-fat diet (16 weeks) | BAT whitening and miR-92a levels in BAT/serum exosomes | — |
| qPCR of serum exosomal miRNA + 18F-FDG PET/CT imaging | human subjects, cohort 1 (n=22) and cohort 2 (n=19); subset with cold acclimation or acute cold exposure | cold acclimation (10-day) / acute cold exposure (1-1.5h) / none | miR-92a and miR-133a levels vs BAT SUV and glucose uptake rate | PET/CT |
- ▲ cAMP treatment increased exosome release from brown adipocytes 4.7-fold
- ▲ cAMP treatment increased exosome release from beige adipocytes but not 3T3-L1, C2C12 or HepG2 cells 10.7-fold
- ▲ Cold-exposure increased exosome release in vivo from BAT and WATi but not WATg, muscle, liver or brain 9.05-fold (BAT), 7.62-fold (WATi)
- – miR-92a and miR-34c* significantly changed in exosomes from cAMP-treated brown adipocytes and serum of BAT-activated mice
- ▼ Cold-exposure reduced exosomal miR-92a in BAT and WATi 6% of control (BAT), 18% of control (WATi)
- ▲ HFD-induced BAT whitening increased miR-92a release from BAT and in serum exosomes 2.75±0.65-fold (BAT), 2.66±0.19-fold (serum)
- ▼ Serum exosomal miR-92a was significantly lower in high-BAT than low-BAT subjects and negatively correlated with BAT SUVmean/SUVmax/glucose uptake in cohort 1
- ▼ In cohort 2, exosomal miR-92a was significantly higher in BAT-negative than BAT-positive individuals and correlated with glucose uptake rate R2=0.40
- fold_change 4.7-fold (cAMP-induced exosome release increase in brown adipocytes)
- fold_change 10.7-fold (cAMP-induced exosome release increase in beige adipocytes)
- fold_change 9.05-fold (BAT), 7.62-fold (WATi) (cold-exposure induced exosome release increase in vivo)
- pvalue P=0.017 (serum exosomal miR-92a lower in high-BAT vs low-BAT group, cohort 1)
- correlation R2=0.28, P=0.011 (Log10 miR-92a vs BAT SUVmax, cohort 1)
- correlation R2=0.26, P=0.016 (Log10 miR-92a vs glucose uptake rate, cohort 1)
- correlation R2=0.40, P=0.004 (Log10 miR-92a vs glucose uptake rate, cohort 2)
- count 8 out of 10 correct predictions (prediction of BAT activity change on cold acclimation from miR-92a change)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper used a multi-stage design combining in vitro murine adipocyte experiments, in vivo mouse cold-exposure and β3-agonist models, and two independent cross-sectional human cohorts (n=22 and n=19) assessed by 18F-FDG PET/CT. Group differences in exosome release and miRNA abundance were tested with unpaired two-tailed Student's t-tests throughout; continuous associations between serum exosomal miR-92a and BAT activity metrics were assessed by Pearson's correlation on log10-transformed miR-92a values; and stepwise multivariable linear regression identified independent predictors of log10 miR-92a in both cohorts. Results were reported as means ± SEM or SD with exact or threshold P values and R² statistics for correlations.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Unpaired two-tailed Student's t-test | Exosome particle counts: cAMP-treated vs untreated brown, beige, white (3T3-L1) adipocytes, muscle (C2C12), and hepatocytes (HepG2) in vitro (Fig. 1c); cold-exposed vs control BAT, WATi, WATg, muscle, liver, and brain in vivo (Fig. 1d) | not stated for individual groups | not stated |
| Unpaired two-tailed Student's t-test | miR-92a, miR-34c*, and miR-133a levels in exosomes from cAMP-treated vs untreated brown adipocytes and serum of CL-treated or cold-exposed vs control mice (Fig. 2b, 2c); exosomal miR-92a across six tissue types after cold exposure (Fig. 2d) | not stated | not stated |
| Unpaired two-tailed Student's t-test | Serum exosomal miR-92a in high BAT vs low BAT group (median SUVmean split, n=11 per group), human cohort 1 (Fig. 3b); P=0.017 | n=11 per group | not stated |
| Unpaired two-tailed Student's t-test | Serum exosomal miR-133a in high BAT vs low BAT group, cohort 1 (Supplementary Fig. 3a); P=0.757 | n=11 per group | not stated |
| Unpaired two-tailed Student's t-test | Exosomal miR-92a in BAT-negative vs BAT-positive individuals, cohort 2 (Fig. 3g) | n=19 total; subgroup sizes not stated | not stated |
| Pearson's correlation | Log10 miR-92a vs BAT SUVmean (R²=0.18, P<0.05, n=22), vs BAT SUVmax (R²=0.28, P=0.011, n=22; log10–log10 form: R²=0.25, P<0.05), vs glucose uptake rate (R²=0.26, P=0.016, n=22), cohort 1 (Figs 3c–e, Supplementary Figs 3c–d) | n=22 | not stated |
| Pearson's correlation | Change in miR-92a vs change in BAT SUVmean after 10-day cold acclimation, cohort 1 subset (Fig. 3f); R²=0.29, P=0.11 | n=10 | not stated |
| Pearson's correlation | Log10 miR-92a vs BAT SUVmean after acute cold exposure, cohort 1 subset (Supplementary Fig. 3e); R²=0.21, P=0.08 | n=15 | not stated |
| Pearson's correlation | Log10 miR-92a vs BAT glucose uptake rate, cohort 2 (Fig. 3h); R²=0.40, P=0.004 | n=19 | not stated |
| Stepwise multivariable linear regression | Log10 miR-92a as dependent variable; age, sex, BMI, fat mass, BAT SUVmean, SUVmax, and glucose uptake rate as candidates, cohort 1 (Supplementary Table 6) | n=22 | not stated |
| Stepwise multivariable linear regression | Log10 miR-92a as dependent variable; age, sex, BMI, fat mass, and glucose uptake rate as candidates, cohort 2 (Supplementary Table 7) | n=19 | not stated |
| Average-linkage hierarchical clustering | 188 detected miRNA expression profiles from in vitro (cAMP-treated brown adipocytes) and in vivo (serum of cold or CL-treated mice) to assess similarity structure (Supplementary Fig. 1e) | 188 miRNAs across experimental groups | na |
| Fold-change threshold filter (>2-fold) with Venn diagram overlap | Screening of 757 profiled murine serum exosomal miRNAs to identify candidates altered by CL treatment or cold exposure relative to controls; cross-referenced with in vitro brown adipocyte profile (Fig. 2a, Supplementary Table 1) | 757 miRNAs profiled | na |
-
miRNA candidates were selected from 757 profiled miRNAs using a >2-fold change threshold without a formal multiple-testing correction↳ Could also: A false discovery rate (FDR) procedure such as Benjamini-Hochberg applied across all profiled miRNAs could also have been used to rank and filter candidates — FDR control provides a probabilistic bound on the expected proportion of false positives among selected candidates, which is particularly informative when simultaneously screening hundreds of features; it would complement fold-change filtering by separating magnitude from statistical reliability
-
Multiple independent t-tests were used to compare exosome release and miR-92a levels across six tissue types between two conditions (cold-exposed vs control)↳ Could also: A two-way ANOVA (tissue × condition) followed by a post-hoc test with multiplicity correction (e.g., Tukey HSD or Bonferroni) could also have been applied — A single ANOVA model accounts for the full family of comparisons across tissue types in one framework and a post-hoc correction controls the experiment-wise error rate; it would also allow formal testing of a tissue-by-condition interaction, which is relevant to the paper's central claim about tissue specificity
-
Cohort 1 subjects were dichotomized into 'high BAT' and 'low BAT' groups at the median SUVmean for a t-test comparison of miR-92a↳ Could also: Treating BAT activity as a continuous predictor in simple linear regression or Pearson's correlation (which the authors also report) could serve as the primary analysis — Median-split dichotomization discards information about the magnitude of BAT activity and reduces statistical power; continuous analyses—which the authors also performed—more fully characterize the dose-response relationship between miR-92a and BAT activity
-
Pearson's correlation was used to associate log10-transformed miR-92a with BAT activity metrics in cohorts of n=22 and n=19↳ Could also: Spearman's rank correlation could also have been used as a distribution-free alternative — Spearman's correlation does not assume bivariate normality and is less sensitive to influential outliers, which can have a large effect in small samples; reporting both coefficients would indicate whether any association is driven by a few extreme observations
-
Stepwise multivariable linear regression was used to select predictors of log10 miR-92a, with up to seven candidate predictors relative to n=19–22 subjects↳ Could also: Regularized regression (e.g., LASSO or ridge regression) could also have been used for simultaneous variable selection and coefficient shrinkage in these small-n, multi-predictor settings — Stepwise selection can be unstable with small samples and correlated predictors, and tends to produce overfitted models; regularized approaches penalize model complexity systematically and typically yield more stable, generalizable estimates
-
Most experimental group results are reported as mean ± SEM↳ Could also: Mean ± SD or 95% confidence intervals could also be used to express the spread of observations — SEM shrinks with increasing sample size and quantifies uncertainty in the mean estimate rather than biological variability across observations; SD and CIs directly describe the distribution of individual measurements and are often preferred for conveying biological spread, especially in small-n experiments
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-27117818
Paper: Chen Y, et al. "Exosomal microRNA miR-92a concentration in serum reflects human brown fat activity." Nat Commun 7:11420 (2016). doi:10.1038/ncomms11420 · PMC4853423.
Note: Co-authored by C. Schlein (UKE Hamburg) — the operator's own group.
Datasets the paper relies on
| ref | what | where | access |
|---|---|---|---|
| GSE79440 | mouse/BA exosomal miRNA TaqMan Low-Density-Array (RT-qPCR panel) Ct → DataAssist v3.01 global-normalized signal, 8 samples | GEO | open |
| Supplementary Tables 3/4/5 (ncomms11420-s1.pdf) | per-subject human data (PET BAT activity + serum exosomal Log10 miR-92a), cohort 1 (n=22) + cohort 2 (n=19) | Nature SI | open |
| Supplementary Tables 6/7 | reported Pearson correlation matrices (cohort 1 / cohort 2) | Nature SI | open |
| Supplementary Table 1/2 | the 7 deregulated miRNA candidates + FC values (Fig 2) | Nature SI | open |
In scope (pipeline-/statistics-derived, reproducible)
| id | result | pipeline | data |
|---|---|---|---|
| RA-1 | Fig 2 / Supp Table 1: miRNAs deregulated >2-fold in CL-316,243 / cold / cAMP vs control; the 7 candidate FCs; Venn overlap | DataAssist global-norm → fold-change (ratio of normalized signal) → 2-fold threshold | GSE79440 |
| RA-2 | Fig 3 c/d/e/f/h: Pearson correlations of serum exosomal Log10 miR-92a vs BAT SUVmean / SUVmax / glucose-uptake (cohort 1 n=22), the cold-acclimation Δ correlation (n=10), and cohort 2 glucose-uptake (n=19) | Pearson correlation (scipy) on published per-subject values | Supp Tables 3/4/5 |
RA-2 is a third-party-tool-on-the-paper's-data reproduction (P16): recomputing the authors' statistics with scipy from their published per-subject tables. Equally valid.
Out of scope (wet-lab / instrument / not pipeline-derived)
- Exosome isolation, electron microscopy, western blots, NTA exosome counts (Fig 1).
- qPCR wet-lab measurement of miR-92a itself (we reuse the reported per-subject values).
- Fig 3a PET/CT image; Fig 3b high/low-BAT group t-test and Fig 3g no-BAT/BAT t-test — depend on an undisclosed BAT-detection threshold / group assignment not derivable from the shipped tables (only 1/22 subjects has SUVmean=0, paper's Fig 3g uses n=6 "no detectable BAT") → not cleanly reproducible; noted, not forced.
- Mouse in-vivo / in-vitro wet-lab assays underlying the miRNA panel (the panel output Ct/normalized signal IS in GSE79440 and is in scope; the bench work is not).
Heavy compute
RA-1 download + analysis runs in a «our HPC» SLURM job with cwd on «infra»
(reproductions/pmid-27117818/). RA-2 is light statistics on small published tables.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
PMID 27117818 (Schlein/Heeren et al.) reproduces well from the authors' own deposited data (GEO GSE79440 + per-subject SI tables): 12 of 14 pinned claims match to 3 decimals and the central claim — serum exosomal miR-92a inversely correlates with BAT activity — is fully confirmed with matching r and P-values (Fig 3d/e/h). The only deviations are two within-tol items with documented causes on our preprocessing side: the Fig 2a Venn down-count (24 vs 29, from undetected-assay flooring) and the Fig 3c regression slope (~7%, from SUVmean ambiguity in the reconstructed n=22 cohort). No fabrication signal; overall a solid reproduction with minor explainable deviations.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.