Comparative profiling of skeletal muscle models reveals heterogeneity of transcriptome and metabolism.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the PIPELINE, and the paper's biology reproduces 1:1 on the deposited data: re-running the deposited R pipeline (limma + per-gene 2-way ANOVA + clusterProfiler GO) on the repo's deposited processed matrices recovers the paper's named marker genes EXACTLY (C2C12 actin/myosin; L6 SLC2A1 + all five ETC complexes), 85-92% of the deposited model-specific gene lists, the Fig 1B PCA, the Fig 1C correlation/heterogeneity structure, and the GO themes (L6=proliferation, C2C12=muscle, HSMC=development). HOWEVER the deposited DATA does not exactly regenerate the deposited Stats/ OUTPUTS: gene-list counts run ~10-25% higher (deposited matrix is larger than the version behind the committed Stats), and Step03's 2-way ANOVA is NOT regenerable from the deposited GENENAME_batch.Rds at all -- the deposited column names are platform-prefixed (GPL17586_HumanCell_GSM..._UC1.cel.gz) and break the deposited gsub Species/Sample parser (-> degenerate design -> all-NA), and even after correcting the parser the exact p-values differ from shipped (deposited batch values differ from those behind the committed stats), though marker-gene significance is concordant. Both data and Stats were committed in the same commit (1bdd58f) yet are mutually inconsistent. Verdict PARTIAL: pipeline + biological conclusions reproduce; exact deposited numeric outputs do not, due to a deposit-hygiene/code-naming problem -- NOT evidence of fabrication (the reported biology is supported by the re-run). NOT attempted: all wet-lab metabolic assays (out of scope, bench experiments), the upstream raw-CEL->normalized RMA step (not shipped as runnable with pinned inputs), Step10 proteomics integration.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 64assessed: 2026-06-16 ⛓ be4cafbb4ca5
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe authors hypothesized that metabolic differences between rat L6, mouse C2C12, and human primary skeletal muscle cell (HSMC) models would be reflected at the mRNA level, and that transcriptomic differences could account for the diverse metabolic and contractile phenotypes observed across these models.
- ★ C2C12 myotubes show enriched mRNA expression of genes coding for actin and myosin compared with L6 and HSMC. finding
- ★ L6 myotubes have the highest mRNA levels of glucose transporter genes and genes encoding the five complexes of the mitochondrial electron transport chain. finding
- ★ Insulin-stimulated glucose uptake and oxidative capacity are greatest in L6 myotubes. finding
- ★ Insulin-induced glycogen synthesis is highest in HSMCs. finding
- ★ C2C12 myotubes have higher baseline glucose oxidation than L6 and HSMC. finding
- ★ All three models respond to electrical pulse stimulation with increased glucose uptake and altered gene expression, but in a slightly different manner. finding
- ★ There is substantial heterogeneity in transcriptomic and metabolic profiles among L6, C2C12, and primary human myotubes, warranting model-specific recommendations for experimental use. finding
- The best correlations of transcript expression occur between cells and tissues from the same species, indicating strong species-specific effects on skeletal muscle mRNA composition. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Transcriptomic profiling (microarray, public GEO datasets) | Mouse C2C12, rat L6, human primary myotubes, and skeletal muscle tissue (mouse/rat/human) | none (untreated/control samples only) | mRNA expression levels, PCA clustering, correlation, top-100 gene overlap | Affymetrix arrays (human GPL570/GPL6244/GPL17586; rat GPL1355/GPL22740/GPL14746; mouse GPL81/GPL1261/GPL17400) |
| Glucose uptake (2-deoxyglucose) | L6, C2C12, HSMC myotubes | insulin (100 nmol/L) | 2-[3H]deoxyglucose uptake normalized to protein | liquid scintillation counter (WinSpectral 1414) |
| Glycogen synthesis | L6, C2C12, HSMC myotubes | insulin (100 nM) | [14C]glucose incorporation into glycogen normalized to protein | liquid scintillation counter |
| Glucose oxidation | L6, C2C12, HSMC myotubes | FCCP (0.5 µM) present/absent | [14C]-CO2 release normalized to protein | liquid scintillation counter |
| Fatty acid oxidation | L6, C2C12, HSMC myotubes | palmitate (25 µM) with [3H]palmitate tracer | released [3H] normalized to protein | liquid scintillation counter |
| Oxygen consumption / extracellular acidification (Seahorse Mito Stress Test) | L6, C2C12, HSMC myotubes | oligomycin, FCCP, rotenone/antimycin A | OCR and ECAR | Seahorse XF (Agilent) |
| Electrical pulse stimulation (EPS) | L6, C2C12, HSMC myotubes | EPS (40 V, 2 ms, 1 Hz, 3 h) | glucose uptake and gene expression (qPCR) | C-Pace EP (Ionoptix) |
| Proliferation assay (imaging surface area + BrdU incorporation) | L6, C2C12, HSMC myoblasts | none | cell surface area over 72 h; BrdU incorporation over 12-24 h | inverted light microscope/ImageJ; BrdU Cell Proliferation Assay Kit |
- ▲ C2C12 myotubes show enriched actin/myosin gene expression relative to L6 and HSMC.
- ▲ L6 myotubes show highest levels of glucose transporter genes and all five mitochondrial ETC complex genes.
- ▲ Insulin-stimulated glucose uptake and oxidative capacity greatest in L6 myotubes.
- ▲ Insulin-induced glycogen synthesis highest in HSMCs.
- ▲ C2C12 myotubes exhibit higher baseline glucose oxidation than L6 and HSMC.
- – All models respond to EPS with increased glucose uptake and altered gene expression, but responses differ slightly between models.
- – Cell and tissue models cluster separately by PCA, and transcript correlations are strongest between cells and tissues of the same species.
- – Significant overlap observed among the top 100 expressed genes when comparing models.
- count n=11 (C2C12), 13 (L6), 11 donors (HSMC) (glucose uptake experiment replicates)
- count n=9 (C2C12), 8 (L6), 12 donors (HSMC) (glycogen synthesis experiment replicates)
- count n=6 (C2C12), 8 (L6), 19 replicates/12 donors (HSMC) (glucose oxidation experiment replicates)
- count n=11 (C2C12), 17 (L6), 37 replicates/12 donors (HSMC) (fatty acid oxidation experiment replicates)
- count n=7 (C2C12), 8 (L6), 13 replicates/8 donors (HSMC) (Seahorse OCR/ECAR experiment replicates)
- pvalue false discovery rate < 0.05 (threshold for gene ontology enrichment analysis)
- mean age 50.4 ± 10 yr; BMI 22.9 ± 2.4 kg/m2 (characteristics of HSMC donors (healthy men))
- other passage 6-9 (passage range of HSMC used for experiments)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
Transcriptomic comparisons used publicly available microarray data processed with robust multi-array and quantile normalization, with differential expression and ranked statistics computed in limma and gene ontology enrichment thresholded at a false discovery rate < 0.05. For the wet-lab metabolic and gene-expression assays, group means across three cell models were compared after testing normality with the Shapiro-Wilk test: normally distributed data were analyzed by ANOVA with Tukey's multiple comparison, and non-normal data by Kruskal-Wallis with Dunn's multiple comparison. Results were summarized as group averages of independent experiments/donors, with sample sizes and tests stated in figure captions.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| One-way ANOVA with Tukey's multiple comparison | normally distributed metabolic/functional assay comparisons across L6, C2C12, and HSMC (per figure captions) | independent experiments per cell model and donor-derived replicates for HSMC (varies by assay, e.g. glucose uptake 11 C2C12/13 L6/11 HSMC) | stated |
| Kruskal-Wallis test with Dunn's multiple comparison | non-normally distributed metabolic/functional assay comparisons across the three models | independent experiments and donor replicates as stated per assay | stated |
| Shapiro-Wilk normality test | pre-test assessment of distribution for assay data to choose parametric vs nonparametric path | — | na |
| limma differential expression with ranked statistics | transcriptomic comparison of cells/tissues across human, rat, mouse platforms | publicly available GEO samples per platform (Supplemental Table S1) | not stated |
| clusterProfiler gene ontology enrichment | genes differentially expressed in one model vs the other two (Fig. 2D) | genes passing FDR < 0.05 | not stated |
| Correlation analysis / principal component analysis | clustering and correlation of transcript profiles across models and tissues (Fig. 1B, 1C) | — | not stated |
-
Normality was assessed with the Shapiro-Wilk test, and the analysis branched to ANOVA/Tukey or Kruskal-Wallis/Dunn accordingly.↳ Could also: A consistently applied nonparametric framework, or a single model with variance-stabilizing transformation, could also be used regardless of the normality test outcome. — With the modest per-group sample sizes here, normality tests have limited power; pre-specifying one approach or transforming the data can simplify interpretation and reduce dependence on a screening test.
-
The three cell models were compared with one-way ANOVA (or Kruskal-Wallis) treating each as an independent group.↳ Could also: A mixed-effects model treating donor as a random effect (especially for HSMC, where some replicates come from the same individuals) could also be applied. — A mixed model can account for repeated sampling from the same donors and unequal replicate structure across models, which would more explicitly represent biological clustering.
-
Group results were summarized as averages of independent experiments and donor replicates.↳ Could also: Reporting a 95% confidence interval or showing individual data points alongside the mean could also be presented. — Confidence intervals and dot plots convey both spread and uncertainty directly, which is often informative for small sample sizes.
-
Transcriptomic differential expression used limma with ranked statistics on microarray data merged across species-specific platforms.↳ Could also: Surrogate variable analysis or explicit batch/platform covariates within limma could also be incorporated. — Modeling platform/species as covariates can help separate technical interarray variability from biological signal, a point the authors themselves note as a consideration.
-
GO enrichment thresholded genes at FDR < 0.05 and used clusterProfiler over-representation analysis.↳ Could also: A threshold-free approach such as Gene Set Enrichment Analysis (GSEA) on the full ranked gene list could also be used. — GSEA uses all genes without a hard significance cutoff, which can detect coordinated shifts in pathways where individual genes do not pass the FDR threshold.
-
Pairwise model differences were controlled with Tukey's or Dunn's post hoc correction within each assay.↳ Could also: Reporting effect sizes (e.g., standardized mean differences) alongside the corrected p-values could also be included. — Effect-size estimates communicate the magnitude of differences between models, complementing the significance testing for readers choosing a model.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
PCA segregates in vitro muscle cell models from skeletal muscle tissue, with best transcript correlations within the same species, revealing transcriptome heterogeneity across models.microarray human rat mouse skeletal muscle 2019×1papers★ This paper is the founder (earliest)
-
Gene ontology enrichment shows rat L6 cells enriched for metabolism- and proliferation-related pathways.microarray rat l6 myotube 2019×1papers★ This paper is the founder (earliest)
-
L6 cells show highest expression of glucose transporters and the five ETC complexes, while C2C12 are enriched for actin/myosin genes.microarray rat l6 myotube up 2019×1papers★ This paper is the founder (earliest)
-
Insulin-induced glycogen synthesis is highest in human (HSMC) myotubes.other human hsmc myotube up 2019×1papers★ This paper is the founder (earliest)
-
All muscle models respond to electrical pulse stimulation with increased glucose uptake and gene expression, with model-specific differences.other human rat mouse skeletal muscle myotube up 2019×1papers★ This paper is the founder (earliest)
-
C2C12 myotubes show higher baseline glucose oxidation.other mouse c2c12 myotube up 2019×1papers★ This paper is the founder (earliest)
-
Insulin-stimulated glucose uptake and oxidative capacity are greatest in L6 myotubes.other rat l6 myotube up 2019×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
- CoINcIDE: A framework for discovery of patient... L1 87/100
- A curated collection of transcriptome datasets... L1 62/100
- An NMF-Based Methodology for Selecting Biomark... L1 84/100
- Unveiling prognostics biomarkers of tyrosine m...⚑ L1 51/100 ⚑
- Meta-analysis of gene expression profiles of l... L1 78/100
- Colorectal Cancer Prediction Based on Weighted...⚑ L1 80/100 ⚑
- Curation of over 10 000 transcriptomic studies... L1 80/100
- Construction and Validation of an Immune Infil...⚑ L1 51/100 ⚑
- Identification of a novel 10 immune-related ge...
- Exploration of the shared diagnostic genes and... L1 76/100
- IRSN-23 gene diagnosis enhances breast cancer... L1 71/100
- Molecular Classification Models for Triple Neg... L1 86/100
- Predicting Bone Metastasis Using Gene Expressi... L1 62/100
- Autoencoder Networks Decipher the Association... L1 74/100
- Comprehensive analysis of a novel RNA modifica... L1 71/100
- Discovery and validation of molecular patterns... L1 83/100
- A curated collection of transcriptome datasets... L1 62/100
- PulmonDB: a curated lung disease gene expressi...⚑ L1 53/100 ⚑
- An NMF-Based Methodology for Selecting Biomark... L1 84/100
- Meta-analysis of gene expression profiles of l... L1 78/100
- Curation of over 10 000 transcriptomic studies... L1 80/100
- Predicting Bone Metastasis Using Gene Expressi... L1 62/100
- Comprehensive analysis of a novel RNA modifica... L1 71/100
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Re-running the deposited limma+ANOVA+clusterProfiler pipeline on the deposited matrices reproduces the paper's biology 1:1 qualitatively — named marker genes are exact and GO/PCA/correlation structure holds — so the central conclusion is confirmed and there is no fabrication signal. The deviations are entirely at the deposit/input level: the deposited Data_Processed is an expanded version that no longer matches the data behind the committed Stats (gene counts ~10-25% high), and the deposited 2-way_ANOVA.txt is not regenerable from the shared data because platform-prefixed column names break the deposited parser. This is an authors'-side deposit-hygiene/code-data inconsistency (q4 red), moderate in magnitude with direction preserved (q6 yellow), leaving an overall solid-but-deviating reproduction (q8 yellow).
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.