Comparative profiling of skeletal muscle models reveals heterogeneity of transcriptome and metabolism.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the PIPELINE, and the paper's biology reproduces 1:1 on the deposited data: re-running the deposited R pipeline (limma + per-gene 2-way ANOVA + clusterProfiler GO) on the repo's deposited processed matrices recovers the paper's named marker genes EXACTLY (C2C12 actin/myosin; L6 SLC2A1 + all five ETC complexes), 85-92% of the deposited model-specific gene lists, the Fig 1B PCA, the Fig 1C correlation/heterogeneity structure, and the GO themes (L6=proliferation, C2C12=muscle, HSMC=development). HOWEVER the deposited DATA does not exactly regenerate the deposited Stats/ OUTPUTS: gene-list counts run ~10-25% higher (deposited matrix is larger than the version behind the committed Stats), and Step03's 2-way ANOVA is NOT regenerable from the deposited GENENAME_batch.Rds at all -- the deposited column names are platform-prefixed (GPL17586_HumanCell_GSM..._UC1.cel.gz) and break the deposited gsub Species/Sample parser (-> degenerate design -> all-NA), and even after correcting the parser the exact p-values differ from shipped (deposited batch values differ from those behind the committed stats), though marker-gene significance is concordant. Both data and Stats were committed in the same commit (1bdd58f) yet are mutually inconsistent. Verdict PARTIAL: pipeline + biological conclusions reproduce; exact deposited numeric outputs do not, due to a deposit-hygiene/code-naming problem -- NOT evidence of fabrication (the reported biology is supported by the re-run). NOT attempted: all wet-lab metabolic assays (out of scope, bench experiments), the upstream raw-CEL->normalized RMA step (not shipped as runnable with pinned inputs), Step10 proteomics integration.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 64assessed: 2026-06-16 ⛓ be4cafbb4ca5
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusMetabolic and contractile differences between commonly used skeletal muscle cell models (rat L6, mouse C2C12, primary human myotubes) may be reflected in and accounted for by differences in their transcriptomic landscapes.
- ★ L6, C2C12, and primary human myotubes show substantial heterogeneity in transcriptomic and metabolic profiles finding
- ★ Actin- and myosin-coding mRNAs are enriched in C2C12, whereas L6 has highest expression of glucose transporters and the five mitochondrial electron transport chain complexes finding
- ★ Insulin-stimulated glucose uptake and oxidative capacity are greatest in L6 myotubes finding
- ★ Insulin-induced glycogen synthesis is highest in human myotubes (HSMCs), while C2C12 has higher baseline glucose oxidation finding
- ★ All three models respond to electrical pulse stimulation with increased glucose uptake and gene expression, but in slightly different manners finding
- Cells and tissues correlate best within the same species, indicating strong species-specific differences in muscle mRNA composition finding
- A merged cross-species transcriptomic profiling pipeline using public GEO data, human ortholog annotation, and ranked statistics enables comparison of muscle models method
- ★ Model-specific transcriptomic/metabolic signatures support recommendations for which model to use for a given hypothesis resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Transcriptomic profiling (microarray meta-analysis) | Human, rat, mouse skeletal muscle tissue and L6/C2C12/HSMC myotubes (public GEO data) | none (untreated control samples) | mRNA expression, PCA clustering, correlation, gene ontology enrichment | Affymetrix arrays GPL570/GPL6244/GPL17586 (human), GPL1355/GPL22740/GPL14746 (rat), GPL81/GPL1261/GPL17400 (mouse) |
| Proliferation (cell surface imaging) | L6, C2C12, HSMC myoblasts | none | surface area occupied by cells over 72 h | Inverted light microscope, ImageJ |
| BrdU incorporation proliferation assay | L6, C2C12, HSMC myoblasts | none | BrdU incorporation (proliferation) | BrdU Cell Proliferation Assay Kit no. 6813, Cell Signaling Technology |
| Radiolabeled 2-deoxyglucose uptake | L6, C2C12, HSMC myotubes | insulin (100 nmol/L) | 2-[1,2-3H]deoxy-D-glucose uptake | Liquid scintillation counter (WinSpectral 1414, Wallac) |
| Glycogen synthesis ([14C]glucose incorporation) | L6, C2C12, HSMC myotubes | insulin (100 nM) | [14C]-labeled glycogen | Liquid scintillation counter; D-[U-14C]glucose (PerkinElmer) |
| Glucose oxidation ([14C]CO2 capture) | L6, C2C12, HSMC myotubes | FCCP (0.5 µM) vs none | [14C]-CO2 released | Liquid scintillation counter; D-[U-14C]glucose |
| Fatty acid oxidation ([3H]palmitate) | L6, C2C12, HSMC myotubes | none (palmitate substrate) | 3H-labeled oxidation products | Liquid scintillation counter; [9,10-3H(N)]palmitate (PerkinElmer) |
| Mitochondrial respiration (Seahorse XF Mito Stress Test) | L6, C2C12, HSMC myotubes | oligomycin (1 µM), FCCP (2 µM), rotenone/antimycin A (0.75 µM) | oxygen consumption rate (OCR) and extracellular acidification rate (ECAR) | Agilent Seahorse XF analyzer |
- – PCA showed clear segregation of cell models from skeletal muscle tissues, and best transcript correlations occurred within the same species
- – C2C12 enriched for actin/myosin genes; L6 highest for glucose transporters and the five ETC complexes
- ▲ Insulin-stimulated glucose uptake and oxidative capacity greatest in L6 myotubes
- ▲ Insulin-induced glycogen synthesis highest in human myotubes (HSMCs)
- ▲ C2C12 myotubes had higher baseline glucose oxidation
- ▲ All models responded to EPS with increased glucose uptake and gene expression, with model-specific differences
- – Significant overlap in top 100 expressed genes across models despite overall divergence top 100 genes
- – Gene ontology enrichment showed rat L6 enriched for metabolism- and proliferation-related pathways
- count 12 healthy men, age 50.4 ± 10 yr, BMI 22.9 ± 2.4 kg/m2 (Primary human muscle cell donors (vastus lateralis biopsies))
- other FDR < 0.05 (Significance threshold for differential expression / gene ontology)
- count top 100 expressed genes (Genes compared for overlap across models (chord diagram))
- count n=11 (C2C12), 13 (L6), 11 (HSMC) (Independent experiments/donors for glucose uptake assay)
- count n=9 (C2C12), 8 (L6), 12 (HSMC) (Independent experiments/donors for glycogen synthesis assay)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
Transcriptomic comparisons used publicly available microarray data processed with robust multi-array and quantile normalization, with differential expression and ranked statistics computed in limma and gene ontology enrichment thresholded at a false discovery rate < 0.05. For the wet-lab metabolic and gene-expression assays, group means across three cell models were compared after testing normality with the Shapiro-Wilk test: normally distributed data were analyzed by ANOVA with Tukey's multiple comparison, and non-normal data by Kruskal-Wallis with Dunn's multiple comparison. Results were summarized as group averages of independent experiments/donors, with sample sizes and tests stated in figure captions.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| One-way ANOVA with Tukey's multiple comparison | normally distributed metabolic/functional assay comparisons across L6, C2C12, and HSMC (per figure captions) | independent experiments per cell model and donor-derived replicates for HSMC (varies by assay, e.g. glucose uptake 11 C2C12/13 L6/11 HSMC) | stated |
| Kruskal-Wallis test with Dunn's multiple comparison | non-normally distributed metabolic/functional assay comparisons across the three models | independent experiments and donor replicates as stated per assay | stated |
| Shapiro-Wilk normality test | pre-test assessment of distribution for assay data to choose parametric vs nonparametric path | — | na |
| limma differential expression with ranked statistics | transcriptomic comparison of cells/tissues across human, rat, mouse platforms | publicly available GEO samples per platform (Supplemental Table S1) | not stated |
| clusterProfiler gene ontology enrichment | genes differentially expressed in one model vs the other two (Fig. 2D) | genes passing FDR < 0.05 | not stated |
| Correlation analysis / principal component analysis | clustering and correlation of transcript profiles across models and tissues (Fig. 1B, 1C) | — | not stated |
-
Normality was assessed with the Shapiro-Wilk test, and the analysis branched to ANOVA/Tukey or Kruskal-Wallis/Dunn accordingly.↳ Could also: A consistently applied nonparametric framework, or a single model with variance-stabilizing transformation, could also be used regardless of the normality test outcome. — With the modest per-group sample sizes here, normality tests have limited power; pre-specifying one approach or transforming the data can simplify interpretation and reduce dependence on a screening test.
-
The three cell models were compared with one-way ANOVA (or Kruskal-Wallis) treating each as an independent group.↳ Could also: A mixed-effects model treating donor as a random effect (especially for HSMC, where some replicates come from the same individuals) could also be applied. — A mixed model can account for repeated sampling from the same donors and unequal replicate structure across models, which would more explicitly represent biological clustering.
-
Group results were summarized as averages of independent experiments and donor replicates.↳ Could also: Reporting a 95% confidence interval or showing individual data points alongside the mean could also be presented. — Confidence intervals and dot plots convey both spread and uncertainty directly, which is often informative for small sample sizes.
-
Transcriptomic differential expression used limma with ranked statistics on microarray data merged across species-specific platforms.↳ Could also: Surrogate variable analysis or explicit batch/platform covariates within limma could also be incorporated. — Modeling platform/species as covariates can help separate technical interarray variability from biological signal, a point the authors themselves note as a consideration.
-
GO enrichment thresholded genes at FDR < 0.05 and used clusterProfiler over-representation analysis.↳ Could also: A threshold-free approach such as Gene Set Enrichment Analysis (GSEA) on the full ranked gene list could also be used. — GSEA uses all genes without a hard significance cutoff, which can detect coordinated shifts in pathways where individual genes do not pass the FDR threshold.
-
Pairwise model differences were controlled with Tukey's or Dunn's post hoc correction within each assay.↳ Could also: Reporting effect sizes (e.g., standardized mean differences) alongside the corrected p-values could also be included. — Effect-size estimates communicate the magnitude of differences between models, complementing the significance testing for readers choosing a model.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
PCA segregates in vitro muscle cell models from skeletal muscle tissue, with best transcript correlations within the same species, revealing transcriptome heterogeneity across models.microarray human rat mouse skeletal muscle 2019×1papers★ This paper is the founder (earliest)
-
Gene ontology enrichment shows rat L6 cells enriched for metabolism- and proliferation-related pathways.microarray rat l6 myotube 2019×1papers★ This paper is the founder (earliest)
-
L6 cells show highest expression of glucose transporters and the five ETC complexes, while C2C12 are enriched for actin/myosin genes.microarray rat l6 myotube up 2019×1papers★ This paper is the founder (earliest)
-
Insulin-induced glycogen synthesis is highest in human (HSMC) myotubes.other human hsmc myotube up 2019×1papers★ This paper is the founder (earliest)
-
All muscle models respond to electrical pulse stimulation with increased glucose uptake and gene expression, with model-specific differences.other human rat mouse skeletal muscle myotube up 2019×1papers★ This paper is the founder (earliest)
-
C2C12 myotubes show higher baseline glucose oxidation.other mouse c2c12 myotube up 2019×1papers★ This paper is the founder (earliest)
-
Insulin-stimulated glucose uptake and oxidative capacity are greatest in L6 myotubes.other rat l6 myotube up 2019×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
- CoINcIDE: A framework for discovery of patient... L1 87/100
- A curated collection of transcriptome datasets... L1 62/100
- An NMF-Based Methodology for Selecting Biomark... L1 84/100
- Unveiling prognostics biomarkers of tyrosine m...⚑ L1 51/100 ⚑
- Meta-analysis of gene expression profiles of l... L1 78/100
- Colorectal Cancer Prediction Based on Weighted...⚑ L1 80/100 ⚑
- Curation of over 10 000 transcriptomic studies... L1 80/100
- Construction and Validation of an Immune Infil...⚑ L1 51/100 ⚑
- Identification of a novel 10 immune-related ge...
- Exploration of the shared diagnostic genes and... L1 76/100
- IRSN-23 gene diagnosis enhances breast cancer... L1 71/100
- Molecular Classification Models for Triple Neg... L1 86/100
- Predicting Bone Metastasis Using Gene Expressi... L1 62/100
- Autoencoder Networks Decipher the Association... L1 74/100
- Comprehensive analysis of a novel RNA modifica... L1 71/100
- Discovery and validation of molecular patterns... L1 83/100
- A curated collection of transcriptome datasets... L1 62/100
- PulmonDB: a curated lung disease gene expressi...⚑ L1 53/100 ⚑
- An NMF-Based Methodology for Selecting Biomark... L1 84/100
- Meta-analysis of gene expression profiles of l... L1 78/100
- Curation of over 10 000 transcriptomic studies... L1 80/100
- Predicting Bone Metastasis Using Gene Expressi... L1 62/100
- Comprehensive analysis of a novel RNA modifica... L1 71/100
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Re-running the deposited limma+ANOVA+clusterProfiler pipeline on the deposited matrices reproduces the paper's biology 1:1 qualitatively — named marker genes are exact and GO/PCA/correlation structure holds — so the central conclusion is confirmed and there is no fabrication signal. The deviations are entirely at the deposit/input level: the deposited Data_Processed is an expanded version that no longer matches the data behind the committed Stats (gene counts ~10-25% high), and the deposited 2-way_ANOVA.txt is not regenerable from the shared data because platform-prefixed column names break the deposited parser. This is an authors'-side deposit-hygiene/code-data inconsistency (q4 red), moderate in magnitude with direction preserved (q6 yellow), leaving an overall solid-but-deviating reproduction (q8 yellow).
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.