Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comparative profiling of skeletal muscle models reveals heterogeneity of transcriptome and metabolism.

Am J Physiol Cell Physiol · 2019
L1 64/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Reported values were directly comparable
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
64/100
Reproducibility score
0.6 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 25% of all assessed papers rank 854 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the PIPELINE, and the paper's biology reproduces 1:1 on the deposited data: re-running the deposited R pipeline (limma + per-gene 2-way ANOVA + clusterProfiler GO) on the repo's deposited processed matrices recovers the paper's named marker genes EXACTLY (C2C12 actin/myosin; L6 SLC2A1 + all five ETC complexes), 85-92% of the deposited model-specific gene lists, the Fig 1B PCA, the Fig 1C correlation/heterogeneity structure, and the GO themes (L6=proliferation, C2C12=muscle, HSMC=development). HOWEVER the deposited DATA does not exactly regenerate the deposited Stats/ OUTPUTS: gene-list counts run ~10-25% higher (deposited matrix is larger than the version behind the committed Stats), and Step03's 2-way ANOVA is NOT regenerable from the deposited GENENAME_batch.Rds at all -- the deposited column names are platform-prefixed (GPL17586_HumanCell_GSM..._UC1.cel.gz) and break the deposited gsub Species/Sample parser (-> degenerate design -> all-NA), and even after correcting the parser the exact p-values differ from shipped (deposited batch values differ from those behind the committed stats), though marker-gene significance is concordant. Both data and Stats were committed in the same commit (1bdd58f) yet are mutually inconsistent. Verdict PARTIAL: pipeline + biological conclusions reproduce; exact deposited numeric outputs do not, due to a deposit-hygiene/code-naming problem -- NOT evidence of fabrication (the reported biology is supported by the re-run). NOT attempted: all wet-lab metabolic assays (out of scope, bench experiments), the upstream raw-CEL->normalized RMA step (not shipped as runnable with pinned inputs), Step10 proteomics integration.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.1246757

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 64
    assessed: 2026-06-16 ⛓ be4cafbb4ca5
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Metabolic and contractile differences between commonly used skeletal muscle cell models (rat L6, mouse C2C12, primary human myotubes) may be reflected in and accounted for by differences in their transcriptomic landscapes.

Core claims
  • L6, C2C12, and primary human myotubes show substantial heterogeneity in transcriptomic and metabolic profiles finding
  • Actin- and myosin-coding mRNAs are enriched in C2C12, whereas L6 has highest expression of glucose transporters and the five mitochondrial electron transport chain complexes finding
  • Insulin-stimulated glucose uptake and oxidative capacity are greatest in L6 myotubes finding
  • Insulin-induced glycogen synthesis is highest in human myotubes (HSMCs), while C2C12 has higher baseline glucose oxidation finding
  • All three models respond to electrical pulse stimulation with increased glucose uptake and gene expression, but in slightly different manners finding
  • Cells and tissues correlate best within the same species, indicating strong species-specific differences in muscle mRNA composition finding
  • A merged cross-species transcriptomic profiling pipeline using public GEO data, human ortholog annotation, and ranked statistics enables comparison of muscle models method
  • Model-specific transcriptomic/metabolic signatures support recommendations for which model to use for a given hypothesis resource
Experimental setups
Assay System Perturbation Readout Platform
Transcriptomic profiling (microarray meta-analysis) Human, rat, mouse skeletal muscle tissue and L6/C2C12/HSMC myotubes (public GEO data) none (untreated control samples) mRNA expression, PCA clustering, correlation, gene ontology enrichment Affymetrix arrays GPL570/GPL6244/GPL17586 (human), GPL1355/GPL22740/GPL14746 (rat), GPL81/GPL1261/GPL17400 (mouse)
Proliferation (cell surface imaging) L6, C2C12, HSMC myoblasts none surface area occupied by cells over 72 h Inverted light microscope, ImageJ
BrdU incorporation proliferation assay L6, C2C12, HSMC myoblasts none BrdU incorporation (proliferation) BrdU Cell Proliferation Assay Kit no. 6813, Cell Signaling Technology
Radiolabeled 2-deoxyglucose uptake L6, C2C12, HSMC myotubes insulin (100 nmol/L) 2-[1,2-3H]deoxy-D-glucose uptake Liquid scintillation counter (WinSpectral 1414, Wallac)
Glycogen synthesis ([14C]glucose incorporation) L6, C2C12, HSMC myotubes insulin (100 nM) [14C]-labeled glycogen Liquid scintillation counter; D-[U-14C]glucose (PerkinElmer)
Glucose oxidation ([14C]CO2 capture) L6, C2C12, HSMC myotubes FCCP (0.5 µM) vs none [14C]-CO2 released Liquid scintillation counter; D-[U-14C]glucose
Fatty acid oxidation ([3H]palmitate) L6, C2C12, HSMC myotubes none (palmitate substrate) 3H-labeled oxidation products Liquid scintillation counter; [9,10-3H(N)]palmitate (PerkinElmer)
Mitochondrial respiration (Seahorse XF Mito Stress Test) L6, C2C12, HSMC myotubes oligomycin (1 µM), FCCP (2 µM), rotenone/antimycin A (0.75 µM) oxygen consumption rate (OCR) and extracellular acidification rate (ECAR) Agilent Seahorse XF analyzer
Key results
  • PCA showed clear segregation of cell models from skeletal muscle tissues, and best transcript correlations occurred within the same species
  • C2C12 enriched for actin/myosin genes; L6 highest for glucose transporters and the five ETC complexes
  • Insulin-stimulated glucose uptake and oxidative capacity greatest in L6 myotubes
  • Insulin-induced glycogen synthesis highest in human myotubes (HSMCs)
  • C2C12 myotubes had higher baseline glucose oxidation
  • All models responded to EPS with increased glucose uptake and gene expression, with model-specific differences
  • Significant overlap in top 100 expressed genes across models despite overall divergence top 100 genes
  • Gene ontology enrichment showed rat L6 enriched for metabolism- and proliferation-related pathways
Key statistics
  • count 12 healthy men, age 50.4 ± 10 yr, BMI 22.9 ± 2.4 kg/m2 (Primary human muscle cell donors (vastus lateralis biopsies))
  • other FDR < 0.05 (Significance threshold for differential expression / gene ontology)
  • count top 100 expressed genes (Genes compared for overlap across models (chord diagram))
  • count n=11 (C2C12), 13 (L6), 11 (HSMC) (Independent experiments/donors for glucose uptake assay)
  • count n=9 (C2C12), 8 (L6), 12 (HSMC) (Independent experiments/donors for glycogen synthesis assay)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

Transcriptomic comparisons used publicly available microarray data processed with robust multi-array and quantile normalization, with differential expression and ranked statistics computed in limma and gene ontology enrichment thresholded at a false discovery rate < 0.05. For the wet-lab metabolic and gene-expression assays, group means across three cell models were compared after testing normality with the Shapiro-Wilk test: normally distributed data were analyzed by ANOVA with Tukey's multiple comparison, and non-normal data by Kruskal-Wallis with Dunn's multiple comparison. Results were summarized as group averages of independent experiments/donors, with sample sizes and tests stated in figure captions.

Replicationmixed Sample sizeStated per assay in the text and figure captions as the number of independent experiments for C2C12 and L6 and the number of replicates/individual donors for HSMC; no formal power calculation described GroupsL6 vs C2C12 vs HSMC myotubes (and cells vs adult tissue, across mouse/rat/human, for transcriptomics) Pairingunpaired Randomization/blindingnot stated Dispersionunclear Multiplicity correctionTukey's HSD (parametric), Dunn's test (nonparametric) for multi-group comparisons; Benjamini-Hochberg-type false discovery rate (FDR < 0.05) for transcriptomic/GO analyses
Statistical tests used
Test Applied to n Assumptions
One-way ANOVA with Tukey's multiple comparison normally distributed metabolic/functional assay comparisons across L6, C2C12, and HSMC (per figure captions) independent experiments per cell model and donor-derived replicates for HSMC (varies by assay, e.g. glucose uptake 11 C2C12/13 L6/11 HSMC) stated
Kruskal-Wallis test with Dunn's multiple comparison non-normally distributed metabolic/functional assay comparisons across the three models independent experiments and donor replicates as stated per assay stated
Shapiro-Wilk normality test pre-test assessment of distribution for assay data to choose parametric vs nonparametric path na
limma differential expression with ranked statistics transcriptomic comparison of cells/tissues across human, rat, mouse platforms publicly available GEO samples per platform (Supplemental Table S1) not stated
clusterProfiler gene ontology enrichment genes differentially expressed in one model vs the other two (Fig. 2D) genes passing FDR < 0.05 not stated
Correlation analysis / principal component analysis clustering and correlation of transcript profiles across models and tissues (Fig. 1B, 1C) not stated
Approaches that could also have been used
  • Normality was assessed with the Shapiro-Wilk test, and the analysis branched to ANOVA/Tukey or Kruskal-Wallis/Dunn accordingly.
    Could also: A consistently applied nonparametric framework, or a single model with variance-stabilizing transformation, could also be used regardless of the normality test outcome. — With the modest per-group sample sizes here, normality tests have limited power; pre-specifying one approach or transforming the data can simplify interpretation and reduce dependence on a screening test.
  • The three cell models were compared with one-way ANOVA (or Kruskal-Wallis) treating each as an independent group.
    Could also: A mixed-effects model treating donor as a random effect (especially for HSMC, where some replicates come from the same individuals) could also be applied. — A mixed model can account for repeated sampling from the same donors and unequal replicate structure across models, which would more explicitly represent biological clustering.
  • Group results were summarized as averages of independent experiments and donor replicates.
    Could also: Reporting a 95% confidence interval or showing individual data points alongside the mean could also be presented. — Confidence intervals and dot plots convey both spread and uncertainty directly, which is often informative for small sample sizes.
  • Transcriptomic differential expression used limma with ranked statistics on microarray data merged across species-specific platforms.
    Could also: Surrogate variable analysis or explicit batch/platform covariates within limma could also be incorporated. — Modeling platform/species as covariates can help separate technical interarray variability from biological signal, a point the authors themselves note as a consideration.
  • GO enrichment thresholded genes at FDR < 0.05 and used clusterProfiler over-representation analysis.
    Could also: A threshold-free approach such as Gene Set Enrichment Analysis (GSEA) on the full ranked gene list could also be used. — GSEA uses all genes without a hard significance cutoff, which can detect coordinated shifts in pathways where individual genes do not pass the FDR threshold.
  • Pairwise model differences were controlled with Tukey's or Dunn's post hoc correction within each assay.
    Could also: Reporting effect sizes (e.g., standardized mean differences) alongside the corrected p-values could also be included. — Effect-size estimates communicate the magnitude of differences between models, complementing the significance testing for readers choosing a model.
Software: R 3.5.2 · R/limma · R/clusterProfiler · R/BioMart · GraphPad Prism 8.1 · ImageJ

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
181
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GPL1261 GEO in Methods (http://purl.org/orb/Methods)
also used by 2 papers:
GPL17586 GEO in Methods (http://purl.org/orb/Methods)
also used by 2 papers:
GPL1355 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GPL14746 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GPL17400 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GPL22740 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GPL81 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Fig 4Fig 3Fig 1CFig 1B
MS-C2C12-actin
Reported
actin enriched in C2C12 (qual.)
Reproduced
ACTN4/ACTN1/ACTR2/ACTN3 in C2C12 list (identical to shipped)
exact
MS-C2C12-myosin
Reported
myosin enriched in C2C12 (qual.)
Reproduced
MYL4/MYL1/MYL6/MYH1/MYH4/MYLK2 (identical to shipped)
exact
MS-L6-glut
Reported
glucose transporters highest in L6 (qual.)
Reproduced
SLC2A1 in L6 list (matches shipped)
exact
MS-L6-etc
Reported
five ETC complexes highest in L6 (qual.)
Reproduced
34 ETC genes spanning CI-CV (shipped 33)
within tolerance
MS-HSMC-n
Reported
2768 (deposited Stats)
Reproduced
3226 (87.6% of shipped recovered)
partial
MS-L6-n
Reported
2684 (deposited Stats)
Reproduced
3494 (85.2% recovered)
partial
MS-C2C12-n
Reported
2040 (deposited Stats)
Reproduced
2276 (92.0% recovered)
partial
ANOVA-table
Reported
Stats/2-way_ANOVA.txt (14749 genes)
Reproduced
all-NA from deposited data (naming bug); not regenerable
did not match
ANOVA-marker-sig
Reported
SLC2A4/MYH1 highly significant
Reproduced
SLC2A4 1e-104, MYH1 3e-71 (corrected parser); concordant
within tolerance
INTRA-pct
Reported
not printed in paper
Reproduced
Human 52.83% / Mouse 6.15% / Rat 44.92% (FDR<0.01)
partial
CORR-fig1c
Reported
6x6 Spearman heatmap 0.7-1.0 (Fig 1C)
Reproduced
6x6 matrix range 0.61-0.84, heterogeneity structure intact
partial
PCA-fig1b
Reported
PCA, models vs tissue (Fig 1B)
Reproduced
PC1 23.5% / PC2 18.8%, separation reproduced
partial
GO-enrich
Reported
L6=proliferation/metabolism, C2C12=muscle, HSMC=development (qual.)
Reproduced
top terms match (L6 mitosis/DNA-repl, C2C12 muscle contraction, HSMC ECM/morphogenesis); 56-70% term recovery
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 64/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

Re-running the deposited limma+ANOVA+clusterProfiler pipeline on the deposited matrices reproduces the paper's biology 1:1 qualitatively — named marker genes are exact and GO/PCA/correlation structure holds — so the central conclusion is confirmed and there is no fabrication signal. The deviations are entirely at the deposit/input level: the deposited Data_Processed is an expanded version that no longer matches the data behind the committed Stats (gene counts ~10-25% high), and the deposited 2-way_ANOVA.txt is not regenerable from the shared data because platform-prefixed column names break the deposited parser. This is an authors'-side deposit-hygiene/code-data inconsistency (q4 red), moderate in magnitude with direction preserved (q6 yellow), leaving an overall solid-but-deviating reproduction (q8 yellow).

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

305.9 k
tokens (I/O) · 26 M incl. cache
59 min
runtime · 0.1 CPU-h
2.1 GB
peak RAM
3 (1 failed)
HPC jobs
hummel
machine