Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Large-scale integration of single-cell transcriptomic data captures transitional progenitor states in mouse skeletal muscle regeneration.

Commun Biol · 2021
L1 92/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
92/100
Reproducibility score
1.0 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 83% of all assessed papers rank 179 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (1:1, no fabrication concern). Paper = McKellar et al. 2021 Commun Biol, the scMuscle integrated atlas (365,011 cells/nuclei from 111 mouse skeletal-muscle samples). This room's assigned accession GSE159500 = 2 of the 23 newly generated samples (GSM4831162=D2_Ev3, GSM4831163=D7_Ev3; 10x Chromium v3, notexin TA injury). EVERY in-scope pipeline-derived result reproduced: (C1) from-raw read counts EXACT (STARsolo processed exactly 208,607,334 / 205,196,263 reads, matching SRA/ENA); (C2) from-raw cell calls WITHIN-TOL (STARsolo EmptyDrops_CR 4846/4917 vs deposited CellRanger 4836/4889 = +0.21%/+0.57% -- sub-1% agreement between two independent cell-callers on the same mm10-2.1.0 27,998-gene reference; 97.5-97.8% valid barcodes); (C3) deposited filtered-matrix dims EXACT (27998x4836 / 27998x4889); (C4) final QC'd annotated cells EXACT (3537/3416 from meta.csv); (S1) headline atlas size EXACT -- the deposited Dryad integrated Seurat object holds exactly 365,011 cells across (S2) 111 samples (confirmed). Two SLURM compute jobs ran on «our HPC» (STARsolo re-quant + atlas load). NOT ATTEMPTED (out of scope): full from-scratch 111-sample re-integration (88 of 111 samples are external public accessions; only 2 are in GSE159500) and the 503,929 pre-QC barcode total (not derivable from the post-QC object). Both datasets profiled (GSE159500 + Dryad atlas, both grade A, delivers_promised=yes). All reported numbers are derivable from the shipped data/code -- no possible-fabrication flags.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 92
    assessed: 2026-06-22 ⛓ 80f83edd57a4
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether large-scale integration of single-cell and single-nucleus transcriptomic datasets can overcome the poor sampling of rare, transient myogenic progenitor cell states in skeletal muscle regeneration and provide spatial context for their occurrence.

Core claims
  • Large-scale integration of 111 sc/snRNAseq datasets captures rare, transitional myogenic progenitor states (commitment and fusion) that are poorly represented in individual datasets. finding
  • Harmony outperforms Scanorama and BBKNN for batch-correcting heterogeneous muscle sc/snRNAseq datasets based on speed, marker gene consistency, and preservation of continuous myogenic structure. finding
  • Integration resolves cell subtypes including endothelial cells by vessel-type origin, functionally distinct fibro-adipogenic progenitors (FAPs), and multiple immune cell populations. finding
  • The number of genes expressed per cell increases through myogenic progenitor states and is suppressed again in mature myofibers. finding
  • Novel surface markers (e.g., Timp1, Clcn5, Igfbp3, Bst2) and transcription factors (e.g., Purb, Zbtb18, Mycl, Scx) specifically mark committed and fusing myogenic cells. finding
  • Spatial RNA sequencing combined with the integrated single-cell reference enables high-resolution local deconvolution of cell subtypes in injured muscle. method
  • Single-cell and single-nucleus preparations show systematic, largely cell-type-independent differences in detection of lncRNAs, mitochondrial, ribosomal, and dissociation-stress genes. finding
  • The integrated dataset and an interactive web tool are provided as a public resource for the field. resource
Experimental setups
Assay System Perturbation Readout Platform
scRNAseq (10x Chromium v3) adult (7 mo) and aged (20 mo) C57BL/6J mouse tibialis anterior muscle notexin injury (multiple timepoints within 1 week) vs. uninjured control single-cell transcriptomes, cell type/subtype identification 10x Chromium v3
scRNAseq/snRNAseq (public data integration) mouse skeletal muscle, various ages (10 days–30 months) cardiotoxin or notexin injury, 0.5–21 dpi, or none cell composition and transcriptomic dynamics across conditions 10x Chromium v2, v3, v3.1
Batch-correction/integration benchmarking merged sc/snRNAseq compendium (365,011 cells) none (computational comparison) integration quality (marker gene expression, local/global structure, sc/sn resolution) Harmony, Scanorama, BBKNN
Differential gene expression (Wilcoxon Rank Sum test) 84,383 myogenic cells, PHATE-binned along differentiation trajectory none genes differentially expressed across myogenic differentiation stages
Pseudotime/trajectory analysis (PHATE dimensional reduction) 84,383 myogenic and myofiber cells none continuous myogenic differentiation trajectory (quiescence to maturation) PHATE
Spatial RNA sequencing mouse muscle, 3 timepoints post-injury injury (timepoint comparison) local deconvolution of cell subtypes using integrated dataset as reference
Ligand-receptor co-expression analysis integrated sc/snRNAseq compendium injury response comparison dynamic cell-cell interaction patterns
Key results
  • Integration of 111 samples yielded 365,011 cells/nuclei after quality control from 503,929 raw barcodes
  • Harmony was the only method to resolve sc/sn differences across all cell types and maintain continuous myogenic structure
  • 84,383 myogenic cells selected; PHATE dimension 1 captured 95.6% of variance and served as differentiation pseudotime 95.6% variance
  • Committed and fusing myogenic cells (bins 8-18) comprised only 0.65% of the compendium, totaling 2,366 cells across 92 of 111 samples 0.65%
  • Genes expressed per cell (normalized to sequencing saturation) increase from quiescence through differentiation, then decline in mature myofibers
  • 2,031 genes differentially expressed across differentiation bins after filtering (adjusted p<10^-10, log2FC>0.5); 67 surface marker genes and 12 transcription factors identified as specific to committed/fusing cells
  • Single-nucleus data showed increased intronic/intergenic reads, lncRNA detection, and decreased mitochondrial/ribosomal transcripts versus single-cell data
  • Average of ~5 committed myoblasts and ~15 fusing myocytes captured per sample, representing 0.2% and 0.5% of all cells respectively 0.2% and 0.5%
Key statistics
  • count 365,011 cells/nuclei after QC and integration (final integrated compendium size from 111 samples)
  • count 503,929 total cell barcodes before QC (raw compendium prior to filtering)
  • count 84,383 myogenic and myofiber cells (cells used for PHATE differentiation trajectory)
  • count 2,366 committed or fusing cells (PHATE bins 8-18, representing 0.65% of compendium)
  • other PHATE_Harmony_1 accounts for 95.6% of variance (used as pseudotime proxy for myogenic differentiation)
  • pvalue adjusted p-value < 10^-10, average log2-fold-change > 0.5 (filtering threshold for differential gene expression across PHATE bins)
  • count 2,031 differentially expressed genes across differentiation (Wilcoxon Rank Sum test between PHATE bins)
  • count 67 surface marker genes; 12 transcription factors (genes enriched specifically in committed and fusing myogenic cells)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study integrates 111 mouse skeletal muscle sc/snRNAseq datasets (23 newly collected, 88 public) totaling ~365,000 cells after quality control, using Harmony batch correction selected over Scanorama and BBKNN based on qualitative and marker-gene criteria. Differential gene expression across 25 evenly spaced bins of a PHATE-derived myogenic differentiation axis was tested with the Wilcoxon Rank Sum test, filtered by adjusted p-value and log2-fold-change thresholds. Results were summarized descriptively through UMAP and PHATE visualizations, canonical marker gene violin/dot plots, boxplots of transcriptomic diversity, and a Simpson diversity index to assess batch-correction mixing quality.

Replicationbiological Sample size23 newly collected samples plus 88 publicly available datasets; 111 individual samples total yielding 365,011 cells/nuclei after QC; no formal power calculation described GroupsMyogenic differentiation states (25 PHATE bins); sc vs. sn preparation; Chromium chemistry version (v2, v3, v3.1); injury time-point (0–21 dpi); age (7 mo adult vs. 20 mo aged); injury model (notexin vs. cardiotoxin) Pairingunpaired Randomization/blindingnot stated DispersionIQR Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionadjusted p-value threshold applied; specific correction method not named in available text
Statistical tests used
Test Applied to n Assumptions
Wilcoxon Rank Sum test Differential gene expression between each of 25 PHATE bins along the myogenic differentiation axis 84,383 myogenic cells distributed across 25 bins not stated
Simpson's diversity index Assessment of sample-source mixing per PHATE bin to evaluate Harmony batch-correction quality 84,383 myogenic cells across 111 samples na
Approaches that could also have been used
  • Differential gene expression across differentiation bins was tested with the Wilcoxon Rank Sum test, treating individual cells as independent observations
    Could also: Pseudo-bulk approaches (e.g., DESeq2 or edgeR on per-sample aggregated counts) could also be used to test bin-level differences — Pseudo-bulk methods account for the non-independence of cells originating from the same biological sample and have been shown to better control false-positive rates in multi-sample single-cell DGE comparisons
  • Batch-correction method selection (Harmony over Scanorama and BBKNN) was based on qualitative assessment of UMAP/PHATE visualizations and marker gene consistency
    Could also: Quantitative integration benchmarking metrics such as kBET, LISI (Local Inverse Simpson's Index), or Average Silhouette Width (ASW) could also be computed for each method — Quantitative metrics provide a reproducible, method-independent basis for comparing integration quality and enable readers to assess mixing independently of visualization choices
  • The specific multiple-testing correction procedure underlying the 'adjusted p-value' threshold is not named in the available text
    Could also: Benjamini-Hochberg FDR correction with an explicit target alpha (e.g., FDR < 0.05) is a standard and widely reported approach for scRNAseq DGE — Naming the correction method and its target error rate aids reproducibility and allows direct comparison with results from other published datasets
  • Transcriptomic diversity (genes detected per cell) was normalized to per-sample sequencing saturation before comparing across bins
    Could also: Downsampling (rarefaction) to a common UMI depth per cell before computing gene detection rates could also equalize sampling effort — Rarefaction is an established ecology-derived approach that directly standardizes sequencing effort without relying on saturation estimates, and its properties are well understood in diversity analysis
  • Dispersion of genes-per-cell across differentiation bins is summarized with boxplots (median and quartiles)
    Could also: 95% confidence intervals around the per-bin mean could additionally be reported alongside or instead of boxplots — CIs convey both within-bin variability and uncertainty in the mean estimate, complementing the distributional summary of boxplots; many journals now recommend CIs for small or heterogeneous group sizes
  • Simpson's diversity index was computed at the bin level on sample identifiers to assess batch-correction quality
    Could also: Cell-neighborhood-level mixing metrics such as the Local Inverse Simpson's Index (LISI) or Shannon entropy per cell could also quantify integration quality — Neighborhood-level metrics capture local mixing structure more granularly than bin-level summaries and are increasingly reported as standard in single-cell integration benchmarks
Software: Seurat · Harmony · Scanorama · BBKNN · SoupX · DoubletFinder · PHATE

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34773081

Paper: McKellar et al. 2021, Commun Biol 4:1280. "Large-scale integration of single-cell transcriptomic data captures transitional progenitor states in mouse skeletal muscle regeneration." (scMuscle) PMCID PMC8589952 · DOI 10.1038/s42003-021-02810-x

This RU's data accession: GEO GSE159500 — "Single-cell transcriptomic sampling of regenerating mouse muscle tissue". This is a small slice of the paper's new data: 2 samples from 7-month-old mice, 10x Chromium v3, notexin injury timepoints (D2_Ev3, D7_Ev3). The paper's other new data live under GSE162172 (20 mo scRNAseq) and GSE161318 (spatial/Visium); the full atlas also folds in 88 public datasets. The brief assigns GSE159500 to this room, so the in-scope reproduction is anchored on those 2 samples.

Code:

  • Brief's code_url = github.com/rvalieris/parallel-fastq-dump — the SRA→FASTQ download utility named in the paper's Methods. Per P16 this counts: a third-party tool applied to the paper's own data is a valid reproduction.
  • Authors' own analysis repo = github.com/mckellardw/scMuscle (Seurat v3.2.1 + Harmony v1.0 integration, SoupX, DoubletFinder, PHATE, CellChat). README is thin on a runnable per-sample pipeline; the heavy lifting (Cell Ranger / count matrices) is described in the paper Methods, not scripted in the repo.

In scope (pipeline-derived, attempted)

# Result Pipeline Feasibility
C1 Raw read counts per run (SRR12820638 D2, SRR12820639 D7) parallel-fastq-dump SRA→FASTQ, count reads HIGH — direct, cheap
C2 Per-sample cell count vs deposited matrix (GSE159500_RAW.tar) STARsolo (10x v3) re-quantification from raw FASTQ → cell calling; compare to authors' deposited filtered matrix MEDIUM — needs mouse ref + STAR index + EmptyDrops cell calling; depends on barcode read (R1) surviving on SRA
C3 Genes detected / matrix dimensions of deposited GSE159500 matrices load deposited CSV/MTX/TSV, report rows×cols HIGH — direct profiling

Stretch (attempted if compute allows)

# Result Approach
S1 Atlas headline 365,011 cells/nuclei (Results) + cluster count download processed Seurat object from Dryad (doi:10.5061/dryad.t4b8gtj34), count cells/clusters. NOTE: this verifies the deposited artifact, not a from-scratch re-integration of 111 samples.

Out of scope (not attempted, stated why)

  • Full atlas re-integration of 111 samples (23 new + 88 public, ~365k cells, Harmony): requires gathering 88 external accessions across 14 groups + very large compute; not feasible as a faithful 1:1 within this room. GSE159500 contributes only 2 of the 111 samples.
  • Wet-lab / experimental results (notexin injury, dissociation, library prep): not computational.
  • Spatial RNA-seq (Visium/Slide-seq) results: different accession (GSE161318), out of this RU's data scope.
  • CellChat / BayesPrism / PHATE downstream biological interpretations: depend on the full integrated object; not pinnable to GSE159500's 2 samples.

Reproduction strategy

  1. front1 («infra»): download GSE159500_RAW.tar (deposited matrices = authors' "truth"), prefetch SRR12820638/39, parallel-fastq-dump → FASTQ, fetch mouse reference (GRCm38/mm10 + GTF) and 10x v3 whitelist.
  2. SLURM (compute only): build STAR index, run STARsolo per sample (CB=16, UMI=12, EmptyDrops_CR cell filter), emit filtered matrix.
  3. Compare reproduced cell/gene counts vs deposited matrices (C2/C3) and read counts vs SRA metadata (C1). Grade provisionally.
Figures / tables: Fig.1
C1a
Reported
208,607,334 reads (SRR12820638 D2_Ev3)
Reproduced
208,607,334 (STARsolo from-raw 'Number of Reads', «job»)
exact
C1b
Reported
205,196,263 reads (SRR12820639 D7_Ev3)
Reproduced
205,196,263 (STARsolo from-raw 'Number of Reads', «job»)
exact
C2a
Reported
4836 cells (D2_Ev3 CellRanger filtered)
Reproduced
4846 (STARsolo EmptyDrops_CR from raw FASTQ; +10, +0.21%)
within tolerance
C2b
Reported
4889 cells (D7_Ev3 CellRanger filtered)
Reproduced
4917 (STARsolo EmptyDrops_CR from raw FASTQ; +28, +0.57%)
within tolerance
C3a
Reported
27998 x 4836 (D2_Ev3 filtered matrix dims)
Reproduced
27998 x 4836
exact
C3b
Reported
27998 x 4889 (D7_Ev3 filtered matrix dims)
Reproduced
27998 x 4889
exact
C4a
Reported
3537 final QC'd annotated cells (D2_Ev3)
Reproduced
3537
exact
C4b
Reported
3416 final QC'd annotated cells (D7_Ev3)
Reproduced
3416
exact
S1
Reported
365,011 cells and nuclei (integrated scMuscle atlas)
Reproduced
365,011 (Dryad scMuscle.slim.seurat meta.data rows, «job»)
exact
S2
Reported
111 samples; 503,929 pre-QC barcodes
Reproduced
111 samples confirmed (sample col = 111 unique); 503,929 uncheckable from post-QC object
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 92/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

621.5 k
tokens (I/O) · 46.1 M incl. cache
161 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.