Single-Cell Transcriptome Profiling Reveals Multicellular Ecosystem of Nucleus Pulposus during Degeneration Progression.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL reproduction of Tu et al. 2022 (NP scRNA-seq, GSE165722). HEADLINE FINDING: the paper's central '39,732 cells after QC' is NOT reproducible from the deposited per-sample matrices using the paper's OWN stated QC (<600 genes, >8% mito) -> 16,445 cells (33% of the 49,637 deposited); a 15-point threshold grid finds no combination near 39,732 -> flagged as possible-fabrication/under-specification. Reproduced exactly: 8 samples (C2), 2000-HVG parameter (X2). Reproduced qualitatively (3rd-party tool Seurat5+batchelor::fastMNN, P16-valid): multiple NPC subpopulations (~10 vs reported 6, C4a), presence of macrophage/neutrophil/G-MDSC lineages (C4b), and the G-MDSC depletion with degeneration (mild>>severe, ~4x vs paper's ~3x, C5) -- direction and order-of-magnitude hold but absolute numbers differ, downstream of the C1 cell-set discrepancy and Seurat-version drift (paper used 2.3.4). NOT attempted: wet-lab/clinical (Pfirrmann grading, flow/IHC/qPCR), FASTQ->matrix alignment (matrices deposited), trajectory/pySCENIC/CellPhoneDB. The qualitative biology holds; the exact published cell counts/proportions do not regenerate from the deposit + documented methods.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 59assessed: 2026-06-21 ⛓ 8a4457cbd630
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-21
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetBulk-tissue transcriptomics masks cellular heterogeneity in the nucleus pulposus (NP), so the paper uses single-cell RNA sequencing to resolve the cellular and molecular complexity of NP cells (NPCs) and infiltrating immune cells across progressive stages of intervertebral disc degeneration (IVDD).
- ★ Human NP contains six transcriptionally distinct NPC subpopulations: HT-CLNPs, effector NPCs, homeostatic NPCs, regulatory NPCs, fibroNPCs, and adhesion NPCs finding
- ★ FibroNPCs represent the dominant subpopulation in end-stage (grade V) degeneration finding
- ★ CD90+ NPCs act as progenitor cells in degenerative NP tissue finding
- ★ NP-infiltrating immune cells include a previously unrecognized population of granulocytic myeloid-derived suppressor cells (G-MDSCs) finding
- ★ Integrin alpha M (CD11b) and OLR1 are surface markers identifying NP-derived G-MDSCs resource
- ★ G-MDSCs are enriched in mildly degenerated (grade II/III) NP tissue relative to severely degenerated (grade IV/V) tissue finding
- ★ G-MDSCs exhibit immunosuppressive function and reduce NPC extracellular matrix degradation in vitro mechanism
- ★ Grade-associated transcription factors (e.g., AR, REL, LMX1A, PRDM1) and pathways (TNF, MAPK, Hippo, antigen processing/presentation) are linked to IVDD severity finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA sequencing | human NP tissue from 8 IVDD patients (grades II-V) | none (natural degeneration grade comparison) | transcriptome-wide gene expression per cell, cell subpopulation identification | BD Rhapsody system |
| RT-qPCR | human NPCs from discs of different degeneration grades | none | expression of subpopulation marker genes (FN1, CRTAC1, COL3A1, MMP2, MSMO1, HMGCS1) | — |
| immunohistochemistry | human NP tissue sections | none | protein-level expression of NPC subpopulation markers | — |
| QuSAGE gene set enrichment / GO analysis | scRNA-seq-derived NPC subpopulation transcriptomes | none | biological process/pathway enrichment per subpopulation | — |
| pseudotime trajectory analysis | HT-CLNP-I and HT-CLNP-II NPC subsets | none | developmental/progression trajectory between subsets | — |
| in vitro functional co-culture assay | isolated G-MDSCs and NPCs | co-culture (immune cell exposure) | immunosuppressive activity and NPC matrix degradation markers | — |
- – 39,732 cells profiled by scRNA-seq from 8 individuals across degeneration grades II-V n=39732 cells, 8 patients
- ▲ Proportion of adhesion NPCs and fibroNPCs increased with IVDD severity from grade II to V
- ▼ Proportion of homeostatic NPCs decreased with increasing degeneration severity
- ▲ FN1, CRTAC1, COL3A1, and MMP2 significantly elevated in NPCs from grade IV/V discs by RT-qPCR
- ▼ MSMO1 and HMGCS1 (homeostatic/effector markers) significantly decreased in grade IV/V discs
- ▼ G-MDSCs enriched in grade II/III (mild) NP tissue compared to grade IV/V (severe) tissue
- ▲ Multiple heat shock protein genes (HSPA1B, HSPH1, HSP90AA1, HSPA8, HSPA1A, HSPB8, HSPD1) elevated in severe vs mild degeneration
- ▲ INHBA and other TGF-beta/stress-related DEGs upregulated in severe degeneration avg_logFC=1.54 (INHBA)
- count 39,732 cells (total cells profiled by scRNA-seq across 8 patients)
- fold_change avg_logFC=1.542242, p=0 (INHBA upregulation in severe vs mild degeneration DEG analysis)
- fold_change avg_logFC=1.355454, p=0 (FHL2 upregulation in severe vs mild degeneration DEG analysis)
- pvalue 2.432 × 10^-11 (GO term 'sterol biosynthetic process' enrichment in effector NPCs)
- pvalue 5.795 × 10^-10 (GO term 'cholesterol biosynthetic process' enrichment in effector NPCs)
- count n=3 (RT-qPCR replicate number for validation of NPC atlas gene markers (mean ± SD))
- other lifetime prevalence up to 84% (background statistic on low back pain prevalence)
- other $85 billion (2008) (background statistic on socioeconomic burden of low back pain)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This single-cell RNA sequencing (scRNA-seq) study profiled 39,732 cells from nucleus pulposus (NP) tissues of eight patients across Pfirrmann degeneration grades II–V (two donors per grade). Unsupervised clustering identified six NPC subpopulations; differentially expressed genes (DEGs) between degeneration severity groups were identified with adjusted p-values and average log-fold changes reported. Pathway enrichment was assessed with QuSAGE and Gene Ontology analyses, and key findings were validated by RT-qPCR (n = 3, mean ± SD) and immunohistochemistry.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| DEG test (specific test not stated in provided text; output format — p_val, avg_logFC, pct.1, pct.2, p_val_adj — is characteristic of Seurat FindMarkers) | Severe (grades IV–V) vs mild (grades II–III) NPC comparison; Table 2 | Cells pooled across 8 donors; exact cell counts not stated in provided text | not stated |
| Unsupervised graph-based clustering (algorithm not named in provided text) | Identification of six NPC subpopulations; Figure 1B–C | 39,732 cells after QC filtering | na |
| QuSAGE (Quantitative Set Analysis for Gene Expression) | Functional enrichment of NPC subpopulation-specific gene expression; Figure 2A | — | not stated |
| Gene Ontology (GO) enrichment (specific statistical test not stated) | Biological process enrichment per NPC subpopulation; Table 3 | — | not stated |
| RT-qPCR (inferential test not stated; results described as 'significantly elevated/decreased') | Validation of representative NPC subpopulation markers across degeneration grades; Figure 1E | n = 3 (biological/technical replication not specified) | not stated |
| Pseudotime trajectory analysis (algorithm not named in provided text) | HT-CLNP-I vs HT-CLNP-II progression; Figure 2H | — | na |
-
DEG analysis was performed by comparing cells pooled across donors within severity groups, treating individual cells as independent observations↳ Could also: Pseudobulk aggregation (e.g., summing or averaging counts per donor, then applying DESeq2 or edgeR) could also be used, treating the donor as the unit of replication — With only two donors per grade, pseudobulk approaches explicitly model donor-level variance, which can reduce inflation of test statistics that arises when many cells from few donors are treated as independent replicates
-
RT-qPCR validation results across degeneration grades were described as statistically significant (Figure 1E, n = 3) but the inferential test applied was not named↳ Could also: A one-way ANOVA with a post-hoc correction (e.g., Tukey HSD) or a Kruskal–Wallis test with Dunn's post-hoc could also be applied across the four grade levels — Naming and reporting the test, its statistic, and the exact p-value supports reproducibility and lets readers assess whether the assumptions of the chosen test were met for n = 3 groups
-
Variability in qPCR data was summarized with SD (n = 3)↳ Could also: A 95% confidence interval could also convey the same spread and additionally communicate the precision of the mean estimate — For small n such as 3, CIs make the uncertainty of the point estimate more transparent and are increasingly requested by journals adopting estimation-statistics reporting guidelines
-
The relationship between degeneration grade and cell subpopulation proportion was described descriptively (Figure 1D)↳ Could also: An ordinal regression or Spearman correlation between grade (ordinal II–V) and subpopulation proportion could also formally quantify the grade-proportion relationship — A formal trend test would distinguish a monotonic grade-associated shift from random variation across the eight donors, and would complement the visual description in Figure 1D
-
Multiple GO enrichment p-values were reported in Table 3 without a stated multiple-testing correction across the many GO terms tested↳ Could also: Benjamini–Hochberg FDR correction across all tested GO terms could also be applied and reported alongside raw p-values — When testing hundreds of GO terms simultaneously, controlling FDR limits the expected proportion of false discoveries, which is standard practice in genome-wide enrichment analyses
-
Cell subpopulation clustering was performed and described without a formal evaluation of cluster stability or optimal cluster number↳ Could also: Silhouette analysis, bootstrapped cluster stability metrics, or assessment across a range of resolution parameters could also be reported — Documenting how cluster number and identity were chosen helps readers understand how robustly the six subpopulations are supported by the data and whether the results are sensitive to resolution parameter choice
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34825784
Paper: Tu J. et al., Single-Cell Transcriptome Profiling Reveals Multicellular Ecosystem of Nucleus Pulposus during Degeneration Progression. Adv Sci (Weinh) 2022. PMID 34825784 · PMCID PMC8787427 · DOI 10.1002/advs.202103631.
Data: GEO GSE165722 (BioProject PRJNA695536, SRA SRP303685).
- Platform: BD Rhapsody scRNA-seq (GPL24676, Illumina NovaSeq 6000), Homo sapiens.
- 8 samples GSM5048708–GSM5048715 (Sample1..8), Pfirrmann grades II–V (2 each).
- Processed data deposited as per-sample matrices in
GSE165722_RAW.tar(~67 MB):GSMxxxx_SampleN.counts.tsv.gz(gene × cell count matrix) +GSMxxxx_SampleN.cellname.txt.gz(cell names). Raw FASTQ only in SRA.
Code link in brief: https://github.com/alanocallaghan/scater — this is the
third-party scater Bioconductor package (the paper used its fastMNN for batch
correction), not the authors' own analysis repo. Per brief rule P16, applying an
existing third-party tool to the paper's own data is an equally valid reproduction.
Default branch devel, HEAD 8c547447c1451ed6fd95c341cd777e12cee9a439 (checked 2026-06-19).
Methods pipeline reported (Methods + Fig S1)
fastp → UMI-tools → STAR (Ensembl 91) → Seurat v2.3.4 (normalize/scale) →
scater fastMNN (batch correction) → graph clustering → manual annotation;
trajectory: scanpy 1.6.0 / Monocle2 / Slingshot / pySCENIC 0.9.5; CellPhoneDB.
QC: exclude cells with <600 genes and >8% mitochondrial content;
PCA on top 2000 HVGs.
In scope (pipeline-derived, reproducible from deposited matrices)
- C1 — Total cells after QC = 39,732 (across 8 samples). Count matrix columns; apply QC (<600 genes, >8% mito) and compare. CLEAR 1:1 data point.
- C2 — N samples = 8 (matrices present). Structural check.
- C3 — Per-sample / total dimensions, genes detected, QC pass rates. Profiling.
- C4 — Clustering → number of NPC subpopulations (6) + immune cell types. Harder: requires Seurat+fastMNN+clustering; annotation is author-judgment.
- C5 — G-MDSC proportion 6.5% (mild II–III) vs 2.2% (severe IV–V), Fig 4E. Hardest: depends on full annotation. Attempt if clustering reproduces.
Out of scope (wet-lab / clinical / manual)
- Pfirrmann grading & patient ages (Table 1) — clinical, not pipeline.
- Flow cytometry / IHC / qPCR validation, animal work — wet-lab.
- Manual biological naming of subpopulations — author judgment (we report cluster count + markers, not adjudicate the names).
- Full FASTQ→matrix alignment (fastp/STAR) — possible from SRA but heavy and redundant since processed matrices are deposited; not the primary path.
Reproduction strategy
- Download
GSE165722_RAW.taron «host» → «infra» (67 MB). Untar. - C1–C3 light QC pass: parse 8 matrices, count cells/genes, mito %, apply QC.
- C4–C5 SLURM job: Seurat/scater pipeline (normalize → HVG → fastMNN → cluster), count clusters, marker-based coarse annotation, G-MDSC proportion.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.