Immunoregulatory Roles of Tumor-Originated Pericytes Identified by Single-Cell Analysis in Glioblastoma.
Part of the results reproduced; minor but material deviations remained.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Any deviation was negligible
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.
▸Reproduction agent’s raw note
DROP. Not reproducible as described, for two compounding reasons. (1) data_restricted (primary): every reported pipeline number derives from the authors' in-house CD146+ pericyte scRNA-seq of 5 GBM patients, deposited under BioProject PRJCA025812 -> HRA007338 in GSA-Human (NGDC). The API record shows isControlledAccess:1, isFilePublic:0, dacId:4207 -- metadata public but files require a Data Access Request/Agreement to the submitting DAC, so they cannot be obtained autonomously. The UMI RNA-seq set (PRJCA035005) is likewise not file-public. (2) no_code (secondary): the paper states verbatim 'This paper does not report original code.' The harvested code link github.com/theMILOlab/SPATAData is a FALSE POSITIVE -- it is the public spatial-data resource the authors downloaded published datasets FROM (Freiburg MILO lab R package), not this Chongqing group's analysis pipeline; no tool versions or parameters are given for the ~17 named packages. NOT a 1:1 reproduction and NOT a divergent reproduction -- a genuine access drop. What I did NOT attempt and why: re-clustering the 8 public reference datasets (GSE141552/GSE138794/GSE162631/GSE103224/GSE141383/GSE182109/GSE163108/GSE173278) would not reproduce any reported value because they are atlas-integrated with the in-house data and no reported number is attributable to a public set alone (no expected result); per brief rule 6 I did not fabricate a public-data result to avoid the drop. No «our HPC»/SLURM compute was submitted -- there is no reproducible target. The 9 headline claims (C1-C9) are recorded in claims.tsv as reported-but-blocked; grade 'partial' here denotes 'identified & pinned but not reproducible due to access', with reproduced values intentionally empty. A human with approved GSA-Human access to HRA007338 could later attempt them.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-15 ⛓ df43bda15b04
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe origin and functional heterogeneity of pericytes in glioblastoma (GBM) are unclear; the study tests whether GBM pericytes arise from both tumor and normal origins and whether distinct pericyte subpopulations (e.g., CD44High) exert immunoregulatory roles by influencing tumor-associated macrophages and tumor growth.
- ★ Human primary GBMs contain both tumor-originated pericytes (T-PCs) and normal-originated pericytes (N-PCs) with distinctive cell-intrinsic features. finding
- ★ T-PCs and N-PCs are disparate populations with distinctive transcription factor regulons and marker expression (e.g., Desmin biased to T-PC, PDGFRβ biased to N-PC). finding
- ★ A T-PC metacluster marked by CD44 is closely associated with tumor-associated macrophages (TAMs). finding
- ★ Coimplantation of GSC-derived CD44High pericytes with GSCs promotes M2 polarization of TAMs and accelerates orthotopic GBM growth. mechanism
- CD146 labels most pericytes with strong specificity and is a suitable surface antigen for FACS sorting of GBM pericytes. method
- An integrated GBM atlas built from in-house and public scRNA-seq data reveals T-PC/N-PC variations in communication with endothelial and immune cells. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| scRNA-seq (single-cell RNA-sequencing) | CD45-CD146+ pericytes from 5 human primary GBM specimens | none | single-cell transcriptomes, cluster identity, cell origin | 10x Genomics Chromium |
| Fluorescence-activated cell sorting (FACS) | dissociated human primary GBM tissue | none | sorted CD45-CD146+ pericytes (proportion of total cells) | — |
| Immunofluorescence | frozen sections of human primary GBM | none | percent pericytes (CD146/PDGFRβ/NG2) adjacent to CD31+ endothelial cells | — |
| Copy number variation (CNV) analysis | 17 CD146+ scRNA-seq clusters (macrophage cluster as normal reference) | none | chromosomal amplification/deletion, CNV score to classify T-PC vs N-PC | infercnvpy |
| Deep-learning malignant-cell classification | 17 CD146+ subclusters from scRNA-seq | none | malignant vs normal cell identity confirming T-PC/N-PC | Cancer-finder algorithm |
| Regulon (transcription factor) activity analysis | T-PC and N-PC scRNA-seq populations | none | regulon activities (e.g., SOX2, E2F1, NR2F2, FOXC1) | — |
| Orthotopic xenograft / coimplantation | GSC-derived CD44High pericytes plus GSCs implanted in mouse brain | coimplantation (CD44High pericytes vs control) | M2 macrophage infiltration and tumor growth | — |
| Bulk transcriptome analysis | purified GSC-derived CD44High pericytes (in vitro) | none | transcriptome comparison to CD44High T-PC in GBM tissue | — |
- – 33,572 CD146+ vascular pericytes obtained from 5 GBM specimens, above 4700 pericytes per sample 33,572 cells (>4700/sample)
- – 35,677 high-quality CD146+ transcriptomes obtained, forming 17 clusters detected in all 5 patients 35,677 cells; 17 clusters
- – Chromosome 7 amplification and chromosome 10 loss detected in several CD146+ clusters, identifying T-PCs
- – T-PC and N-PC each constitute about half of total pericytes overall, but one population dominates in any individual patient ~50% each
- – PDGFRβ-high pericytes are mostly N-PCs and Desmin-expressing pericytes are mostly T-PCs, while CD146 is highly expressed in both
- – T-PC shows higher SOX2 and E2F1 regulon activity; N-PC shows higher NR2F2 and FOXC1 regulon activity
- ▲ Coimplantation of GSC-derived CD44High pericytes increased M2-subtype macrophage infiltration and accelerated tumor growth
- – CD146+ cells accounted for about 2-10% of total cells in individual GBM samples 2-10%
- count 33,572 vascular pericytes (CD146+ pericytes sorted from 5 GBM specimens)
- count 35,677 high-quality transcriptomes (CD146+ cells after preprocessing, ~5000-10000 per patient)
- count 17 distinct CD146+ clusters (unsupervised clustering, detected in all 5 patients)
- count >4700 pericytes per sample (per-sample pericyte yield)
- other 2-10% of total cells (CD45-CD146+ cells sorted from individual GBM samples)
- count n = 5 sections (immunofluorescent quantification of pericyte markers, two-tailed unpaired Student's t-test)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper used scRNA-seq of FACS-sorted CD146+ pericytes from 5 human primary GBM specimens, processed with Seurat and Harmony for batch correction, with unsupervised clustering and CNV inference (infercnvpy plus Cancer-finder deep learning) to classify tumor- versus normal-originated pericyte populations. Classical immunofluorescence quantifications of pericyte marker coverage were compared with two-tailed unpaired Student's t-tests reported as mean ± SD. Functional validation included orthotopic xenograft co-implantation experiments; statistical methods for those endpoints are not present in the provided text excerpt.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| two-tailed unpaired Student's t-test | Figure 1C: percentages of CD31+ endothelial cells covered by CD146, PDGFRβ, or NG2+ pericytes compared across marker groups | n = 5 sections | not stated |
| two-tailed unpaired Student's t-test | Figure 1E: percentage of CD146+ pericytes with PDGFRβ staining versus percentage of PDGFRβ+ pericytes with CD146 staining | n = 5 sections | not stated |
| CNV inference — infercnvpy (HMM-based copy number scoring) | Figure 1I–K: classification of 17 CD146+ clusters into tumor- or normal-originated pericytes by CNV score, using macrophage cluster C17 as normal reference | 35,677 cells across 5 patients | na |
| Cancer-finder deep learning algorithm | Figure S1J: confirmatory malignant/normal cell classification of T-PC and N-PC subclusters | 35,677 cells | na |
| Unsupervised graph-based clustering (Seurat pipeline with Harmony batch correction) | Figure 1F: identification of 17 distinct CD146+ cell clusters from integrated scRNA-seq data | 35,677 cells from 5 GBM patients | na |
| Transcription factor regulon activity analysis (tool not named in available text) | Figure 2A, Figure S2A: differential regulon activities (SOX2, E2F1, NR2F2, FOXC1) between T-PC and N-PC | 33,572 pericytes | na |
-
Two-tailed unpaired Student's t-test was used to compare immunofluorescence marker coverage percentages at n = 5 sections per group↳ Could also: A Mann-Whitney U (Wilcoxon rank-sum) test could also be applied — With n = 5 per group, the normality assumption underlying the t-test cannot be reliably assessed; a nonparametric rank-based test makes no distributional assumption and is a standard choice for small immunofluorescence quantification datasets
-
Multiple separate t-tests were performed across three pericyte marker comparisons in Figure 1C without a stated multiplicity correction↳ Could also: A one-way ANOVA followed by Tukey HSD post-hoc correction could also be applied to treat the marker comparisons as a single family — Grouping related comparisons into one omnibus test with a post-hoc correction maintains the family-wise error rate; this is a standard alternative when three or more conditions are compared simultaneously in the same experiment
-
Dispersion was reported as mean ± s.d. for groups of n = 5 sections↳ Could also: 95% confidence intervals could also be reported alongside or instead of SD — CIs convey both variability and precision of the estimated mean; for small n they can aid readers in directly assessing the uncertainty around the reported central estimate
-
CNV inference with infercnvpy was used to distinguish tumor- from normal-originated pericytes, validated by a second deep-learning classifier (Cancer-finder)↳ Could also: CopyKAT or Numbat (Bayesian/HMM frameworks) could also be applied for CNV calling from scRNA-seq — Different CNV inference tools use distinct statistical models and can produce complementary confidence metrics; benchmarking across two or more CNV-based methods alongside the expression classifier can further substantiate T-PC/N-PC assignments
-
Batch correction was performed with Harmony prior to unsupervised clustering of the 5-patient scRNA-seq dataset↳ Could also: Probabilistic integration methods such as scVI (variational autoencoder) or Scanorama could also be used for multi-sample integration — Different integration strategies make different assumptions about shared versus patient-specific variation; comparing cluster stability across methods is a common way to assess robustness of identified pericyte subpopulations across patients
-
Pericyte subpopulation identity was based on marker gene expression and CNV scoring within the scRNA-seq clustering↳ Could also: Somatic variant calling from scRNA-seq reads (e.g., Monopogen or SCReadCounts) could also be used as a mutation-based lineage assignment strategy — Mutation-based evidence is orthogonal to both CNV amplitude and expression-classifier signals, providing an independent molecular anchor for confirming tumor origin, particularly for cells with low-amplitude or ambiguous CNV profiles
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41001759
Title: Immunoregulatory Roles of Tumor-Originated Pericytes Identified by Single-Cell Analysis in Glioblastoma. (Chu C, … Bian XW, Wu HB, Zhang A, Zhou W.) Adv Sci (Weinh) 2025 · PMID 41001759 · PMCID PMC12713092 · DOI 10.1002/advs.202511856 Group: Army Medical University (Third Military Medical University), Chongqing — not the MILO lab.
Verbatim Data & Code Availability (from PMC full text)
"scRNA-seq datasets have been deposited at the Genome Sequence Archive at the National Genomics Data Center (Beijing, China) that are accessible under BioProject PRJCA025812. UMI RNA-seq datasets have been deposited … under BioProject PRJCA035005. This paper does not report original code."
What the reproducible (pipeline-derived) results are, and where the data lives
ALL headline computational results derive from the authors' in-house CD146+ pericyte scRNA-seq of 5 GBM patients, integrated with 8 public datasets into a "GBM atlas". Core reported pipeline outputs:
- 35,677 high-quality CD146+ transcriptomes (≈5–10k per patient); 17 CD146+ clusters.
- 33,572 cells retained as true pericytes after removing C4/C16 (endothelial), C17 (macrophage).
- T-PC vs N-PC ≈ 50/50 overall; T-PC → 10 clusters → 5 metaclusters (T-M1..T-M5); N-PC → 9 clusters → 3 metaclusters (N-M1..N-M3).
- CellChat: 4 communication modes; T-M2→macrophage top signals CSF1/SPP1/CCL2; N-M3→microglia top signals APOE/CCL3/TGFB1. CNV: chr7 amp / chr10 loss.
- Tools named (NO versions/params): Seurat, Harmony, infercnvpy, pySCENIC, aPEAR, clusterProfiler, scFEA, PyCoGAPS, CellChat, NMF, NicheNet, Tangram, Scanpy, COMMOT, mistyR, ROGUE, Cancer-Finder.
In scope vs out of scope
- In scope (pipeline-derived): the clustering / pericyte-calling / sub-clustering / CellChat / CNV results above. All require the in-house data (PRJCA025812 / HRA007338).
- Out of scope: wet-lab (IHC, THP-1 M2 polarization, conditioned-medium assays), TCGA survival stratification (external), protein-level validation.
Blocking screening findings
- Primary data is controlled-access. PRJCA025812 → HRA007338 in GSA-Human
(NGDC). API:
isControlledAccess:1,isFilePublic:0, DAC id 4207 — metadata public, files require a Data Access Request/Agreement to the submitting DAC. Cannot be downloaded autonomously. UMI RNA-seq PRJCA035005 not file-public either. - No original code. Paper states verbatim it "does not report original code."
The harvested code link
github.com/theMILOlab/SPATADatais not the paper's analysis pipeline — it is the public data resource from which the authors downloaded published spatial transcriptomics datasets. (SPATAData = Freiburg MILO lab R data-access package; unrelated to this Chongqing group's analysis.) This is a code-link false positive for "analysis code". - No public-data-only claim. The 8 public reference datasets (GSE141552, GSE138794, GSE162631, GSE103224, GSE141383, GSE182109, GSE163108, GSE173278) are obtainable, but every reported number is an atlas integration of in-house + public data — no specific reported value is attributable to a public dataset alone, so re-running a public set maps to no expected result to compare against.
Decision
DROP — drop_reason = data_restricted (controlled-access in-house data,
files not public, request-to-DAC), compounded by no_code (no original code).
No clearly-specified pipeline output is reproducible from publicly obtainable
inputs. No «our HPC» compute submitted (nothing to reproduce). Not fabricating a
public-data result to avoid the drop, per brief rule 6.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a legitimate access drop, not a quality or integrity problem. Every reported value (35677 cells, 17 clusters, 33572 true pericytes, T-PC/N-PC metaclusters, CellChat CSF1/SPP1/CCL2, CNV chr7 gain/chr10 loss) depends on in-house CD146+ scRNA-seq deposited as controlled-access (HRA007338, isFilePublic:0), and the paper explicitly reports no original code — the harvested theMILOlab/SPATAData link is a false-positive data resource. Nothing could therefore be derived, compared, or confirmed; the blocker is on the data-availability side, not our methodology nor an authors' defect. The reported numbers are internally consistent with no fabrication signal, so q5/q7/q8 are yellow (undetermined/limited) rather than red.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.