Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Immunoregulatory Roles of Tumor-Originated Pericytes Identified by Single-Cell Analysis in Glioblastoma.

Adv Sci (Weinh) · 2025
L1 No data access 2/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Any deviation was negligible
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
No data access Data access not granted

This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.

Reproduction agent’s raw note

DROP. Not reproducible as described, for two compounding reasons. (1) data_restricted (primary): every reported pipeline number derives from the authors' in-house CD146+ pericyte scRNA-seq of 5 GBM patients, deposited under BioProject PRJCA025812 -> HRA007338 in GSA-Human (NGDC). The API record shows isControlledAccess:1, isFilePublic:0, dacId:4207 -- metadata public but files require a Data Access Request/Agreement to the submitting DAC, so they cannot be obtained autonomously. The UMI RNA-seq set (PRJCA035005) is likewise not file-public. (2) no_code (secondary): the paper states verbatim 'This paper does not report original code.' The harvested code link github.com/theMILOlab/SPATAData is a FALSE POSITIVE -- it is the public spatial-data resource the authors downloaded published datasets FROM (Freiburg MILO lab R package), not this Chongqing group's analysis pipeline; no tool versions or parameters are given for the ~17 named packages. NOT a 1:1 reproduction and NOT a divergent reproduction -- a genuine access drop. What I did NOT attempt and why: re-clustering the 8 public reference datasets (GSE141552/GSE138794/GSE162631/GSE103224/GSE141383/GSE182109/GSE163108/GSE173278) would not reproduce any reported value because they are atlas-integrated with the in-house data and no reported number is attributable to a public set alone (no expected result); per brief rule 6 I did not fabricate a public-data result to avoid the drop. No «our HPC»/SLURM compute was submitted -- there is no reproducible target. The 9 headline claims (C1-C9) are recorded in claims.tsv as reported-but-blocked; grade 'partial' here denotes 'identified & pinned but not reproducible due to access', with reproduced values intentionally empty. A human with approved GSA-Human access to HRA007338 could later attempt them.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-15 ⛓ df43bda15b04
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The origin and functional heterogeneity of pericytes in glioblastoma (GBM) are unclear; the study tests whether GBM pericytes arise from both tumor and normal origins and whether distinct pericyte subpopulations (e.g., CD44High) exert immunoregulatory roles by influencing tumor-associated macrophages and tumor growth.

Core claims
  • Human primary GBMs contain both tumor-originated pericytes (T-PCs) and normal-originated pericytes (N-PCs) with distinctive cell-intrinsic features. finding
  • T-PCs and N-PCs are disparate populations with distinctive transcription factor regulons and marker expression (e.g., Desmin biased to T-PC, PDGFRβ biased to N-PC). finding
  • A T-PC metacluster marked by CD44 is closely associated with tumor-associated macrophages (TAMs). finding
  • Coimplantation of GSC-derived CD44High pericytes with GSCs promotes M2 polarization of TAMs and accelerates orthotopic GBM growth. mechanism
  • CD146 labels most pericytes with strong specificity and is a suitable surface antigen for FACS sorting of GBM pericytes. method
  • An integrated GBM atlas built from in-house and public scRNA-seq data reveals T-PC/N-PC variations in communication with endothelial and immune cells. resource
Experimental setups
Assay System Perturbation Readout Platform
scRNA-seq (single-cell RNA-sequencing) CD45-CD146+ pericytes from 5 human primary GBM specimens none single-cell transcriptomes, cluster identity, cell origin 10x Genomics Chromium
Fluorescence-activated cell sorting (FACS) dissociated human primary GBM tissue none sorted CD45-CD146+ pericytes (proportion of total cells)
Immunofluorescence frozen sections of human primary GBM none percent pericytes (CD146/PDGFRβ/NG2) adjacent to CD31+ endothelial cells
Copy number variation (CNV) analysis 17 CD146+ scRNA-seq clusters (macrophage cluster as normal reference) none chromosomal amplification/deletion, CNV score to classify T-PC vs N-PC infercnvpy
Deep-learning malignant-cell classification 17 CD146+ subclusters from scRNA-seq none malignant vs normal cell identity confirming T-PC/N-PC Cancer-finder algorithm
Regulon (transcription factor) activity analysis T-PC and N-PC scRNA-seq populations none regulon activities (e.g., SOX2, E2F1, NR2F2, FOXC1)
Orthotopic xenograft / coimplantation GSC-derived CD44High pericytes plus GSCs implanted in mouse brain coimplantation (CD44High pericytes vs control) M2 macrophage infiltration and tumor growth
Bulk transcriptome analysis purified GSC-derived CD44High pericytes (in vitro) none transcriptome comparison to CD44High T-PC in GBM tissue
Key results
  • 33,572 CD146+ vascular pericytes obtained from 5 GBM specimens, above 4700 pericytes per sample 33,572 cells (>4700/sample)
  • 35,677 high-quality CD146+ transcriptomes obtained, forming 17 clusters detected in all 5 patients 35,677 cells; 17 clusters
  • Chromosome 7 amplification and chromosome 10 loss detected in several CD146+ clusters, identifying T-PCs
  • T-PC and N-PC each constitute about half of total pericytes overall, but one population dominates in any individual patient ~50% each
  • PDGFRβ-high pericytes are mostly N-PCs and Desmin-expressing pericytes are mostly T-PCs, while CD146 is highly expressed in both
  • T-PC shows higher SOX2 and E2F1 regulon activity; N-PC shows higher NR2F2 and FOXC1 regulon activity
  • Coimplantation of GSC-derived CD44High pericytes increased M2-subtype macrophage infiltration and accelerated tumor growth
  • CD146+ cells accounted for about 2-10% of total cells in individual GBM samples 2-10%
Key statistics
  • count 33,572 vascular pericytes (CD146+ pericytes sorted from 5 GBM specimens)
  • count 35,677 high-quality transcriptomes (CD146+ cells after preprocessing, ~5000-10000 per patient)
  • count 17 distinct CD146+ clusters (unsupervised clustering, detected in all 5 patients)
  • count >4700 pericytes per sample (per-sample pericyte yield)
  • other 2-10% of total cells (CD45-CD146+ cells sorted from individual GBM samples)
  • count n = 5 sections (immunofluorescent quantification of pericyte markers, two-tailed unpaired Student's t-test)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper used scRNA-seq of FACS-sorted CD146+ pericytes from 5 human primary GBM specimens, processed with Seurat and Harmony for batch correction, with unsupervised clustering and CNV inference (infercnvpy plus Cancer-finder deep learning) to classify tumor- versus normal-originated pericyte populations. Classical immunofluorescence quantifications of pericyte marker coverage were compared with two-tailed unpaired Student's t-tests reported as mean ± SD. Functional validation included orthotopic xenograft co-implantation experiments; statistical methods for those endpoints are not present in the provided text excerpt.

Replicationbiological Sample size5 human primary GBM surgical specimens for scRNA-seq; n = 5 sections for immunofluorescence quantification; sample sizes for in vivo xenograft and in vitro experiments not present in available text excerpt GroupsT-PCs vs N-PCs; CD146 vs PDGFRβ vs NG2 pericyte marker coverage; CD44High vs control GSC-derived pericytes in orthotopic co-implantation Pairingunpaired Randomization/blindingnot stated DispersionSD Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
two-tailed unpaired Student's t-test Figure 1C: percentages of CD31+ endothelial cells covered by CD146, PDGFRβ, or NG2+ pericytes compared across marker groups n = 5 sections not stated
two-tailed unpaired Student's t-test Figure 1E: percentage of CD146+ pericytes with PDGFRβ staining versus percentage of PDGFRβ+ pericytes with CD146 staining n = 5 sections not stated
CNV inference — infercnvpy (HMM-based copy number scoring) Figure 1I–K: classification of 17 CD146+ clusters into tumor- or normal-originated pericytes by CNV score, using macrophage cluster C17 as normal reference 35,677 cells across 5 patients na
Cancer-finder deep learning algorithm Figure S1J: confirmatory malignant/normal cell classification of T-PC and N-PC subclusters 35,677 cells na
Unsupervised graph-based clustering (Seurat pipeline with Harmony batch correction) Figure 1F: identification of 17 distinct CD146+ cell clusters from integrated scRNA-seq data 35,677 cells from 5 GBM patients na
Transcription factor regulon activity analysis (tool not named in available text) Figure 2A, Figure S2A: differential regulon activities (SOX2, E2F1, NR2F2, FOXC1) between T-PC and N-PC 33,572 pericytes na
Approaches that could also have been used
  • Two-tailed unpaired Student's t-test was used to compare immunofluorescence marker coverage percentages at n = 5 sections per group
    Could also: A Mann-Whitney U (Wilcoxon rank-sum) test could also be applied — With n = 5 per group, the normality assumption underlying the t-test cannot be reliably assessed; a nonparametric rank-based test makes no distributional assumption and is a standard choice for small immunofluorescence quantification datasets
  • Multiple separate t-tests were performed across three pericyte marker comparisons in Figure 1C without a stated multiplicity correction
    Could also: A one-way ANOVA followed by Tukey HSD post-hoc correction could also be applied to treat the marker comparisons as a single family — Grouping related comparisons into one omnibus test with a post-hoc correction maintains the family-wise error rate; this is a standard alternative when three or more conditions are compared simultaneously in the same experiment
  • Dispersion was reported as mean ± s.d. for groups of n = 5 sections
    Could also: 95% confidence intervals could also be reported alongside or instead of SD — CIs convey both variability and precision of the estimated mean; for small n they can aid readers in directly assessing the uncertainty around the reported central estimate
  • CNV inference with infercnvpy was used to distinguish tumor- from normal-originated pericytes, validated by a second deep-learning classifier (Cancer-finder)
    Could also: CopyKAT or Numbat (Bayesian/HMM frameworks) could also be applied for CNV calling from scRNA-seq — Different CNV inference tools use distinct statistical models and can produce complementary confidence metrics; benchmarking across two or more CNV-based methods alongside the expression classifier can further substantiate T-PC/N-PC assignments
  • Batch correction was performed with Harmony prior to unsupervised clustering of the 5-patient scRNA-seq dataset
    Could also: Probabilistic integration methods such as scVI (variational autoencoder) or Scanorama could also be used for multi-sample integration — Different integration strategies make different assumptions about shared versus patient-specific variation; comparing cluster stability across methods is a common way to assess robustness of identified pericyte subpopulations across patients
  • Pericyte subpopulation identity was based on marker gene expression and CNV scoring within the scRNA-seq clustering
    Could also: Somatic variant calling from scRNA-seq reads (e.g., Monopogen or SCReadCounts) could also be used as a mutation-based lineage assignment strategy — Mutation-based evidence is orthogonal to both CNV amplitude and expression-classifier signals, providing an independent molecular anchor for confirming tumor origin, particularly for cells with low-amplitude or ambiguous CNV profiles
Software: 10x Genomics Chromium · Seurat · Harmony · infercnvpy · Cancer-finder

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 63/100
partly built on non-reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41001759

Title: Immunoregulatory Roles of Tumor-Originated Pericytes Identified by Single-Cell Analysis in Glioblastoma. (Chu C, … Bian XW, Wu HB, Zhang A, Zhou W.) Adv Sci (Weinh) 2025 · PMID 41001759 · PMCID PMC12713092 · DOI 10.1002/advs.202511856 Group: Army Medical University (Third Military Medical University), Chongqing — not the MILO lab.

Verbatim Data & Code Availability (from PMC full text)

"scRNA-seq datasets have been deposited at the Genome Sequence Archive at the National Genomics Data Center (Beijing, China) that are accessible under BioProject PRJCA025812. UMI RNA-seq datasets have been deposited … under BioProject PRJCA035005. This paper does not report original code."

What the reproducible (pipeline-derived) results are, and where the data lives

ALL headline computational results derive from the authors' in-house CD146+ pericyte scRNA-seq of 5 GBM patients, integrated with 8 public datasets into a "GBM atlas". Core reported pipeline outputs:

  • 35,677 high-quality CD146+ transcriptomes (≈5–10k per patient); 17 CD146+ clusters.
  • 33,572 cells retained as true pericytes after removing C4/C16 (endothelial), C17 (macrophage).
  • T-PC vs N-PC ≈ 50/50 overall; T-PC → 10 clusters → 5 metaclusters (T-M1..T-M5); N-PC → 9 clusters → 3 metaclusters (N-M1..N-M3).
  • CellChat: 4 communication modes; T-M2→macrophage top signals CSF1/SPP1/CCL2; N-M3→microglia top signals APOE/CCL3/TGFB1. CNV: chr7 amp / chr10 loss.
  • Tools named (NO versions/params): Seurat, Harmony, infercnvpy, pySCENIC, aPEAR, clusterProfiler, scFEA, PyCoGAPS, CellChat, NMF, NicheNet, Tangram, Scanpy, COMMOT, mistyR, ROGUE, Cancer-Finder.

In scope vs out of scope

  • In scope (pipeline-derived): the clustering / pericyte-calling / sub-clustering / CellChat / CNV results above. All require the in-house data (PRJCA025812 / HRA007338).
  • Out of scope: wet-lab (IHC, THP-1 M2 polarization, conditioned-medium assays), TCGA survival stratification (external), protein-level validation.

Blocking screening findings

  1. Primary data is controlled-access. PRJCA025812 → HRA007338 in GSA-Human (NGDC). API: isControlledAccess:1, isFilePublic:0, DAC id 4207 — metadata public, files require a Data Access Request/Agreement to the submitting DAC. Cannot be downloaded autonomously. UMI RNA-seq PRJCA035005 not file-public either.
  2. No original code. Paper states verbatim it "does not report original code." The harvested code link github.com/theMILOlab/SPATAData is not the paper's analysis pipeline — it is the public data resource from which the authors downloaded published spatial transcriptomics datasets. (SPATAData = Freiburg MILO lab R data-access package; unrelated to this Chongqing group's analysis.) This is a code-link false positive for "analysis code".
  3. No public-data-only claim. The 8 public reference datasets (GSE141552, GSE138794, GSE162631, GSE103224, GSE141383, GSE182109, GSE163108, GSE173278) are obtainable, but every reported number is an atlas integration of in-house + public data — no specific reported value is attributable to a public dataset alone, so re-running a public set maps to no expected result to compare against.

Decision

DROPdrop_reason = data_restricted (controlled-access in-house data, files not public, request-to-DAC), compounded by no_code (no original code). No clearly-specified pipeline output is reproducible from publicly obtainable inputs. No «our HPC» compute submitted (nothing to reproduce). Not fabricating a public-data result to avoid the drop, per brief rule 6.

Figures / tables: Fig 1Fig 4Fig 5
C1
Reported
35677 high-quality CD146+ transcriptomes (5 patients)
Reproduced
partial
C2
Reported
17 CD146+ clusters
Reproduced
partial
C3
Reported
33572 true pericytes after removing C4/C16/C17
Reproduced
partial
C4
Reported
T-PC: 10 clusters -> 5 metaclusters (T-M1..T-M5)
Reproduced
partial
C5
Reported
N-PC: 9 clusters -> 3 metaclusters (N-M1..N-M3)
Reproduced
partial
C6
Reported
CellChat: 4 communication modes
Reproduced
partial
C7
Reported
T-M2->macrophage top signals CSF1, SPP1, CCL2
Reproduced
partial
C8
Reported
N-M3->microglia top signals APOE, CCL3, TGFB1
Reproduced
partial
C9
Reported
CNV: chr7 amplification, chr10 loss
Reproduced
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 44/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is a legitimate access drop, not a quality or integrity problem. Every reported value (35677 cells, 17 clusters, 33572 true pericytes, T-PC/N-PC metaclusters, CellChat CSF1/SPP1/CCL2, CNV chr7 gain/chr10 loss) depends on in-house CD146+ scRNA-seq deposited as controlled-access (HRA007338, isFilePublic:0), and the paper explicitly reports no original code — the harvested theMILOlab/SPATAData link is a false-positive data resource. Nothing could therefore be derived, compared, or confirmed; the blocker is on the data-availability side, not our methodology nor an authors' defect. The reported numbers are internally consistent with no fabrication signal, so q5/q7/q8 are yellow (undetermined/limited) rather than red.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

66.6 k
tokens (I/O) · 2.7 M incl. cache
6 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.