Assessing personalized molecular portraits underlying endothelial-to-mesenchymal transition within pulmonary arterial hypertension.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Strong partial reproduction via a standard Seurat v4 pipeline (P16) on the paper's own public GSE169471 10x H5 matrices. CORE scRNA results reproduce cleanly: C1 31,439 vs 31,444 (within-tol); C2 28,906 vs 28,906 (EXACT); C3 21 vs 21 clusters @res0.5/20PC (EXACT); C4 9 vs 9 cell types, identical lineages (EXACT); C5 3,304 vs 3,165 EC cells (within-tol, +4.4%). The unstated 9-of-11 sample selection (all 3 IPAH + 6 controls, dropping the two lobes SC155/SC156 of one donor) was recovered deterministically and reproduces both reported totals, supporting the paper's counts. Hardest specified scRNA claim H1 partially reproduces: EC re-clusters to exactly 3 subclusters (EC1/EC2/EC3) and one carries 689 significant markers vs reported EC3=602 (+14.4%). The deep downstream chain (hdWGCNA modules -> 6 ETPGs -> caret PETS score, H2/H3/H5) and the separate bulk-RNA-seq DEGs (H4) were not attempted: under-specified/stochastic and partly on data not provided. No fabrication indicators. All heavy compute on «our HPC» SLURM.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 81assessed: 2026-06-20 ⛓ 62c5fc5041ca
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether integrating single-cell and bulk RNA-seq data can reveal a pathogenic endothelial-to-mesenchymal transition (EndMT) process and specific cell populations driving pulmonary arterial hypertension (PAH), and whether EndMT pattern genes derived from this process can be used to build a machine-learning-based diagnostic molecular signature (PETS) for PAH.
- ★ scRNA-seq of PAH and control lung tissue identifies nine distinct cell populations with high heterogeneity in composition, function, distribution, and communication finding
- ★ Endothelial cells show the most prominent variation among cell types across multiple analytical perspectives in PAH finding
- ★ Endothelial cells undergo endothelial-to-mesenchymal transition (EndMT) in PAH, with a distinct EC subgroup (EC3) exhibiting a contrasting mesenchymal-like phenotype finding
- ★ EndMT pattern genes (ETPGs) were derived from pivotal hdWGCNA module genes overlapping with EC3 marker genes method
- ★ A nine-algorithm machine-learning program built on ETPGs produces the PAH Endothelial-mesenchymal Transition Signature (PETS), with glmNet as the optimal model for discriminating PAH from healthy individuals resource
- In PAH, ECs shift from signaling receiver to signaling sender within the cell-cell communication network finding
- ECs, SMCs, and fibroblasts increase in proportion in PAH lung tissue relative to control finding
- Macrophages/monocytes and ECs contribute most to PAH-associated transcriptomic differences among cell types finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| scRNA-seq (10X Genomics) | human lung tissue (3 PAH, 6 control samples, GSE169471) | none (disease state: PAH vs control) | cell type composition, proportions, tissue preference (Ro/e), contribution score, pathway activity (AUCell) | 10X Genomics |
| bulk RNA-seq / microarray differential expression and GSEA | human PAH lung/blood cohort (GSE117261) | none (disease state: PAH vs control) | differentially expressed genes, enriched pathways, ML signature performance | — |
| cell-cell communication analysis (CellChat) | scRNA-seq-derived cell populations, human lung tissue | none (disease state: PAH vs control) | ligand-receptor interaction number, strength, signaling pathway activity, in/out-degree | CellChat/CellChatDB |
| pseudotime trajectory inference (Slingshot, Monocle2) | 3165 endothelial cells, human lung tissue (PAH and control) | none (disease state) | EC differentiation trajectory, EndMT progression, DEGs along trajectory (qval<0.1) | — |
| high-dimensional weighted gene co-expression network analysis (hdWGCNA) | endothelial cell metacells, human lung tissue | none (disease state) | module eigengenes, pivotal module gene identification | hdWGCNA R package |
| machine learning modeling (9 algorithms: glmNet, Bagged CART, NB, pls, NNet, KNN, RF, glmBoost, CART) | human PAH cohort (GSE117261), discovery:testing = 7:3 | none (disease state classification) | accuracy, C-index, F1-score, precision, recall, RMSR | — |
| RT-qPCR | primary human and mouse pulmonary artery endothelial cells (PAECs) | hypoxia (1% O2) | relative mRNA expression of Rarres2/RARRES2, Tc2n/TC2N, CBY1, CKS1B, MSRB3, SMAGP | SYBR Green Mix (Takara) |
| immunofluorescence staining | mouse hypoxic PAH model lung tissue | hypoxia (10% O2, 4 weeks) | CD31 and chemerin protein expression/localization | — |
- – 28,906 cells passed QC and were classified via UMAP into 21 clusters annotated as 9 cell types 28906 cells; 21 clusters; 9 cell types
- ▲ ECs, SMCs, and fibroblasts increased in proportion in the PAH group
- – Macrophages/monocytes and ECs showed the greatest contribution scores to PAH-driven differences; NK cells, T cells, mast cells showed limited impact
- ▲ Endothelial cells were most enriched in oxidative phosphorylation and MYC target V1 pathways among stromal cells
- ▲ Intercellular interaction number and strength were elevated in the PAH group, with enhanced EC/fibroblast interactions with immune cells
- – ECs exhibited lower incoming and higher outgoing interactions in PAH, indicating a shift from signal receiver to sender
- – Endothelial cells resolved into EC1, EC2, EC3 subpopulations; pseudotime trajectory placed EC1 at the start and EC3 at the endpoint, with EC3 showing marked heterogeneity 602 marker genes for EC3
- – The glmNet model was identified as the optimal machine learning scheme for the PETS signature
- count 31,444 cells (3 PAH + 6 control samples) (initial scRNA-seq dataset from GSE169471)
- count 28,906 cells retained after QC (post-quality-control cell count)
- count 21 cell clusters annotated into 9 cell types (UMAP clustering resolution = 0.5)
- count 602 significant marker genes (EC3 subset marker genes vs EC1/EC2)
- count 3165 endothelial cells (input for Slingshot pseudotime trajectory analysis)
- fold_change average log2 Fold Change > 1 (threshold for EC3 significant altered genes used to define ETPGs)
- pvalue adjusted P value < 0.05 (significance threshold for EC3 marker gene selection)
- other 7:3 split ratio (GSE117261 cohort divided into discovery and testing sets)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper combines single-cell RNA-seq and bulk RNA-seq analyses of PAH versus control samples with a battery of bioinformatic pipelines (Seurat, limma, CellChat, Slingshot/Monocle2, hdWGCNA, AUCell, scMetabolism) to characterize cell populations and an EndMT gene signature, followed by qPCR and a mouse hypoxia model for validation. Group comparisons relied primarily on a T test for continuous variables, a Kruskal-Wallis test across three endothelial subsets, a chi-square test for tissue-preference (Ro/e) analysis, and empirical Bayes statistics for pathway-activity scores, with limma used for bulk differential expression. A nine-algorithm machine-learning pipeline with 10-fold, 10-repeat cross-validation on a 7:3 split cohort was used to build and internally evaluate the PETS signature. Significance was defined globally as two-sided P<0.05 combined with FDR<0.05.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Student's t-test (referred to as "T test") | comparison of continuous variables (e.g., experimental/validation measurements) | — | not stated |
| Kruskal-Wallis test | oxidative phosphorylation activity levels compared among EC1, EC2, EC3 endothelial subsets | — | not stated |
| Chi-square test | predicted vs. observed cell numbers per cell type/group for Ro/e tissue-preference analysis | — | not stated |
| limma differential expression analysis | bulk RNA-seq DEG identification between PAH and control | — | not stated |
| empirical Bayes statistics | comparison of AUCell pathway activity scores between PAH and control groups | — | not stated |
| Gene set enrichment analysis (GSEA) | pathway-level analysis of ranked bulk RNA-seq expression differences | — | na |
-
Continuous variables were compared using a T test without stating whether normality was checked or whether the test was paired or unpaired.↳ Could also: A non-parametric alternative such as the Mann-Whitney U test (unpaired) or Wilcoxon signed-rank test (paired), or an explicit normality check (e.g., Shapiro-Wilk) before selecting a parametric test — would also yield valid inference without relying on the normality assumption, which can matter for the smaller sample sizes typical of qPCR or animal-model experiments
-
The Kruskal-Wallis test was used to compare oxidative phosphorylation levels across EC1, EC2, and EC3, giving an overall (omnibus) result.↳ Could also: A post-hoc pairwise test such as Dunn's test with a multiple-comparison correction (e.g., Benjamini-Hochberg or Bonferroni) — would also indicate which specific pairs of endothelial subsets differ, complementing the omnibus significance result
-
Single-cell differential expression (e.g., EC3 marker genes via FindAllMarkers) was derived from thousands of cells drawn from only 3 PAH and 6 control biological samples.↳ Could also: A pseudobulk approach (aggregating counts per sample before testing, e.g. with DESeq2/edgeR) or a mixed-effects model with sample as a random effect — would also explicitly account for correlation among cells from the same biological sample, an aspect that per-cell-level tests otherwise treat as fully independent observations
-
A combined two-sided P<0.05 and FDR<0.05 threshold is described as applying to "all statistical tests" without naming the correction algorithm or the comparison family for each analysis.↳ Could also: Explicitly naming the FDR procedure (e.g., Benjamini-Hochberg) and stating the comparison family for each analysis (e.g., per DEG list, per pathway set, per marker-gene test) — would also make the multiplicity-control approach fully reproducible and let readers confirm exactly which comparisons were adjusted together
-
The PETS signature was developed and evaluated using a 7:3 split of a single cohort (GSE117261) with 10-fold, 10-repeat cross-validation.↳ Could also: Validation in a fully independent external cohort not used at all during model training or tuning — would also provide an additional check on generalizability beyond resampling within one dataset
-
Effect magnitude for expression differences was conveyed via log2 fold-change thresholds combined with adjusted P-value cutoffs, without a standardized effect size for the T-test/Kruskal-Wallis comparisons.↳ Could also: Reporting a standardized effect size (e.g., Cohen's d for the T-test, or a rank-based measure such as epsilon-squared for Kruskal-Wallis) alongside the p-value — would also convey the magnitude of the observed difference independent of sample size, complementing the significance threshold
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-39462326
Paper: Wu et al. 2024, Mol Med. "Assessing personalized molecular portraits underlying endothelial-to-mesenchymal transition within pulmonary arterial hypertension." PMID 39462326 · PMCID PMC11513636 · DOI 10.1186/s10020-024-00963-z
Listed code: https://github.com/zhanghao-njmu/SCP (SCP = a generic third-party single-cell pipeline R package; per P16 applying it / equivalent Seurat steps to the paper's data is equally valid). The Methods actually describe a standard Seurat v4 workflow (LogNormalize, vst 2000 HVG, IntegrateData, 20 PCs, FindClusters res=0.5, UMAP) plus downstream tools (Slingshot, Monocle2, AUCell, scMetabolism, CellChat, hdWGCNA, caret ML).
Primary data: GEO GSE169471 — droplet 10x scRNA-seq of human lung.
GEO holds 11 samples (3 IPAH + 8 control-ish) as per-GSM .h5 count
matrices inside GSE169471_RAW.tar (310 MB). The paper used 9 (3 PAH + 6
control). Raw FASTQ for controls is access-restricted, but the H5 count
matrices are public → reproduction uses those.
In scope (pipeline-derived, tractable — 80/20 priority)
| ID | Reported claim | Paper location | Pipeline | Tractability |
|---|---|---|---|---|
| C1 | 31,444 cells from 3 PAH + 6 control samples analyzed | Results/Methods | load 9 H5 matrices, count cells | HIGH — deterministic |
| C2 | 28,906 cells retained after QC | Results | QC: >200 & <4000 genes/cell, >3 cells/gene, <10% MT | HIGH — deterministic given thresholds |
| C3 | 21 clusters (FindClusters resolution=0.5, top 20 PCs) | Methods/Results | Seurat integrate→PCA→Louvain | MED — Seurat-version sensitive |
| C4 | 9 cell types (B, EC, epithelial, fibroblast, macro/mono, mast, NK, SMC, T) | Results | marker-based annotation | MED — annotation is judgement |
| C5 | 3165 endothelial cells subset for trajectory | Results | subset EC cluster | MED — depends on C4 |
Hard last ~20% (attempt only if cheap; else documented as not-attempted)
| ID | Reported claim | Why hard |
|---|---|---|
| H1 | 602 significant EC3 marker genes | depends on EC sub-clustering into EC1-3 (under-specified resolution) |
| H2 | 6 ETPGs: RARRES2, CBY1, MSRB3, TC2N, CKS1B, SMAGP (66 greenyellow ∩ 602) | needs hdWGCNA β=24 → 13 modules + EC3 markers; long chain |
| H3 | hdWGCNA: soft β=24, 13 modules, greenyellow module | hdWGCNA params under-specified, stochastic |
| H4 | Bulk RNA-seq: 38 up + 30 down DEGs | uses a SEPARATE bulk dataset (not named in scope brief); out of the GSE169471 pipeline |
| H5 | PETS = 2.6926723·RARRES2 + 0.7410611·TC2N (caret ML, 9 learners) | downstream of H2-H4; ML CV stochastic |
Out of scope (not pipeline / not deposited)
- Wet-lab validation, IHC, clinical PAH cohort phenotypes.
- Bulk RNA-seq DEGs (H4): the bulk accession is not given in the brief; the EndMT signature validation is a separate dataset chain — noted, not attempted unless the accession surfaces cheaply.
Plan
- «our HPC» job: download
GSE169471_RAW.tarto «infra», extract H5s, report per-sample cell counts (raw) → resolve which 9 samples = 31,444 (C1). - QC filter per paper thresholds → C2.
- Seurat integrate (CCA), 20 PCs, FindClusters res=0.5 → C3; marker annotation → C4.
- Subset EC → C5. Stop at the 80% line; document H1-H5 as not-attempted/hard.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The core scRNA-seq pipeline reproduces strongly on the paper's own public GSE169471 H5 matrices — C2 (28,906), C3 (21 clusters), C4 (9 lineages) are exact, and C1/C5 are within-tol; H1's EC3 marker count is close (689 vs 602, +14.4%). All reproduced values are derivable from shared data with no fabrication indicators. The main caveats are on our/authors' methodology and completeness: the paper never states its 9-of-11 sample selection (recovered deterministically by us), and the central novel claims — 6 ETPGs, hdWGCNA module, and the PETS ML formula — were not attempted because the downstream chain is under-specified/stochastic and partly on data not provided. Net: a solid partial reproduction with explainable deviations, not a substantive discrepancy.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.