Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Assessing personalized molecular portraits underlying endothelial-to-mesenchymal transition within pulmonary arterial hypertension.

Mol Med · 2024
L1 81/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
81/100
Reproducibility score
0.4 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 59% of all assessed papers rank 468 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Strong partial reproduction via a standard Seurat v4 pipeline (P16) on the paper's own public GSE169471 10x H5 matrices. CORE scRNA results reproduce cleanly: C1 31,439 vs 31,444 (within-tol); C2 28,906 vs 28,906 (EXACT); C3 21 vs 21 clusters @res0.5/20PC (EXACT); C4 9 vs 9 cell types, identical lineages (EXACT); C5 3,304 vs 3,165 EC cells (within-tol, +4.4%). The unstated 9-of-11 sample selection (all 3 IPAH + 6 controls, dropping the two lobes SC155/SC156 of one donor) was recovered deterministically and reproduces both reported totals, supporting the paper's counts. Hardest specified scRNA claim H1 partially reproduces: EC re-clusters to exactly 3 subclusters (EC1/EC2/EC3) and one carries 689 significant markers vs reported EC3=602 (+14.4%). The deep downstream chain (hdWGCNA modules -> 6 ETPGs -> caret PETS score, H2/H3/H5) and the separate bulk-RNA-seq DEGs (H4) were not attempted: under-specified/stochastic and partly on data not provided. No fabrication indicators. All heavy compute on «our HPC» SLURM.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 81
    assessed: 2026-06-20 ⛓ 62c5fc5041ca
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether integrating single-cell and bulk RNA-seq data can reveal a pathogenic endothelial-to-mesenchymal transition (EndMT) process and specific cell populations driving pulmonary arterial hypertension (PAH), and whether EndMT pattern genes derived from this process can be used to build a machine-learning-based diagnostic molecular signature (PETS) for PAH.

Core claims
  • scRNA-seq of PAH and control lung tissue identifies nine distinct cell populations with high heterogeneity in composition, function, distribution, and communication finding
  • Endothelial cells show the most prominent variation among cell types across multiple analytical perspectives in PAH finding
  • Endothelial cells undergo endothelial-to-mesenchymal transition (EndMT) in PAH, with a distinct EC subgroup (EC3) exhibiting a contrasting mesenchymal-like phenotype finding
  • EndMT pattern genes (ETPGs) were derived from pivotal hdWGCNA module genes overlapping with EC3 marker genes method
  • A nine-algorithm machine-learning program built on ETPGs produces the PAH Endothelial-mesenchymal Transition Signature (PETS), with glmNet as the optimal model for discriminating PAH from healthy individuals resource
  • In PAH, ECs shift from signaling receiver to signaling sender within the cell-cell communication network finding
  • ECs, SMCs, and fibroblasts increase in proportion in PAH lung tissue relative to control finding
  • Macrophages/monocytes and ECs contribute most to PAH-associated transcriptomic differences among cell types finding
Experimental setups
Assay System Perturbation Readout Platform
scRNA-seq (10X Genomics) human lung tissue (3 PAH, 6 control samples, GSE169471) none (disease state: PAH vs control) cell type composition, proportions, tissue preference (Ro/e), contribution score, pathway activity (AUCell) 10X Genomics
bulk RNA-seq / microarray differential expression and GSEA human PAH lung/blood cohort (GSE117261) none (disease state: PAH vs control) differentially expressed genes, enriched pathways, ML signature performance
cell-cell communication analysis (CellChat) scRNA-seq-derived cell populations, human lung tissue none (disease state: PAH vs control) ligand-receptor interaction number, strength, signaling pathway activity, in/out-degree CellChat/CellChatDB
pseudotime trajectory inference (Slingshot, Monocle2) 3165 endothelial cells, human lung tissue (PAH and control) none (disease state) EC differentiation trajectory, EndMT progression, DEGs along trajectory (qval<0.1)
high-dimensional weighted gene co-expression network analysis (hdWGCNA) endothelial cell metacells, human lung tissue none (disease state) module eigengenes, pivotal module gene identification hdWGCNA R package
machine learning modeling (9 algorithms: glmNet, Bagged CART, NB, pls, NNet, KNN, RF, glmBoost, CART) human PAH cohort (GSE117261), discovery:testing = 7:3 none (disease state classification) accuracy, C-index, F1-score, precision, recall, RMSR
RT-qPCR primary human and mouse pulmonary artery endothelial cells (PAECs) hypoxia (1% O2) relative mRNA expression of Rarres2/RARRES2, Tc2n/TC2N, CBY1, CKS1B, MSRB3, SMAGP SYBR Green Mix (Takara)
immunofluorescence staining mouse hypoxic PAH model lung tissue hypoxia (10% O2, 4 weeks) CD31 and chemerin protein expression/localization
Key results
  • 28,906 cells passed QC and were classified via UMAP into 21 clusters annotated as 9 cell types 28906 cells; 21 clusters; 9 cell types
  • ECs, SMCs, and fibroblasts increased in proportion in the PAH group
  • Macrophages/monocytes and ECs showed the greatest contribution scores to PAH-driven differences; NK cells, T cells, mast cells showed limited impact
  • Endothelial cells were most enriched in oxidative phosphorylation and MYC target V1 pathways among stromal cells
  • Intercellular interaction number and strength were elevated in the PAH group, with enhanced EC/fibroblast interactions with immune cells
  • ECs exhibited lower incoming and higher outgoing interactions in PAH, indicating a shift from signal receiver to sender
  • Endothelial cells resolved into EC1, EC2, EC3 subpopulations; pseudotime trajectory placed EC1 at the start and EC3 at the endpoint, with EC3 showing marked heterogeneity 602 marker genes for EC3
  • The glmNet model was identified as the optimal machine learning scheme for the PETS signature
Key statistics
  • count 31,444 cells (3 PAH + 6 control samples) (initial scRNA-seq dataset from GSE169471)
  • count 28,906 cells retained after QC (post-quality-control cell count)
  • count 21 cell clusters annotated into 9 cell types (UMAP clustering resolution = 0.5)
  • count 602 significant marker genes (EC3 subset marker genes vs EC1/EC2)
  • count 3165 endothelial cells (input for Slingshot pseudotime trajectory analysis)
  • fold_change average log2 Fold Change > 1 (threshold for EC3 significant altered genes used to define ETPGs)
  • pvalue adjusted P value < 0.05 (significance threshold for EC3 marker gene selection)
  • other 7:3 split ratio (GSE117261 cohort divided into discovery and testing sets)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper combines single-cell RNA-seq and bulk RNA-seq analyses of PAH versus control samples with a battery of bioinformatic pipelines (Seurat, limma, CellChat, Slingshot/Monocle2, hdWGCNA, AUCell, scMetabolism) to characterize cell populations and an EndMT gene signature, followed by qPCR and a mouse hypoxia model for validation. Group comparisons relied primarily on a T test for continuous variables, a Kruskal-Wallis test across three endothelial subsets, a chi-square test for tissue-preference (Ro/e) analysis, and empirical Bayes statistics for pathway-activity scores, with limma used for bulk differential expression. A nine-algorithm machine-learning pipeline with 10-fold, 10-repeat cross-validation on a 7:3 split cohort was used to build and internally evaluate the PETS signature. Significance was defined globally as two-sided P<0.05 combined with FDR<0.05.

Replicationmixed Sample sizeSample sizes are given for the scRNA-seq cohort (3 PAH, 6 control donor samples, 31,444 total cells) and for the machine-learning cohort split (GSE117261 divided 7:3 into discovery/testing), but no a priori power or sample-size calculation is described; n for the qPCR and mouse hypoxia-model experiments is not stated in the provided text. GroupsPAH vs. control (scRNA-seq, bulk RNA-seq, cell-cell communication); EC1 vs. EC2 vs. EC3 endothelial subsets; hypoxia vs. normoxia in vitro/mouse model Pairingunclear Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesyes Multiplicity correctionFDR (false discovery rate); the specific correction algorithm, e.g. Benjamini-Hochberg, is not named
Statistical tests used
Test Applied to n Assumptions
Student's t-test (referred to as "T test") comparison of continuous variables (e.g., experimental/validation measurements) not stated
Kruskal-Wallis test oxidative phosphorylation activity levels compared among EC1, EC2, EC3 endothelial subsets not stated
Chi-square test predicted vs. observed cell numbers per cell type/group for Ro/e tissue-preference analysis not stated
limma differential expression analysis bulk RNA-seq DEG identification between PAH and control not stated
empirical Bayes statistics comparison of AUCell pathway activity scores between PAH and control groups not stated
Gene set enrichment analysis (GSEA) pathway-level analysis of ranked bulk RNA-seq expression differences na
Approaches that could also have been used
  • Continuous variables were compared using a T test without stating whether normality was checked or whether the test was paired or unpaired.
    Could also: A non-parametric alternative such as the Mann-Whitney U test (unpaired) or Wilcoxon signed-rank test (paired), or an explicit normality check (e.g., Shapiro-Wilk) before selecting a parametric test — would also yield valid inference without relying on the normality assumption, which can matter for the smaller sample sizes typical of qPCR or animal-model experiments
  • The Kruskal-Wallis test was used to compare oxidative phosphorylation levels across EC1, EC2, and EC3, giving an overall (omnibus) result.
    Could also: A post-hoc pairwise test such as Dunn's test with a multiple-comparison correction (e.g., Benjamini-Hochberg or Bonferroni) — would also indicate which specific pairs of endothelial subsets differ, complementing the omnibus significance result
  • Single-cell differential expression (e.g., EC3 marker genes via FindAllMarkers) was derived from thousands of cells drawn from only 3 PAH and 6 control biological samples.
    Could also: A pseudobulk approach (aggregating counts per sample before testing, e.g. with DESeq2/edgeR) or a mixed-effects model with sample as a random effect — would also explicitly account for correlation among cells from the same biological sample, an aspect that per-cell-level tests otherwise treat as fully independent observations
  • A combined two-sided P<0.05 and FDR<0.05 threshold is described as applying to "all statistical tests" without naming the correction algorithm or the comparison family for each analysis.
    Could also: Explicitly naming the FDR procedure (e.g., Benjamini-Hochberg) and stating the comparison family for each analysis (e.g., per DEG list, per pathway set, per marker-gene test) — would also make the multiplicity-control approach fully reproducible and let readers confirm exactly which comparisons were adjusted together
  • The PETS signature was developed and evaluated using a 7:3 split of a single cohort (GSE117261) with 10-fold, 10-repeat cross-validation.
    Could also: Validation in a fully independent external cohort not used at all during model training or tuning — would also provide an additional check on generalizability beyond resampling within one dataset
  • Effect magnitude for expression differences was conveyed via log2 fold-change thresholds combined with adjusted P-value cutoffs, without a standardized effect size for the T-test/Kruskal-Wallis comparisons.
    Could also: Reporting a standardized effect size (e.g., Cohen's d for the T-test, or a rank-based measure such as epsilon-squared for Kruskal-Wallis) alongside the p-value — would also convey the magnitude of the observed difference independent of sample size, complementing the significance threshold
Software: R 4.1.2 · Seurat · limma · clusterProfiler · CellChat · Slingshot · Monocle2 · hdWGCNA · scMetabolism · AUCell · SCP package

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39462326

Paper: Wu et al. 2024, Mol Med. "Assessing personalized molecular portraits underlying endothelial-to-mesenchymal transition within pulmonary arterial hypertension." PMID 39462326 · PMCID PMC11513636 · DOI 10.1186/s10020-024-00963-z

Listed code: https://github.com/zhanghao-njmu/SCP (SCP = a generic third-party single-cell pipeline R package; per P16 applying it / equivalent Seurat steps to the paper's data is equally valid). The Methods actually describe a standard Seurat v4 workflow (LogNormalize, vst 2000 HVG, IntegrateData, 20 PCs, FindClusters res=0.5, UMAP) plus downstream tools (Slingshot, Monocle2, AUCell, scMetabolism, CellChat, hdWGCNA, caret ML).

Primary data: GEO GSE169471 — droplet 10x scRNA-seq of human lung. GEO holds 11 samples (3 IPAH + 8 control-ish) as per-GSM .h5 count matrices inside GSE169471_RAW.tar (310 MB). The paper used 9 (3 PAH + 6 control). Raw FASTQ for controls is access-restricted, but the H5 count matrices are public → reproduction uses those.

In scope (pipeline-derived, tractable — 80/20 priority)

ID Reported claim Paper location Pipeline Tractability
C1 31,444 cells from 3 PAH + 6 control samples analyzed Results/Methods load 9 H5 matrices, count cells HIGH — deterministic
C2 28,906 cells retained after QC Results QC: >200 & <4000 genes/cell, >3 cells/gene, <10% MT HIGH — deterministic given thresholds
C3 21 clusters (FindClusters resolution=0.5, top 20 PCs) Methods/Results Seurat integrate→PCA→Louvain MED — Seurat-version sensitive
C4 9 cell types (B, EC, epithelial, fibroblast, macro/mono, mast, NK, SMC, T) Results marker-based annotation MED — annotation is judgement
C5 3165 endothelial cells subset for trajectory Results subset EC cluster MED — depends on C4

Hard last ~20% (attempt only if cheap; else documented as not-attempted)

ID Reported claim Why hard
H1 602 significant EC3 marker genes depends on EC sub-clustering into EC1-3 (under-specified resolution)
H2 6 ETPGs: RARRES2, CBY1, MSRB3, TC2N, CKS1B, SMAGP (66 greenyellow ∩ 602) needs hdWGCNA β=24 → 13 modules + EC3 markers; long chain
H3 hdWGCNA: soft β=24, 13 modules, greenyellow module hdWGCNA params under-specified, stochastic
H4 Bulk RNA-seq: 38 up + 30 down DEGs uses a SEPARATE bulk dataset (not named in scope brief); out of the GSE169471 pipeline
H5 PETS = 2.6926723·RARRES2 + 0.7410611·TC2N (caret ML, 9 learners) downstream of H2-H4; ML CV stochastic

Out of scope (not pipeline / not deposited)

  • Wet-lab validation, IHC, clinical PAH cohort phenotypes.
  • Bulk RNA-seq DEGs (H4): the bulk accession is not given in the brief; the EndMT signature validation is a separate dataset chain — noted, not attempted unless the accession surfaces cheaply.

Plan

  1. «our HPC» job: download GSE169471_RAW.tar to «infra», extract H5s, report per-sample cell counts (raw) → resolve which 9 samples = 31,444 (C1).
  2. QC filter per paper thresholds → C2.
  3. Seurat integrate (CCA), 20 PCs, FindClusters res=0.5 → C3; marker annotation → C4.
  4. Subset EC → C5. Stop at the 80% line; document H1-H5 as not-attempted/hard.
C1
Reported
31,444 cells (3 PAH + 6 control) loaded
Reproduced
31,439
within tolerance
C2
Reported
28,906 cells after QC
Reproduced
28,906
exact
C3
Reported
21 clusters @ res=0.5, 20 PCs
Reproduced
21
exact
C4
Reported
9 cell types (B, EC, epithelial, fibroblast, macro/mono, mast, NK, SMC, T)
Reproduced
9 (Bcell, EC, Epithelial, Fibroblast, MacroMono, Mast, NK, SMC, Tcell)
exact
C5
Reported
3,165 EC cells subset for trajectory
Reproduced
3,304 (clusters 5/9/18)
within tolerance
H1
Reported
EC3 = 602 significant marker genes
Reproduced
3 EC subclusters @res0.02-0.04 (2978/236/90); sig markers {181,689,143}; closest to EC3 = 689 (+14.4%)
partial
H2-H5
Reported
6 ETPGs / hdWGCNA beta=24 13 modules / bulk 38up+30down / PETS ML formula
Reproduced
NOT_ATTEMPTED
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 81/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

The core scRNA-seq pipeline reproduces strongly on the paper's own public GSE169471 H5 matrices — C2 (28,906), C3 (21 clusters), C4 (9 lineages) are exact, and C1/C5 are within-tol; H1's EC3 marker count is close (689 vs 602, +14.4%). All reproduced values are derivable from shared data with no fabrication indicators. The main caveats are on our/authors' methodology and completeness: the paper never states its 9-of-11 sample selection (recovered deterministically by us), and the central novel claims — 6 ETPGs, hdWGCNA module, and the PETS ML formula — were not attempted because the downstream chain is under-specified/stochastic and partly on data not provided. Net: a solid partial reproduction with explainable deviations, not a substantive discrepancy.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

512.2 k
tokens (I/O) · 37 M incl. cache
177 min
runtime · 0.27 CPU-h
17.1 GB
peak RAM
4
HPC jobs
hummel
machine