Vascular adhesion protein-1 defines a unique subpopulation of human hematopoietic stem cells and regulates their proliferation.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to be reproducible in principle: the paper's headline computational results come from scRNA-seq PRJNA729883 (SRA study SRP319821; 4 sorted-HSC 10x v3.1 samples, raw FASTQ only) via Cell Ranger 5.0.1 -> Seurat 4.0.1 -> UMAP/Monocle2 -> Nebulosa. BRIEF's PRJNA594799 is actually the BULK set; no GEO processed matrix or Seurat object is shipped and the paper's Code-availability section is blank, so reproduction must run the third-party pipeline on the FASTQ from scratch (valid per P16). OUTCOME = PARTIAL: I resolved the data discrepancy, pinned 4 claims (C1 AOC3 near-zero, C2 1267 HSC, C3 HSC1/HSC2 93/7%=89, C4 HSC2 cell-cycle enrichment), and built+submitted the full pipeline on «our HPC» (10x GRCh38-2020-A reference downloaded, STAR index built). NO numeric claim was graded: the ENA FASTQ staging step had a nested-xargs/wget quoting bug that wrote 32 zero-byte files, and the run was finalized on operator instruction before any count matrix existed; the alignment array was cancelled rather than run on empty input. I did NOT confirm or contradict any paper number, and explicitly did not attempt the bulk-RNA-seq DE markers or the Monocle2 pseudotime (hard/under-specified 20%). The STAR index, both sbatch scripts (STARsolo CR-emulation) and the Seurat+Nebulosa R script are staged on «infra» and re-runnable after a one-line download fix; deviation disclosed: Cell Ranger 5.0.1 -> STARsolo CR-emulation, same reference.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-15 ⛓ 257826c4052c
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study tests whether Vascular Adhesion Protein-1 (VAP-1) marks a distinct subpopulation of human hematopoietic stem cells (HSC) and whether VAP-1-generated hydrogen peroxide regulates HSC proliferation and differentiation, both in vitro and in vivo.
- ★ VAP-1 is expressed on a subset of human HSC and on bone marrow vasculature, forming a hematogenic niche. finding
- ★ VAP-1+ HSC are a transcriptionally unique small subset of differentiated, proliferating HSC, whereas VAP-1- HSC are the most primitive HSC. finding
- ★ VAP-1-generated hydrogen peroxide acts via the p53 signaling pathway to regulate HSC proliferation. mechanism
- ★ VAP-1 functions as a check point-like inhibitor of HSC differentiation; its inhibition enhances HSC expansion and differentiation into colony-forming units. finding
- ★ VAP-1 in bone marrow vasculature supports HSC expansion, confirmed using VAP-1 knockout mice, enzymatically inactive VAP-1 knock-in mice, and an enzyme inhibitor. finding
- ★ VAP-1 expression enables characterization and prospective isolation of a new subset of human HSC. resource
- VAP-1 inhibitor (LJP-1586) can be used to expand HSC for potential clinical use. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Bulk RNA-seq | FACS-sorted VAP-1+ and VAP-1- HSC from human bone marrow (n=4 donors) | none (VAP-1+ vs VAP-1- sorting) | genome-wide gene expression / differentially expressed genes | SMART-Seq v4 Ultra Low Input RNA Kit (Takara), Nextera XT (Illumina), HiSeq 3000 |
| Single-cell RNA-seq | Sorted VAP-1+ and VAP-1- CD34+ Lin- human bone marrow cells (1 donor) | none (VAP-1+ vs VAP-1- sorting) | single-cell transcriptomes, clustering, trajectory | 10x Genomics Chromium Single Cell 3' v3.1; Illumina NovaSeq 6000; Cell Ranger 5.0.1 |
| Flow cytometry / FACS sorting | Human BM and cord blood cells; mouse BM cells | none | VAP-1 and HSC marker expression; cell sorting | FACSAria IIu, LSR Fortessa, Sony SH800 |
| Quantitative RT-PCR | Sorted human BM VAP-1+ and VAP-1- HSC (Lin-CD34+CD38-CD45RA-CD90+) | none | expression of TPX2, CDCA8, PCNA, MLLT3, TYMS normalized to B2M | TaqMan Fast Advanced Master Mix; 7900HT Fast Real-Time PCR System |
| Immunocytochemistry / immunofluorescence | FACS-sorted CD34+ cord blood cells; murine tissue sections; mouse femur whole-mount | none | VAP-1, CD31, CD150, Lineage localization | Olympus BX60, Zeiss LSM780 confocal, 3i Marianas spinning disk confocal |
| Colony-forming unit (CFU) assay | Human CB/BM CD34+ cells; mouse BM and peripheral blood cells | VAP-1 inhibitor LJP-1586 treatment; VAP-1-KO vs WT mice | number/type of colonies (CFUs) | MethoCult H4435/H4436/M3434 (STEMCELL Technologies) |
| Long-term culture-initiating cell (LTC-IC) assay | BM cells from WT and VAP-1-KO mice on irradiated stromal feeder layers | VAP-1 knockout | LTC-IC frequency (positive/negative scoring) | MyeloCult M5300 / MethoCult GF M3434 |
| ROS production measurement | Human CD34+ BM cells in liquid culture (9 days) | VAP-1 inhibitor LJP-1586 | reactive oxygen species production | — |
- – VAP-1+ HSC represent a transcriptionally distinct, small subset of differentiated and proliferating HSC, while VAP-1- HSC are the most primitive HSC.
- – Bulk RNAseq identified 687 VAP-1+ enriched and 378 VAP-1- enriched genes (fold change >1, p<0.05). 687 vs 378 genes
- – scRNAseq identified 371 VAP-1+ and 50 VAP-1- marker genes. 371 vs 50 genes
- ▲ HSC expansion and differentiation into colony-forming units are enhanced by inhibition of VAP-1.
- – VAP-1-generated hydrogen peroxide regulates HSC proliferation via the p53 signaling pathway.
- – Contribution of VAP-1 to HSC proliferation confirmed with VAP-1-deficient mice, mutated (enzymatically inactive) VAP-1 mice, and enzyme inhibitor treatment.
- count 687 VAP-1+ and 378 VAP-1- genes (fold change >1, p-value <0.05) (DEGs from bulk RNAseq used for Metascape analysis)
- count 371 VAP-1+ and 50 VAP-1- genes (marker genes from scRNAseq used for analysis)
- count 14,943 expressed genes (down from 60,619 total) (genes remaining after low-expression filtering in bulk RNAseq)
- count 117–150 VAP-1+ and VAP-1- HSC per individual (cells used for bulk RNA sequencing)
- count 11 donors (human bone marrow donors providing fresh BM cells)
- count 260–320 M reads/lane (sequencing depth on HiSeq3000 run)
- other ~10,000 cells per sample targeted; 400 cells/μl loaded (10x Genomics scRNAseq loading)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper characterizes VAP-1+ versus VAP-1− human HSC subpopulations using a multi-modal design: bulk RNA-seq (n=4 BM donors, TMM-normalized, GSEA pathway analysis), single-cell RNA-seq (one donor, two technical replicates per group, Seurat/Wilcoxon-based marker identification), qPCR validation, and functional assays (CFU, LTC-IC, liquid culture) in both human cells and VAP-1 KO/KI/WT mouse models. Differential gene expression in bulk RNA-seq is reported at fold change >1 and p<0.05 thresholds; scRNA-seq cluster markers are identified via Wilcoxon rank-sum test through Seurat's FindMarkers. The paper text is truncated before the functional assay statistics sections, so those tests cannot be fully characterized.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Wilcoxon rank-sum test (Seurat FindMarkers, default parameters) | scRNA-seq differential expression for cluster marker gene identification | Cells from one biological donor; two technical replicates per group (VAP-1+ and VAP-1−) | not stated |
| GSEA pre-ranked | Pathway and gene set enrichment analysis on bulk and scRNA-seq gene lists | 687 VAP-1+ and 378 VAP-1− genes from bulk RNAseq; 371 VAP-1+ and 50 VAP-1− genes from scRNAseq | not stated |
| Bulk RNA-seq differential expression (specific test not named; TMM normalization applied; likely edgeR or limma-voom per cited Kumar et al. pipeline) | VAP-1+ vs VAP-1− HSC bulk transcriptome comparison | n=4 human BM donors (Lonza); 117–150 cells per individual per group | not stated |
| PCA and UMAP dimensional reduction (non-statistical inference tests) | scRNA-seq dimensionality reduction and visualization | Cells from one donor; four samples (two technical replicates × two populations) | na |
| Monocle 2 trajectory inference (pseudotime ordering) | Developmental trajectory construction from scRNA-seq data | Cells from one donor | not stated |
| Metascape GO and pathway over-representation analysis | Gene Ontology biological process and pathway enrichment on DE gene lists | Gene lists derived from bulk and scRNA-seq analyses | not stated |
-
Bulk RNA-seq used TMM normalization (edgeR framework per the cited Kumar et al. pipeline) for differential expression between VAP-1+ and VAP-1− HSC↳ Could also: DESeq2 with median-of-ratios normalization and negative binomial Wald test — DESeq2 is another widely adopted framework that models count overdispersion explicitly; comparing results across both frameworks is a common robustness check in low-n transcriptomic studies (n=4 here)
-
scRNA-seq differential expression used the Wilcoxon rank-sum test via Seurat's FindMarkers with default parameters↳ Could also: MAST (Model-based Analysis of Single-cell Transcriptomics) or a mixed-effects model incorporating donor as a random effect — MAST accounts for the bimodal dropout structure of scRNA-seq data; a mixed-effects formulation would additionally accommodate the fact that all cells originate from a single donor, making the effective independent unit the cell rather than the individual
-
The scRNA-seq experiment used one biological donor with two technical replicates per group (VAP-1+ and VAP-1−)↳ Could also: Multiple biological donors with pseudo-bulk aggregation per donor before differential testing — Pseudo-bulk approaches (e.g., summing counts per donor then applying bulk-RNA-seq methods) treat the donor as the unit of replication, which better separates inter-individual from within-individual cell-to-cell variance and is increasingly recommended for single-cell differential expression
-
Trajectory analysis used Monocle 2 to order cells along a pseudotime axis↳ Could also: Monocle 3 or PAGA (partition-based graph abstraction, implemented in Scanpy) — Monocle 3 and PAGA use different graph-based algorithms that can model branching topologies more flexibly; comparing trajectories across methods is a standard robustness check given sensitivity of pseudotime to algorithm choice
-
Bulk RNA-seq DE gene lists were generated at a p<0.05 and FC>1 threshold with no explicitly described multiple-testing correction↳ Could also: Benjamini-Hochberg false discovery rate (FDR) correction at a defined q-value threshold (e.g., q<0.05 or q<0.10) — With ~15,000 expressed genes tested, the expected number of false positives under a nominal p<0.05 threshold alone is substantial; FDR control explicitly quantifies and limits the proportion of false discoveries in the reported gene list
-
GSEA pre-ranked was used for pathway enrichment on the full ranked gene list from bulk and scRNA-seq comparisons↳ Could also: Over-representation analysis (ORA) using a hypergeometric or Fisher's exact test on a discretized DE gene set — ORA is computationally simpler and more interpretable when a well-defined threshold gene list is already available; GSEA pre-ranked is more sensitive to coordinated moderate shifts across a pathway but requires a meaningful ranking metric, making the two approaches complementary
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34719737
Paper: Iftakhar-E-Khuda et al. (2021) Vascular adhesion protein-1 defines a unique subpopulation of human hematopoietic stem cells and regulates their proliferation. Cell Mol Life Sci. PMID 34719737 · PMCID PMC8629906 · DOI 10.1007/s00018-021-03977-6.
Data accessions (clarified — the BRIEF pointer is the BULK set)
- PRJNA594799 = BULK RNA-seq (this is the accession in BRIEF.md).
- PRJNA729883 = single-cell RNA-seq (SRA study SRP319821) — the source of the
headline computational/Nebulosa results. 4 sorted-HSC 10x samples (v3.1 dual-index),
raw FASTQ only:
- S2
210011_2_VAP_1_neg_replicate2(VAP-1⁻ rep2) - S3
210011_3_VAP_1_neg_replicate3(VAP-1⁻ rep3) - S4
210011_4_VAP_1_pos_replicate1(VAP-1⁺ rep1) - S5
210011_5_VAP_1_pos_replicate2(VAP-1⁺ rep2)
- S2
- No GEO Series / no processed count matrix / no Seurat object is shipped (GDS search empty; SRA records carry no GSM cross-ref). "Code availability" section is blank — no author code repo. → reproduction = apply the described third-party pipeline to the deposited FASTQ (valid per BRIEF rule P16).
Reported pipeline (Methods)
10x Chromium 3' v3.1 → Cell Ranger 5.0.1 (GRCh38) → Seurat 4.0.1 (R 4.0.5; FindVariableFeatures/ScaleData/RunPCA/FindMarkers Wilcoxon) → UMAP → Monocle 2 v2.10.1 (trajectory) → Nebulosa (kernel density estimation of gene expression). Nebulosa is the third-party GitHub tool in the BRIEF (powellgenomicslab/Nebulosa) and IS genuinely used.
IN SCOPE (pipeline-derived, reproducible)
- C1 AOC3/VAP-1 mRNA "practically negative" in the HSC scRNA-seq (the central Nebulosa-shown point). Robust, low-QC-sensitivity. → reproduce.
- C2 ~1267 high-quality HSC after QC (aggregate). → reproduce (approx; exact count is QC-threshold-sensitive = hard 20%).
- C3 Two clusters HSC1 (93%) / HSC2 (7%, ≈89 cells). → reproduce (proportion).
- C4 HSC2 enriched for proliferation / cell-cycle (S.Score, G2M.Score). → reproduce via Seurat CellCycleScoring.
DEVIATION (transparent)
- Cell Ranger 5.0.1 → STARsolo in CellRanger-emulation mode. Cell Ranger old-version
binaries are license-gated (no public direct URL). STARsolo with CR-emulation flags
(CB16/UMI12, 3M-february-2018 whitelist,
--soloUMIdedup 1MM_CR,--soloCBmatchWLtype 1MM_multi_Nbase_pseudocounts,--soloUMIfiltering MultiGeneUMI_CR,--clipAdapterType CellRanger4,--soloCellFilter EmptyDrops_CR) is the accepted free equivalent. Genome built from the public 10x GRCh38-2020-A fasta+GTF (same reference build the paper used). Counts will not be byte-identical to Cell Ranger; this is disclosed and graded accordingly.
OUT OF SCOPE (not attempted / hard 20%)
- Bulk RNA-seq DE marker derivation (PRJNA594799) — specific reported value not pinnable from the open text; would be a separate alignment+DESeq2 effort.
- Monocle 2 pseudotime trajectory — depends on exact cell set & is highly parameter-/seed- sensitive; not a crisp reported number.
- Exact 1267 / 89 cell counts to the unit — QC thresholds (mito%, nFeature) underspecified; we report our values and grade as partial/within-order rather than claim exact.
- All wet-lab / FACS / functional-proliferation assays — non-pipeline.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is an honest partial / setup-only reproduction: the agent correctly resolved a data-accession mismatch (bulk PRJNA594799 vs the real scRNA-seq PRJNA729883), rebuilt the GRCh38-2020-A reference + STAR index, and staged the full STARsolo→Seurat→Nebulosa pipeline, but an our-side download bug (32 zero-byte FASTQ files) meant no count matrix and zero graded numeric claims. The deviations that exist are on our/data-availability side — no processed matrix shipped, blank code-availability section, Cell-Ranger→STARsolo substitution — not evidence against the authors. The internally-consistent arithmetic (89 = 7% of 1267) is neither confirmed nor refuted, so there is no fabrication signal; criticality is yellow purely because the reproduction is incomplete, not because any claim broke.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.