Zebrafish functional xenograft vasculature platform identifies PF-502 as a durable vasculature normalization drug.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough; reproducible. The paper has exactly one computational dataset: bulk RNA-seq of HUVECs treated with vehicle (DMSO) vs PF-502 0.15 uM for 12 h (GSE224106, n=3/group); everything else (zebrafish xenograft imaging, drug screen, RT-PCR, IHC, in-vivo) is wet-lab and out of scope. The code statement says no original code, so per BRIEF P16 the standard described pipeline (DESeq2 + hallmark GSEA) was run on the paper's own GEO matrix. This is a FRESH re-run (2026-06-22, «our HPC» SLURM «job») after the room was re-queued: the re-queue was triggered because the prior 2026-06-16 ROOM_RESULT.json lacked the now-mandatory qc_room and datasets blocks (the underlying reproduction was sound). The GEO matrix re-downloaded byte-identical (sha256 79abb23e...). CORE result is a clean 1:1: the matrix ships both raw counts and the authors' own DESeq2 stats, and counting DEGs at the paper's literal threshold (p.adj<0.5 & |log2FC|>0.5) gives EXACTLY 1364 down / 738 up as reported; an independent DESeq2 1.42.0 re-run from the raw counts reproduces log2FoldChange essentially perfectly (Pearson r=1.0000, n=8625) and the DEG counts within 3% (1398 down / 717 up). No fabrication indicators: every reported DEG number is exactly recoverable from the deposited data. The HARDER GSEA NES values are PARTIAL: fgsea reproduces the paper's direction, FDR0, and rank order (E2F -3.54 > G2M -3.18, mTORC1 -3.28; literal PI3K/AKT/mTOR -1.37) but with larger NES magnitudes than the paper's -2.10/-2.31/-2.34. Two independent engines (fgsea + gseapy, the latter from the 2026-06-16 run) agree within ~0.05 NES, so the gap is a ranking-metric/MSigDB-version difference (gsea-3.0.jar defaults to Signal2Noise on expression with phenotype labels, not the DESeq2 stat); the faithful classic S2N+phenotype route is degenerate at n=3/group and cannot recover the numbers either. NOT ATTEMPTED: from-FASTQ STAR alignment to hg19 (redundant given the shipped matrix + r=1.0 concordance); clusterProfiler GO/KEGG (qualitative only); RT-PCR/zebrafish/in-vivo (wet-lab).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 70assessed: 2026-06-16 ⛓ 169ae47bfa82
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether a modified zebrafish tumor xenograft platform (zFXVP) can visualize and quantify functional tumor vasculature well enough to screen for drugs that durably normalize tumor blood vessels, using PI3K/mTOR inhibitors as a test compound class.
- ★ The zebrafish functional xenograft vasculature platform (zFXVP) enables visualization and quantification of structurally and functionally realistic tumor vasculature formation. resource
- ★ zFXVP-based screening of 11 PI3K-pathway inhibitors identified PF-502 (PF-04691502, S2743) as the most effective vascular normalization agent, uniquely combining reduced vessel density with enhanced blood perfusion. finding
- ★ PF-502 induces durable vascular normalization compared to the narrow normalization window of VEGFR2 TKI (nintedanib). finding
- ★ PF-502 induces endothelial cell-cycle arrest, consistent with reduced neo-vessel density. mechanism
- ★ PF-502 induces S3 cleavage of Notch1, releases NICD, and activates downstream Notch target genes (Hes1, Hey1) in endothelial cells. mechanism
- ★ PF-502's vascular lumen-stabilizing effect depends mainly on activation of Notch signaling, as shown by pharmacological/genetic Notch blockade. mechanism
- ★ Low-dose PF-502 induces durable vascular normalization and improves drug delivery in mouse tumor models, synergizing with chemotherapy. finding
- Tumor vasculature in zFXVP xenografts recapitulates key structural and functional abnormalities of mouse and human tumor vasculature (hyperproliferation, dilated/disorganized capillaries, hyperpermeability, impaired perfusion). finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Fluorescence confocal/stereomicroscope imaging of xenograft vasculature | Zebrafish embryos (kdrl:EGFP, gata1:dsRed, flk-N:EGFP transgenics) implanted with CT26, Hepa1-6, U87, Gl261, LL2, or B16 tumor cells | tumor cell implantation into perivitelline space | vascular branch density, vessel diameter, perfusion | spinning disk confocal (SpinSR10, Olympus) plus VAST BioImager-like automated handling; Imaris/FIJI image analysis |
| EdU proliferation assay | Zebrafish CT26 xenograft endothelial cells vs neighboring intestinal vessels | none (baseline characterization) | number of EdU-positive endothelial cells | — |
| Dextran-tomato (70k MW) angiography | Zebrafish CT26 xenograft (kdrl:EGFP) | none (baseline characterization of vascular permeability) | extravasation/leakage of fluorescent tracer from neo-vessels vs normal vessels | — |
| Live imaging of blood cell flow | Zebrafish xenograft (gata1:dsRed/kdrl:EGFP double transgenic, BFP-labeled tumor) | none (baseline characterization) | percentage of blood-perfused vessels, blood flow behavior | — |
| Compound screen with vascular imaging/quantification | Zebrafish tumor xenografts | 11 PI3K-pathway inhibitor candidates (1 μM, 2-5 dpi) | vascular density, ratio of blood-perfused vessels, vascular branch points, zebrafish developmental toxicity | zFXVP-VAST-Imaris/FIJI batch analysis system |
| Dose-response vascular imaging | Zebrafish tumor xenografts | PF-502 vs nintedanib (VEGFR2 TKI) at varying doses | vessel density, normalization window | — |
| Molecular analysis of Notch1 cleavage and target gene expression | Endothelial cells | PF-502 treatment | Notch1 S3 cleavage, NICD release, Hes1/Hey1 expression | — |
| Tumor vasculature and drug delivery analysis | Mouse tumor model | low-dose (sub-maximum tolerated dose) PF-502, alone or with chemotherapy | vascular normalization, drug delivery efficiency, therapeutic synergy | — |
- ▲ Endothelial cell proliferation in tumor xenograft vessels increased relative to neighboring intestinal vessels 4-fold
- – Among 11 PI3K-pathway inhibitors screened, PF-502 significantly reduced xenograft vascular complexity while greatly enhancing blood perfusion rate
- – PF-502 had relatively minor impact on zebrafish embryonic developmental stage compared to other tested compounds
- – Manual re-analysis confirmed PF-502's vascular normalization effect (vessel diameter, density, functional vessel percentage)
- – PF-502 normalization effect confirmed in additional xenograft models (LL2, GL261)
- ▼ Both PF-502 and nintedanib reduced vessel density dose-dependently, but nintedanib caused complete vascular inhibition within a much narrower concentration range than PF-502
- ▲ Dextran-tomato tracer extravasated from tumor xenograft neo-vessels but was retained in normal vessels of the same fish
- – Blood perfusion percentage differed significantly between tumor and normal vessels
- fold_change 4 times increase (endothelial cell proliferation in tumor xenograft vessels vs neighboring intestinal vessels (EdU assay))
- pvalue p < 0.5 (as printed) (EdU-positive endothelial cell count comparison, Student's t test)
- pvalue p < 0.0001 (percentage of blood perfusion in tumor vessels, Student's t test)
- pvalue p < 0.0001, p < 0.01, p < 0.05 (vascular density, blood-perfused vessel ratio, and branch points across 11 tested compounds vs control)
- pvalue p < 0.0001, p < 0.01 (vessel diameter, blood vessel density, percentage of functional vessels with vs without PF-502 treatment)
- count n = 20 (zebrafish used to assess percentage of normal development after 5 days of compound treatment)
- count n = 5 (distribution of endothelial nuclei per 100 μm blood vessel length, tumor vs normal vessels)
- other 7-40 times faster proliferation (cited literature value for proliferating endothelial cell ratio in mouse/human tumors vs normal tissue)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used a zebrafish functional xenograft vasculature platform (zFXVP) combined with spinning-disk confocal imaging and batch image analysis (Imaris/FIJI) to screen 11 PI3K inhibitor compounds for vascular normalization activity, followed by individual confirmation experiments. Group differences were assessed primarily with Student's t-tests, comparing drug-treated versus DMSO-treated zebrafish xenografts or tumor versus normal tissue. Results were reported as means ± SEM with threshold significance symbols, and key findings were extended to murine models.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Student's t-test (two-sample, unpaired implied) | Figure 2B — EdU-positive endothelial cells in tumor xenograft vs. neighboring intestinal vessels | n = 3 | not stated |
| Student's t-test (two-sample, unpaired implied) | Figure 2G — percentage of blood perfusion in tumor vessels vs. normal vessels | not stated | not stated |
| Student's t-test (two-sample, unpaired implied) | Figure 3B — vascular branch density, blood-perfusion ratio, and branch points for each of 11 compounds vs. DMSO control | 6–10 animals per condition | not stated |
| Student's t-test (two-sample, unpaired implied) | Figure 3E — vessel diameter, blood vessel density, and percentage functional vessels in xenografts with vs. without PF-502 | not stated | not stated |
-
Eleven compounds were each compared to a single DMSO control using separate Student's t-tests across multiple vascular endpoints↳ Could also: A one-way ANOVA (or two-way ANOVA for compound × endpoint) followed by a post-hoc correction such as Dunnett's test (many-vs-one) or Tukey HSD (all pairwise) could also be applied — An ANOVA framework would account for the family of comparisons made simultaneously, keeping the experiment-wise Type I error rate at the nominal level; Dunnett's test is specifically designed for the many-treatments-vs-one-control structure used here
-
Dispersion was reported as SEM throughout, including for groups with n = 3 and n = 5↳ Could also: SD or a 95% confidence interval could also be used to express spread — SD describes the variability of the observations themselves rather than the precision of the mean estimate; with small n, the distinction between SEM and SD is large, and many reporting guidelines (e.g., Nature Methods, ARRIVE) recommend SD or CI for biological replicates so readers can judge biological variability
-
Significance was reported only as threshold symbols (p < 0.05, p < 0.01, p < 0.0001) rather than exact p-values↳ Could also: Exact p-values (e.g., p = 0.012) could also be reported — Exact p-values allow readers and meta-analysts to apply their own threshold judgments, facilitate effect-size calculations, and are recommended by many journals and statistical guidelines
-
Group differences were assessed with Student's t-tests with sample sizes as small as n = 3↳ Could also: A non-parametric alternative such as the Mann-Whitney U test could also be applied — With very small n, the normality assumption underlying the t-test cannot be verified empirically; non-parametric tests make no distributional assumption, which is one reason they are sometimes preferred when n < 5–6 per group
-
Results were summarized as mean differences without reporting standardized effect sizes↳ Could also: Cohen's d or a ratio-based effect size (fold change with CI) could also be reported alongside p-values — Effect sizes allow readers to judge the biological magnitude of differences independently of sample size, and are increasingly requested by journals to aid interpretation and future power calculations
-
Randomization and blinding procedures for zebrafish assignment to treatment groups and for image analysis were not described↳ Could also: Pre-specified randomization of larvae to treatment groups and blinded image quantification could also be documented — Documenting these procedures (per ARRIVE 2.0 guidelines for animal studies) allows readers to assess the risk of allocation and assessment bias; blinded quantification is particularly relevant here because image-based vascular metrics involve subjective threshold choices
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37680473
Paper: Zhong et al. 2023, iScience. "Zebrafish functional xenograft vasculature platform identifies PF-502 as a durable vasculature normalization drug." DOI 10.1016/j.isci.2023.107734 · PMID 37680473 · PMCID PMC10480778.
The single RNA-seq experiment (the only computational dataset)
- GSE224106 — bulk RNA-seq of primary HUVECs, 2 conditions × 3 replicates:
- GSM7011986/87/88 = ctrl (DMSO)
- GSM7011989/90/91 = PF-502 (0.15 μM, 12 h)
- Platform GPL24676 (Illumina NovaSeq 6000), 150 bp paired-end.
- BioProject PRJNA929912 (raw FASTQ in SRA); GEO ships a processed matrix
GSE224106_All.RNAseq.data.csv.gz(~847 KB).
Pipeline described in Methods
fastp (QC/trim) → STAR 2.6.0.a (align to hg19) → GenomicFeatures + Rsamtools (quantify) → DESeq2 (normalize + stats) → thresholds p.adj < 0.5 and |log2FoldChange| > 0.5 → clusterProfiler (GO/KEGG) + GSEA (gsea-3.0.jar, hallmark sets).
NOTE: the paper states "p.adj < 0.5" (not 0.05) — unusual; we test the literal value and also the conventional 0.05, and report both.
IN SCOPE (pipeline-derived, attempted)
| id | reported result | location | pipeline | difficulty |
|---|---|---|---|---|
| C1 | 1,364 downregulated genes (p.adj<0.5, log2FC<-0.5) | Results / Fig 5 | DESeq2 on count matrix | CORE (80%) |
| C2 | 738 upregulated genes (p.adj<0.5, log2FC>0.5) | Results / Fig 5 | DESeq2 on count matrix | CORE (80%) |
| C3 | GSEA PI3K/AKT/mTOR NES = -2.10, FDR q = 0.0 | Results / Fig 5 | GSEA hallmark | HARDER |
| C4 | GSEA G2M checkpoint NES = -2.31, FDR q = 0.0 | Results / Fig 5 | GSEA hallmark | HARDER |
| C5 | GSEA E2F targets NES = -2.34, FDR q = 0.0 | Results / Fig 5 | GSEA hallmark | HARDER |
| C6 | Notch pathway up (DLL4, NOTCH1, NOTCH4, HES1, HEY2 direction) | Results / Fig 5 | DESeq2 direction | STRETCH |
OUT OF SCOPE (wet-lab / manual / not pipeline-derived — not attempted)
- Zebrafish xenograft vasculature imaging, tumor growth, microscopy quantification.
- RT-PCR validation of DLL4/Notch genes (wet-lab; we only check DEG direction).
- Drug screening / PF-502 identification, IHC, survival, all in-vivo assays.
Code statement
Paper says "This paper does not report original code." The BRIEF lists github.com/pangxueyu233/Pipeline-of-transcriptome (corresponding-author's generic transcriptome pipeline). Per BRIEF P16, applying that / a standard DESeq2+GSEA pipeline to the paper's own data is equally valid. We reproduce by running the described standard pipeline (DESeq2, fgsea/GSEA hallmark) on the paper's matrix.
Strategy
- CORE: DESeq2 on the GEO processed count matrix → DEG counts (C1, C2). This is the most direct 1:1 and the quick minimum.
- HARDER: rank genes, run fgsea/GSEA against MSigDB hallmark → NES (C3–C5).
- STRETCH: full STAR(hg19) alignment from SRA FASTQ → counts → DESeq2, if the processed-matrix route leaves ambiguity.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The single computational dataset (GSE224106) ships raw counts plus the authors' own DESeq2 output, so the core DEG claims reproduce 1:1: 1364 down / 738 up exactly at the literal padj<0.5 threshold, with log2FC concordance r=1.0000 on an independent re-run — no fabrication indicators. The only deviation is in the GSEA NES magnitudes (paper -2.1/-2.3 vs reproduced -3.2/-3.5), which is on the methodology/version axis (our DESeq2-Wald preranked route vs the authors' likely S2N/gsea-3.0.jar + older MSigDB), since the faithful classic route is degenerate at n=3/group. Direction, FDR0, and rank order all hold, so the central conclusion is confirmed; the residual gap is moderate and explainable, hence an overall yellow rather than green.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.