Bioinformatics Strategies to Identify Shared Molecular Biomarkers That Link Ischemic Stroke and Moyamoya Disease with Glioblastoma.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🔴The deviation was non-trivial in magnitude
- 🔴The central claim did not (fully) hold under reproduction
- 🔴Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DIFFERENT (clear mismatch), not a drop. The data and tools are reproducible and the result is unambiguous: applying the paper's own named tools (DESeq2 on the shipped GSE106804 raw counts; RMA+limma on the shipped GSE131293 CEL files) with the paper's stated cutoff (p<0.05 & |log2FC|>1) yields DEG counts FAR below those reported. GBM: 611 vs 3585 reported (and the up/down skew is reversed: reproduced 521up/90down vs reported 1038up/2547down). Moyamoya: 60 probesets / 46 genes vs 1382 reported. Shared GBM-moyamoya genes: 0 vs 50 reported. A sensitivity sweep (p-only, |logFC|-only, FC>1.5, padj) shows NO reasonable threshold interpretation reaches the reported totals (moyamoya's 1382 even exceeds p-only=891 on a 3-vs-3 array, which is implausible). The repo ships only generic boilerplate R scripts that hardcode UNRELATED datasets (GSE104174 scleroderma/T2D, GSE42546 schizophrenia/depression) and never reference this paper's data or diseases, so the authors' exact code cannot reproduce the numbers either. Per brief P16 we used the named third-party tools on the paper's own data. NOT ATTEMPTED (the hard 20%): I.stroke DEGs (GSE56267) because GEO ships only gene-fusion reports, no expression count matrix (DESeq2 not directly runnable; would need SRA FASTQ alignment undescribed in Methods); and all web-tool downstream steps (EnrichR enrichment, STRING/NetworkAnalyst PPI, cytoHubba hub genes, TF/miRNA, DrugBank drugs) which are non-scriptable manual web steps. The reproducible foundational claims are a clear mismatch -> flagged as possible fabrication for human review.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 19assessed: 2026-06-15 ⛓ b38d0e01ded3
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe paper hypothesizes that glioblastoma (GBM) shares common molecular biomarkers, dysregulated genes, pathways, and protein–protein interactions with ischemic stroke and moyamoya disease, and uses a bioinformatics framework to identify and validate these shared links to illuminate the diseases' interrelated mechanisms.
- ★ Shared differentially expressed genes link glioblastoma with ischemic stroke and with moyamoya disease, revealing molecular associations among the diseases. finding
- ★ A generalized network-based bioinformatics workflow integrating DEG analysis, GO/pathway enrichment, PPI, TF/miRNA regulation, and gold-standard validation can identify shared biomarkers across the three diseases. method
- ★ 3585 DEGs were identified in glioblastoma (1038 up, 2547 down). finding
- ★ 1465 significant DEGs were identified in ischemic stroke (1120 up, 345 down). finding
- ★ 1382 significant DEGs were identified in moyamoya disease (715 up, 667 down). finding
- ★ Enriched pathways from KEGG, WikiPathways, and BioCarta were found significantly linked to DEGs shared between GBM–ischemic stroke and GBM–moyamoya pairs. finding
- Candidate drugs for glioblastoma and its associated diseases were predicted from hub proteins using DrugBank. resource
- Hypoxia is proposed as a shared mechanism predisposing both glioma/tumors and cerebral ischemia. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq differential expression analysis | glioblastoma patient extracellular vesicles (13 patients, 6 healthy controls) | none (disease vs healthy) | differentially expressed genes (DEGs) | DESeq2 (R package), GEO GSE106804 |
| bulk RNA-seq differential expression analysis | ischemic stroke patient cortical tissue (7 patients, 6 healthy controls) | none (disease vs healthy) | differentially expressed genes (DEGs) | DESeq2 (R package), GEO GSE56267 |
| microarray differential expression analysis | moyamoya patient neural crest stem cells (3 patients, 3 controls) | none (disease vs healthy) | differentially expressed genes (DEGs) | Limma (R package), GEO GSE131293 |
| Gene Ontology and pathway enrichment analysis (GSEA) | shared DEGs across disease pairs | none | enriched GO BP terms and pathways | EnrichR; KEGG, BioCarta, Reactome, WikiPathways |
| protein–protein interaction network analysis | shared DEG-encoded proteins | none | PPI network hub proteins | STRING (confidence 800, degree>15), Network Analyst |
| transcription factor and miRNA regulatory analysis | identified DEGs | none | DEG-TF and DEG-miRNA interactions | JASPAR, ENCODE, TarBase, miRTarBase, Cytoscape Network Analyzer, Network Analyst |
| drug–protein interaction prediction | hub proteins shared across disease pairs | none | predicted drugs | DrugBank v5.0, Network Analyst |
| gold-standard validation | biomarker genes and pathways | none | validation against curated disease databases | DisGeNET, dbGaP, Rare-Diseases-AutoRIF |
- – 3585 DEGs identified in glioblastoma 1038 up / 2547 down
- – 1465 significant DEGs identified in ischemic stroke 1120 up / 345 down
- – 1382 significant DEGs identified in moyamoya disease 715 up / 667 down
- – 50 genes shared between glioblastoma and its associated disease(s) used for pathway enrichment 50 genes
- – Enriched pathways from KEGG, WikiPathways, and BioCarta significantly linked to shared DEGs
- count 3585 DEGs (glioblastoma DEGs (p<0.05, |logFC|>1); dataset GSE106804, 13 patients + 6 controls)
- count 1465 DEGs (ischemic stroke DEGs; dataset GSE56267, 7 patients + 6 controls)
- count 1382 DEGs (moyamoya DEGs; dataset GSE131293, 3 patients + 3 controls)
- pvalue p < 0.05 (significance threshold for DEGs and enriched pathways)
- fold_change absolute log2 fold change > 1 (DEG selection cutoff for up/downregulation)
- count 50 shared genes (common DEGs between GBM and associated diseases used in pathway enrichment)
- other STRING confidence 800, degree > 15 (PPI network construction thresholds)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This bioinformatics study applied differential expression analysis (DESeq2 for RNA-seq; limma for microarray) to three public GEO datasets representing glioblastoma, ischemic stroke, and moyamoya disease, filtering DEGs by p < 0.05 and |log2FC| ≥ 1. Shared DEGs between disease pairs were identified by set intersection, and downstream analyses included pathway/GO enrichment (EnrichR/GSEA), protein–protein interaction network analysis, transcription factor and miRNA regulatory analysis, and computational drug prediction. Results were reported as lists of significant genes, enriched pathways, and network hub nodes, with validation against DisGeNET, dbGaP, and Rare-Diseases-AutoRIF.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| DESeq2 Wald test (negative binomial GLM) | Differential expression analysis of RNA-seq data for glioblastoma (GSE106804) and ischemic stroke (GSE56267) | GBM: 13 patients + 6 controls; I. stroke: 7 patients + 6 controls | not stated |
| limma moderated t-test | Differential expression analysis of microarray data for moyamoya disease (GSE131293) | 3 patients + 3 controls | not stated |
| EnrichR GSEA (gene set enrichment analysis) | Pathway and Gene Ontology enrichment for shared DEGs between GBM–I.stroke and GBM–moyamoya pairs; databases: KEGG, WikiPathways, BioCarta, Reactome, GO:BP | — | not stated |
| Jaccard coefficient | Co-occurrence scoring for edge prediction in the gene-disease network | — | na |
-
A p < 0.05 threshold was applied to filter DEGs; whether this refers to raw or BH-adjusted p-values is not specified in the text↳ Could also: An explicitly stated Benjamini-Hochberg FDR-adjusted p-value threshold (e.g., q < 0.05 or q < 0.10) could also be applied and reported — With thousands of simultaneous statistical tests in a genome-wide setting, explicitly applying and reporting FDR correction is a common practice that quantifies the expected proportion of false positives among the reported DEGs and aids cross-study comparability
-
Differential expression was analyzed separately for RNA-seq (DESeq2) and microarray (limma) datasets, and DEGs were then overlapped across platforms↳ Could also: A cross-platform meta-analysis approach (e.g., RankProd, MetaDE, or limma with ComBat batch correction) could also be used to combine evidence across all three datasets simultaneously — Integrated analysis can increase statistical power by pooling samples and explicitly modeling platform as a covariate, potentially identifying more reproducible shared signals while accounting for technical differences between RNA-seq and microarray platforms
-
Enrichment analysis used the discrete overlapping DEG list (filtered by p < 0.05 and |logFC| ≥ 1) as input to EnrichR↳ Could also: Preranked GSEA using the full continuous ranked gene list (e.g., ranked by signed Wald statistic or −log10(p) × sign(FC)) could also be applied — Preranked GSEA avoids a hard threshold and can detect coordinated but moderate pathway-level shifts that might be missed when only genes exceeding a discrete fold-change or significance cutoff are considered
-
The moyamoya dataset contained n=3 per group, the minimum accepted by the study's own inclusion criterion↳ Could also: A formal statistical power analysis or sensitivity analysis for the smallest-n group could also be reported alongside the differential expression results — Documenting expected power at n=3 for a representative effect size contextualizes the likely completeness of the moyamoya DEG list and helps readers interpret downstream overlap results involving that disease
-
Shared biomarkers were identified by simple set intersection of DEG lists derived from separately analyzed datasets↳ Could also: A hypergeometric test or Fisher's exact test on the overlap size, given the DEG list sizes and total gene universe, could also be reported — A statistical test on the intersection itself provides a measure of whether the shared gene count exceeds chance expectation, adding a quantitative basis for interpreting the biological relevance of the overlap
-
Computational validation was performed using literature-curated databases (DisGeNET, dbGaP, Rare-Diseases-AutoRIF)↳ Could also: Validation in independent held-out GEO datasets for each disease, or permutation-based resampling of the DEG overlap, could also be employed as complementary validation strategies — Independent dataset validation provides orthogonal evidence not derived from the same literature corpus as the curated databases, allowing assessment of whether identified biomarkers are reproducible across different patient cohorts or experimental contexts
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
7 downstream papers · 3 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Engineered nanointerfaces for microfluidic isolation... 2018 · 241 cites
- Comprehensive <i>In Silico</i> Analysis of a Novel S... 2021 · 15 cites
- A tumor microenvironment model for glioma diagnosis... 2025 · 0 cites
- Targeting WTAP/ROR1/WNT5A-Mediated Crosstalk Between... 2026 · 0 cites
- A molecular brain atlas reveals cellular shifts duri... 2025 · 19 cites
- Cross-organ metabolite production and consumption in... 2025 · 7 cites
- Screening of key functional components of Taohong Si... 2023 · 5 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36015199
Paper: Bioinformatics Strategies to Identify Shared Molecular Biomarkers That Link Ischemic Stroke and Moyamoya Disease with Glioblastoma. Pharmaceutics 2022, 14(8):1573. PMID 36015199 · PMCID PMC9413912 · DOI 10.3390/pharmaceutics14081573.
Repo: https://github.com/hiddenntreasure/glioblastoma @ commit pinned in
reproduction/environment.lock (default branch main, pushed 2022-07-11).
What the paper does (pipeline)
A standard "shared-biomarker / disease-association" transcriptomics workflow over three GEO datasets (one per disease):
| disease | accession | platform | design | DEG tool (methods) |
|---|---|---|---|---|
| Glioblastoma (GBM) | GSE106804 | RNA-seq (EV) GPL? | 13 GBM vs 6 healthy | DESeq2 |
| Ischemic stroke | GSE56267 | RNA-seq GPL11154 (Illumina HiSeq2000) | 7 stroke vs 6 healthy cortex | DESeq2 |
| Moyamoya (mm) | GSE131293 | microarray, Affy HG-U133 Plus 2.0 | 3 MMD vs 3 control (iNCSC-VSMC) | Limma |
DEG cutoff (Methods §2.2): p-value < 0.05 and |log2FC| > 1. Then: shared DEGs (Venn) for the two pairs GBM∩stroke and GBM∩mm; downstream enrichment (EnrichR), PPI (STRING/NetworkAnalyst), hub genes (Cytoscape cytoHubba), TF/miRNA, drugs (DrugBank via NetworkAnalyst).
In scope (pipeline-derived, scriptable, clearly-specified) — ATTEMPTED
- GBM DEG counts (Table 1) via DESeq2 on the shipped raw count matrix
GSE106804_Gene_counts.txt.gz. Reported: 3585 total (1038 up, 2547 down). - Moyamoya DEG counts (Table 1) via RMA + Limma on the shipped CEL files. Reported: 1382 total (715 up, 667 down).
- Shared DEG count GBM∩mm (Fig 9 Venn) = 50, derivable by intersecting (1)&(2).
These are the foundational, low-hanging, fully-specified numeric outputs (80%).
Out of scope / not attempted (documented, not dropped)
- Stroke DEGs (GSE56267) — reproducibility gap. GEO supplementary ships only
*.FusionReport.txt.gz(gene-fusion detection output), no gene-level count matrix. Standard DESeq2 differential expression cannot be run on fusion reports; it would require pulling raw FASTQ from SRA (SRP040622), aligning + quantifying — a pipeline not described anywhere in the Methods (which only say "DESeq2"). This is the hard 20% and is undescribed → not attempted; recorded as a gap. Hence the GBM∩stroke Venn (59) and all stroke-side downstream numbers are not attempted. - Enrichment (EnrichR), PPI (STRING/NetworkAnalyst), hub genes (cytoHubba), TF/miRNA, drug prediction (DrugBank/NetworkAnalyst) — interactive web tools, no shipped code/parameters, manual curation ("Manual curation was used to limit pathways"). Non-scriptable, depend on exact uploaded gene list → out of scope.
Repo caveat (recorded for the auditor)
The repo's two R scripts are generic boilerplate, not this paper's analysis:
01. Limma.R hardcodes setwd(.../Scleroderma/SSc/GSE104174) with a CTRL/T2D
(type-2-diabetes) contrast; 02. DEseq-2.R hardcodes .../Schizophreniea/scz/ GSE42546 and a depression contrast. Neither mentions GSE106804/GSE56267/GSE131293,
stroke, moyamoya, or glioblastoma. So the shipped code does not reproduce the
paper's numbers as-is; per the brief (P16) we apply the named tools (DESeq2,
Limma) to the paper's own data with the paper's stated cutoff — an equally valid
reproduction route.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Every reproducible foundational count is a severe mismatch: GBM DEGs 611 vs reported 3585 (with the up/down skew reversed, 521up/90down vs 1038up/2547down), moyamoya 46 genes vs 1382, and 0 vs 50 shared GBM–moyamoya genes — so the paper's central biomarker-linkage claim does not hold. A sensitivity sweep shows no threshold interpretation reaches the reported totals (moyamoya's 1382 exceeds even p-only=891 on a 3-vs-3 array), and the repo ships only boilerplate scripts hardcoding unrelated datasets, so the numbers are not derivable from the shipped data or code. This sits squarely on the authors' side (value-not-derivable, fabrication-suspect), not on our methodology; the only data-availability caveat is the stroke arm (GSE56267), which lacks a count matrix and was out of scope.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.