Bioinformatics Strategies to Identify Shared Molecular Biomarkers That Link Ischemic Stroke and Moyamoya Disease with Glioblastoma.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🔴The deviation was non-trivial in magnitude
- 🔴The central claim did not (fully) hold under reproduction
- 🔴Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DIFFERENT (clear mismatch), not a drop. The data and tools are reproducible and the result is unambiguous: applying the paper's own named tools (DESeq2 on the shipped GSE106804 raw counts; RMA+limma on the shipped GSE131293 CEL files) with the paper's stated cutoff (p<0.05 & |log2FC|>1) yields DEG counts FAR below those reported. GBM: 611 vs 3585 reported (and the up/down skew is reversed: reproduced 521up/90down vs reported 1038up/2547down). Moyamoya: 60 probesets / 46 genes vs 1382 reported. Shared GBM-moyamoya genes: 0 vs 50 reported. A sensitivity sweep (p-only, |logFC|-only, FC>1.5, padj) shows NO reasonable threshold interpretation reaches the reported totals (moyamoya's 1382 even exceeds p-only=891 on a 3-vs-3 array, which is implausible). The repo ships only generic boilerplate R scripts that hardcode UNRELATED datasets (GSE104174 scleroderma/T2D, GSE42546 schizophrenia/depression) and never reference this paper's data or diseases, so the authors' exact code cannot reproduce the numbers either. Per brief P16 we used the named third-party tools on the paper's own data. NOT ATTEMPTED (the hard 20%): I.stroke DEGs (GSE56267) because GEO ships only gene-fusion reports, no expression count matrix (DESeq2 not directly runnable; would need SRA FASTQ alignment undescribed in Methods); and all web-tool downstream steps (EnrichR enrichment, STRING/NetworkAnalyst PPI, cytoHubba hub genes, TF/miRNA, DrugBank drugs) which are non-scriptable manual web steps. The reproducible foundational claims are a clear mismatch -> flagged as possible fabrication for human review.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 19assessed: 2026-06-15 ⛓ b38d0e01ded3
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetGlioblastoma is pathologically and molecularly connected to ischemic stroke and moyamoya disease, and these connections (shared genes, pathways, protein interactions) can be identified and characterized using a bioinformatics/network-based framework.
- ★ Glioblastoma, ischemic stroke, and moyamoya disease share molecular biomarkers and pathways indicating a pathological interconnection finding
- ★ A bioinformatics pipeline integrating differential expression, GO/pathway enrichment, PPI, TF, and miRNA analysis can identify shared disease biomarkers linking these three diseases method
- ★ Common differentially expressed genes (DEGs) exist between glioblastoma and ischemic stroke finding
- ★ Common DEGs exist between glioblastoma and moyamoya disease finding
- ★ Shared DEGs between disease pairs are enriched in specific KEGG, BioCarta, WikiPathways pathways and GO biological process terms finding
- Hub proteins identified from shared protein-protein interaction networks were used to predict candidate drugs via DrugBank resource
- DEG-transcription factor and DEG-miRNA regulatory relationships were identified for the shared biomarkers finding
- ★ Findings were cross-validated against gold-standard databases DisGeNET, dbGaP, and Rare-Diseases-AutoRIF method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RNA-seq differential expression analysis (DESeq2) | Extracellular vesicles from glioblastoma patients (n=13) vs healthy controls (n=6), dataset GSE106804 | disease (GBM) vs healthy | differentially expressed genes (up/downregulated) | DESeq2 (R package), Wald test, Cook's distance filtering |
| RNA-seq differential expression analysis (DESeq2) | Cortical tissue from ischemic stroke patients (n=7) vs healthy controls (n=6), dataset GSE56267 | disease (ischemic stroke) vs healthy | differentially expressed genes (up/downregulated) | DESeq2 (R package) |
| Microarray differential expression analysis | Neural crest stem cells from moyamoya patients (n=3) vs stable controls (n=3), dataset GSE131293 | disease (moyamoya) vs healthy | differentially expressed genes (up/downregulated) | Limma (R package), t-test |
| Gene Ontology and pathway enrichment analysis (GSEA-type) | Shared DEGs between GBM-ischemic stroke and GBM-moyamoya pairs | none | enriched biological pathways and GO Biological Process terms | EnrichR (KEGG, BioCarta, Reactome, WikiPathways, GO BP) |
| Protein-protein interaction network analysis | Shared DEGs/proteins between disease pairs | none | PPI network hub proteins (degree >15) | STRING database (confidence score 800), Network Analyst |
| Transcription factor and miRNA regulatory network analysis | Shared DEGs between disease pairs | none | DEG-TF and DEG-miRNA regulatory interactions | JASPAR, ENCODE, TarBase, miRTarBase, Cytoscape Network Analyzer/NetworkAnalyst |
| Drug-target prediction | Hub proteins from shared PPI networks (GBM-I.stroke, GBM-mm) | none | predicted protein-drug interactions | DrugBank database v5.0 via Network Analyst |
- – 3585 DEGs identified in glioblastoma (p<0.05, |logFC|>1) 1038 upregulated, 2547 downregulated
- – 1465 significant DEGs identified in ischemic stroke 1120 upregulated, 345 downregulated
- – 1382 significant DEGs identified in moyamoya disease 715 upregulated, 667 downregulated
- – 50 shared genes identified and used for pathway enrichment analysis between disease pairs 50 genes
- count 3585 DEGs (1038 up, 2547 down) (Glioblastoma DEGs, GSE106804 (13 patients, 6 controls))
- count 1465 DEGs (1120 up, 345 down) (Ischemic stroke DEGs, GSE56267 (7 patients, 6 controls))
- count 1382 DEGs (715 up, 667 down) (Moyamoya DEGs, GSE131293 (3 patients, 3 controls))
- pvalue p < 0.05 (Significance threshold for DEG selection)
- fold_change |log2FC| > 1 (Fold-change threshold for up/downregulated DEG classification)
- other confidence score 800, node degree > 15 (STRING PPI network construction criteria)
- count 50 shared genes (Shared genes used for pathway enrichment analysis between disease pairs)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This bioinformatics study applied differential expression analysis (DESeq2 for RNA-seq; limma for microarray) to three public GEO datasets representing glioblastoma, ischemic stroke, and moyamoya disease, filtering DEGs by p < 0.05 and |log2FC| ≥ 1. Shared DEGs between disease pairs were identified by set intersection, and downstream analyses included pathway/GO enrichment (EnrichR/GSEA), protein–protein interaction network analysis, transcription factor and miRNA regulatory analysis, and computational drug prediction. Results were reported as lists of significant genes, enriched pathways, and network hub nodes, with validation against DisGeNET, dbGaP, and Rare-Diseases-AutoRIF.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| DESeq2 Wald test (negative binomial GLM) | Differential expression analysis of RNA-seq data for glioblastoma (GSE106804) and ischemic stroke (GSE56267) | GBM: 13 patients + 6 controls; I. stroke: 7 patients + 6 controls | not stated |
| limma moderated t-test | Differential expression analysis of microarray data for moyamoya disease (GSE131293) | 3 patients + 3 controls | not stated |
| EnrichR GSEA (gene set enrichment analysis) | Pathway and Gene Ontology enrichment for shared DEGs between GBM–I.stroke and GBM–moyamoya pairs; databases: KEGG, WikiPathways, BioCarta, Reactome, GO:BP | — | not stated |
| Jaccard coefficient | Co-occurrence scoring for edge prediction in the gene-disease network | — | na |
-
A p < 0.05 threshold was applied to filter DEGs; whether this refers to raw or BH-adjusted p-values is not specified in the text↳ Could also: An explicitly stated Benjamini-Hochberg FDR-adjusted p-value threshold (e.g., q < 0.05 or q < 0.10) could also be applied and reported — With thousands of simultaneous statistical tests in a genome-wide setting, explicitly applying and reporting FDR correction is a common practice that quantifies the expected proportion of false positives among the reported DEGs and aids cross-study comparability
-
Differential expression was analyzed separately for RNA-seq (DESeq2) and microarray (limma) datasets, and DEGs were then overlapped across platforms↳ Could also: A cross-platform meta-analysis approach (e.g., RankProd, MetaDE, or limma with ComBat batch correction) could also be used to combine evidence across all three datasets simultaneously — Integrated analysis can increase statistical power by pooling samples and explicitly modeling platform as a covariate, potentially identifying more reproducible shared signals while accounting for technical differences between RNA-seq and microarray platforms
-
Enrichment analysis used the discrete overlapping DEG list (filtered by p < 0.05 and |logFC| ≥ 1) as input to EnrichR↳ Could also: Preranked GSEA using the full continuous ranked gene list (e.g., ranked by signed Wald statistic or −log10(p) × sign(FC)) could also be applied — Preranked GSEA avoids a hard threshold and can detect coordinated but moderate pathway-level shifts that might be missed when only genes exceeding a discrete fold-change or significance cutoff are considered
-
The moyamoya dataset contained n=3 per group, the minimum accepted by the study's own inclusion criterion↳ Could also: A formal statistical power analysis or sensitivity analysis for the smallest-n group could also be reported alongside the differential expression results — Documenting expected power at n=3 for a representative effect size contextualizes the likely completeness of the moyamoya DEG list and helps readers interpret downstream overlap results involving that disease
-
Shared biomarkers were identified by simple set intersection of DEG lists derived from separately analyzed datasets↳ Could also: A hypergeometric test or Fisher's exact test on the overlap size, given the DEG list sizes and total gene universe, could also be reported — A statistical test on the intersection itself provides a measure of whether the shared gene count exceeds chance expectation, adding a quantitative basis for interpreting the biological relevance of the overlap
-
Computational validation was performed using literature-curated databases (DisGeNET, dbGaP, Rare-Diseases-AutoRIF)↳ Could also: Validation in independent held-out GEO datasets for each disease, or permutation-based resampling of the DEG overlap, could also be employed as complementary validation strategies — Independent dataset validation provides orthogonal evidence not derived from the same literature corpus as the curated databases, allowing assessment of whether identified biomarkers are reproducible across different patient cohorts or experimental contexts
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
7 downstream papers · 3 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Engineered nanointerfaces for microfluidic isolation... 2018 · 241 cites
- Comprehensive <i>In Silico</i> Analysis of a Novel S... 2021 · 15 cites
- A tumor microenvironment model for glioma diagnosis... 2025 · 0 cites
- Targeting WTAP/ROR1/WNT5A-Mediated Crosstalk Between... 2026 · 0 cites
- A molecular brain atlas reveals cellular shifts duri... 2025 · 19 cites
- Cross-organ metabolite production and consumption in... 2025 · 7 cites
- Screening of key functional components of Taohong Si... 2023 · 5 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36015199
Paper: Bioinformatics Strategies to Identify Shared Molecular Biomarkers That Link Ischemic Stroke and Moyamoya Disease with Glioblastoma. Pharmaceutics 2022, 14(8):1573. PMID 36015199 · PMCID PMC9413912 · DOI 10.3390/pharmaceutics14081573.
Repo: https://github.com/hiddenntreasure/glioblastoma @ commit pinned in
reproduction/environment.lock (default branch main, pushed 2022-07-11).
What the paper does (pipeline)
A standard "shared-biomarker / disease-association" transcriptomics workflow over three GEO datasets (one per disease):
| disease | accession | platform | design | DEG tool (methods) |
|---|---|---|---|---|
| Glioblastoma (GBM) | GSE106804 | RNA-seq (EV) GPL? | 13 GBM vs 6 healthy | DESeq2 |
| Ischemic stroke | GSE56267 | RNA-seq GPL11154 (Illumina HiSeq2000) | 7 stroke vs 6 healthy cortex | DESeq2 |
| Moyamoya (mm) | GSE131293 | microarray, Affy HG-U133 Plus 2.0 | 3 MMD vs 3 control (iNCSC-VSMC) | Limma |
DEG cutoff (Methods §2.2): p-value < 0.05 and |log2FC| > 1. Then: shared DEGs (Venn) for the two pairs GBM∩stroke and GBM∩mm; downstream enrichment (EnrichR), PPI (STRING/NetworkAnalyst), hub genes (Cytoscape cytoHubba), TF/miRNA, drugs (DrugBank via NetworkAnalyst).
In scope (pipeline-derived, scriptable, clearly-specified) — ATTEMPTED
- GBM DEG counts (Table 1) via DESeq2 on the shipped raw count matrix
GSE106804_Gene_counts.txt.gz. Reported: 3585 total (1038 up, 2547 down). - Moyamoya DEG counts (Table 1) via RMA + Limma on the shipped CEL files. Reported: 1382 total (715 up, 667 down).
- Shared DEG count GBM∩mm (Fig 9 Venn) = 50, derivable by intersecting (1)&(2).
These are the foundational, low-hanging, fully-specified numeric outputs (80%).
Out of scope / not attempted (documented, not dropped)
- Stroke DEGs (GSE56267) — reproducibility gap. GEO supplementary ships only
*.FusionReport.txt.gz(gene-fusion detection output), no gene-level count matrix. Standard DESeq2 differential expression cannot be run on fusion reports; it would require pulling raw FASTQ from SRA (SRP040622), aligning + quantifying — a pipeline not described anywhere in the Methods (which only say "DESeq2"). This is the hard 20% and is undescribed → not attempted; recorded as a gap. Hence the GBM∩stroke Venn (59) and all stroke-side downstream numbers are not attempted. - Enrichment (EnrichR), PPI (STRING/NetworkAnalyst), hub genes (cytoHubba), TF/miRNA, drug prediction (DrugBank/NetworkAnalyst) — interactive web tools, no shipped code/parameters, manual curation ("Manual curation was used to limit pathways"). Non-scriptable, depend on exact uploaded gene list → out of scope.
Repo caveat (recorded for the auditor)
The repo's two R scripts are generic boilerplate, not this paper's analysis:
01. Limma.R hardcodes setwd(.../Scleroderma/SSc/GSE104174) with a CTRL/T2D
(type-2-diabetes) contrast; 02. DEseq-2.R hardcodes .../Schizophreniea/scz/ GSE42546 and a depression contrast. Neither mentions GSE106804/GSE56267/GSE131293,
stroke, moyamoya, or glioblastoma. So the shipped code does not reproduce the
paper's numbers as-is; per the brief (P16) we apply the named tools (DESeq2,
Limma) to the paper's own data with the paper's stated cutoff — an equally valid
reproduction route.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Every reproducible foundational count is a severe mismatch: GBM DEGs 611 vs reported 3585 (with the up/down skew reversed, 521up/90down vs 1038up/2547down), moyamoya 46 genes vs 1382, and 0 vs 50 shared GBM–moyamoya genes — so the paper's central biomarker-linkage claim does not hold. A sensitivity sweep shows no threshold interpretation reaches the reported totals (moyamoya's 1382 exceeds even p-only=891 on a 3-vs-3 array), and the repo ships only boilerplate scripts hardcoding unrelated datasets, so the numbers are not derivable from the shipped data or code. This sits squarely on the authors' side (value-not-derivable, fabrication-suspect), not on our methodology; the only data-availability caveat is the stroke arm (GSE56267), which lacks a count matrix and was out of scope.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.