Network Controllability Reveals Key Mitigation Points for Tumor-Promoting Signaling in Tumor-Educated Platelets.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Any deviation was negligible
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH + 1:1 on the headline. Real repo is ozgeosmanoglu/Platelet_NSCLC (the scaffold's plotly/plotly.R code_url was a harvester false-positive, corrected). The repo ships data/myDGEListFiltF (filtered/collapsed/outlier-removed DGEList, 14545x776, sha256 0a291a58...b7f05) = the exact DE input, so the differential-expression result reproduces WITHOUT re-aligning the 826 SRA runs. On «our HPC» (R 4.5.3, limma 3.66.0, RUVSeq 1.44.0, DESeq2 1.50.2) I reran the authors' RUVSeq(k=5)->limma-voom and ->DESeq2-LRT chain on that DGEList. limma-voom: 110 up / 107 down = 217 DEGs vs reported 111/108=219 -> off by 1 gene per direction (~0.9%), a clean within-tolerance match; NSCLC sample count 402 reproduced EXACTLY. Important honest finding: the repo's DESeq2-LRT path gives 593 DEGs, so the paper's reported figure is specifically the limma result (and the downstream network correctly uses limma logFC). NOT ATTEMPTED (declared out of scope, the hard ~20%): network-controllability node classes (62 indispensable/127 neutral/212 dispensable/86 critical), network size (401 nodes/962 edges), the 5 key genes, and drug repurposing -- Step7.1.2 only READS the classification from NetworkAnalysis.xlsx (output of the CytoCtrlAnalyzer Cytoscape GUI app, NOT shipped), and the network build requires a live/time-varying OmnipathR pull plus an unshipped proteomics file; these numbers are not derivable from the deposited artifacts and are flagged for the human reviewer, not asserted.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 84assessed: 2026-06-14 ⛓ 4e1ec9c8d2c8
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan integrative transcriptome and network controllability analysis of tumor-educated platelets (TEPs) in NSCLC identify key signaling nodes and FDA-approved drugs that therapeutically disrupt metastasis-promoting platelet–tumor signaling (e.g., ITAM, P2Y12)?
- ★ TEPs in NSCLC show 111 upregulated and 108 downregulated genes versus non-cancer control platelets, enriched in ECM interaction, cytoskeleton, immune signaling, and platelet activation pathways. finding
- ★ Four complementary strategies identify five high-confidence central TEP signaling genes: ITGA2B, FLNA, GRB2, FCGR2A, and APP, all targetable by FDA-approved drugs. finding
- ★ Fostamatinib, an SYK inhibitor, is the top candidate drug to selectively disrupt ITAM-mediated platelet activation. finding
- ★ Network controllability analysis of a TEP-specific signaling network identifies critical and indispensable nodes as key mitigation points for tumor-promoting signaling. method
- ★ A low-dose combination therapy of fostamatinib, aducanumab, and acetylsalicylic acid (aspirin) may control TEP effects. finding
- ★ A TEP-specific signaling network was constructed by integrating GSE89843 transcriptome data with directed, signed protein–protein interactions from OmniPath. resource
- The reconstructed platelet signaling network contains 962 interactions among 401 platelet-specific proteins, with 86 critical nodes and 62 indispensable nodes. finding
- k-means clustering of DEGs yields nine modules; four upregulated modules represent protein homeostasis, platelet activation, metabolism, and cytoskeletal remodeling. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq / transcriptome differential expression analysis | platelets from NSCLC patients vs non-cancer donors (human) | none (disease vs control comparison) | differentially expressed genes (up/downregulated) | GEO dataset GSE89843 |
| pathway enrichment analysis (KEGG, GO Biological Process, Reactome) | DEGs from human TEPs | none | enriched pathways/processes | — |
| k-means / hierarchical clustering of DEGs | human platelet samples across NSCLC, healthy, and inflammatory conditions | none | gene expression modules (9 modules) | — |
| gene set enrichment analysis (GSEA) | human TEP transcriptome | none | normalized enrichment score (NES) of Reactome pathways | — |
| protein–protein interaction network reconstruction and controllability/topology analysis | platelet-specific proteins (human signaling network) | in silico node removal/control | critical, indispensable, intermittent, redundant nodes; MDS/MSS membership; control capacity | OmniPath database |
| drug target screening | TEP signaling subnetworks / module gene products | in silico pharmacological targeting | FDA-approved druggable targets | — |
- ▲ 111 genes upregulated in TEPs vs non-cancer platelets 111 genes
- ▼ 108 genes downregulated in TEPs vs non-cancer platelets 108 genes
- – Upregulated genes enriched in ECM–receptor interaction, focal adhesion, cytoskeleton, platelet activation; downregulated in immune signaling, apoptosis, ribosome
- – Platelet signaling network comprises 962 interactions among 401 proteins 962 interactions / 401 proteins
- – 86 critical nodes, 196 intermittent nodes, 119 redundant nodes identified 86 / 196 / 119
- – 62 indispensable nodes, 127 neutral nodes, 212 dispensable nodes identified 62 / 127 / 212
- – Chronic pancreatitis DEG profile most similar to NSCLC among inflammatory conditions
- – Indispensable nodes have a control capacity of zero, unlike critical nodes
- count 111 upregulated genes (DEGs upregulated in TEPs vs control platelets)
- count 108 downregulated genes (DEGs downregulated in TEPs vs control platelets)
- count 962 interactions among 401 proteins (reconstructed platelet signaling network from OmniPath)
- count 86 critical, 196 intermittent, 119 redundant nodes (node controllability classification)
- count 62 indispensable, 127 neutral, 212 dispensable nodes (node indispensability classification)
- other NES of at least 2 (threshold for top upregulated Reactome pathways in GSEA)
- other 60% healthy, 40% inflammatory conditions (composition of non-cancer donor controls)
- count 9 modules (k-means clustering of DEGs)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a computational/bioinformatics study that re-analyzed a publicly available GEO transcriptome dataset (GSE89843) comparing platelet gene expression between NSCLC patients and non-cancer donors. Differential gene expression was used to identify 219 DEGs, which were then subjected to k-means clustering, GSEA, and KEGG/GO/Reactome pathway enrichment. A platelet-specific signaling network was constructed from OmniPath-derived protein–protein interactions, and network controllability analysis (minimum driver node sets, indispensability classification) was used to prioritize therapeutic targets. No primary wet-lab experiments were performed; all statistics are derived from the secondary analysis of the existing dataset.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential gene expression analysis (specific algorithm not stated in the available text) | NSCLC platelets vs. non-cancer donor platelets (Figure 1A, Datasheet S1) | — | not stated |
| K-means clustering | Clustering of DEGs into 9 modules to separate NSCLC from inflammatory-condition samples (Figure 2A) | 219 DEGs | not stated |
| Gene Set Enrichment Analysis (GSEA) with NES ≥ 2 threshold | Top upregulated REACTOME pathways in TEP transcriptome | — | not stated |
| Over-representation / pathway enrichment analysis (KEGG, Gene Ontology BP/MF/CC, Reactome) | Upregulated and downregulated DEG sets (Figure 1C, Figure S1, Datasheets S2–S4) | 111 upregulated, 108 downregulated genes | not stated |
| Network controllability analysis (Minimum Driver Set / Minimum Steering Node Set classification; indispensability scoring) | Platelet signaling network of 401 nodes / 962 interactions; classification of critical (86), intermittent (196), redundant (119), indispensable (62), neutral (127), dispensable (212) nodes | 401 nodes, 962 interactions | na |
-
The DEG analysis method (test/package) is not specified in the available text↳ Could also: Standard packages such as DESeq2 (negative binomial Wald test), limma-voom, or edgeR could be used for RNA-seq count data, each with BH-FDR adjusted p-values and log2-fold-change reporting — Naming the specific method and its FDR threshold allows readers to assess the stringency of DEG calling and to reproduce the analysis; each package makes different distributional assumptions suited to different data structures
-
GSEA pathway significance was filtered by a hard NES ≥ 2 threshold↳ Could also: Reporting FDR-adjusted p-values (q-values) alongside NES, as is standard in GSEA output (e.g., q < 0.05), would be an additional or alternative criterion — NES alone does not account for sampling variability; an FDR threshold provides a probabilistic statement about false discovery that is interpretable across datasets of different sizes
-
K-means clustering was used to separate DEG modules, with k = 9 chosen↳ Could also: Consensus clustering, hierarchical clustering with dynamic tree-cutting (e.g., WGCNA), or silhouette-based k selection could also be used — K-means requires pre-specifying k and is sensitive to initialization; consensus or hierarchical approaches can provide data-driven cluster-number selection and stability estimates
-
Network importance was assessed via controllability (MDS/MSS framework)↳ Could also: Classical topological centrality metrics — betweenness centrality, eigenvector centrality, or PageRank — are widely used and computationally simpler alternatives for node prioritization — Controllability analysis captures dynamic reachability properties; centrality measures are more transparent, easier to benchmark, and allow direct comparison across published platelet network studies
-
The study uses a single publicly available dataset (GSE89843) without a held-out validation cohort↳ Could also: Cross-validation against an independent TEP or platelet transcriptome dataset (e.g., other GEO series for cancer platelets) could be used to assess reproducibility of the DEG set — A single-dataset analysis cannot distinguish dataset-specific from disease-specific signals; replication in an independent cohort strengthens confidence in the identified gene modules and drug targets
-
DEG results and module memberships are reported as categorical (up/down, module ID) without accompanying effect sizes or adjusted p-values in the main text↳ Could also: Volcano plots or MA-plots reporting log2 fold-change and -log10(adjusted p-value) for each gene, together with a fold-change threshold, are standard complements to DEG lists — Effect size (fold-change) combined with statistical significance allows readers to gauge biological magnitude alongside confidence; genes with large fold-changes but high variance are distinguished from those with modest but reliable changes
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41226816
Paper: Osmanoglu Ö, Özer E, Gupta SK, Heinze KG, Schulze H, Dandekar T. Network Controllability Reveals Key Mitigation Points for Tumor-Promoting Signaling in Tumor-Educated Platelets. Int J Mol Sci 2025. PMID 41226816 · PMCID PMC12609506 · DOI 10.3390/ijms262110780.
Code & data pointers (corrected)
- Real analysis repo: https://github.com/ozgeosmanoglu/Platelet_NSCLC
(commit
fd4ee1259ec98c5d44b1db6e0daf8cf908f0fbee, pushed 2024-11-08, public, no LICENSE file). Thecode_urlin the scaffold metadata (plotly/plotly.R) is a harvester false-positive and is WRONG — corrected here. - Data: GEO
GSE89843= ENA PRJNA353588 (Best et al. 2017 TEP RNA-seq); 779 samples (402 NSCLC / 377 non-cancer), Illumina HiSeq 2500, Kallisto-quantified. - Shipped artifact (key):
data/myDGEListFiltF(19 MB) — an RDGEListthat is the filtered + technical-replicate-collapsed + outlier-removed count matrix, i.e. the exact input to the RUVSeq + differential-expression step. This lets the DE result be reproduced without re-aligning the 826 SRA runs.
Pipeline map (scripts/)
| step | script | produces |
|---|---|---|
| 0 | Step0_PrepMetaData.R | phenodata (class = NSCLC vs Non-cancer) |
| 1 | Step1_ImportCounts.R | Kallisto→gene counts (tximport) |
| 2 | Step2_DataExploration.R | filter/collapse/outlier-remove → myDGEListFiltF; RUVSeq k=5 → W_1..W_5 |
| 3.1 | Step3.1_DGEanalysisDeseq2.R | DESeq2 LRT (design ~W_1..W_5+class, reduced ~W_1..W_5), fitType glmGamPoi → DEGs |
| 3.2 | Step3.2_DGEanalysisLimma.R | limma-voom (TMM, same RUV design) → DEGs |
| 7 | Step7_Interactions.R | OmnipathR PPI + platelet proteome filter → platelet network |
| 7.1.1 | Step7.1.1_BuildGraph_GeneScores.R | igraph network + SANTA gene scores |
| 7.1.2 | Step7.1.2_NetworkControllability.R | reads NetworkAnalysis.xlsx (classes) + drug repurposing |
| 7.1.3 | Step7.1.3_NetworkStats.R | control-centrality stats |
IN SCOPE (reproduce, clear 1:1)
R1 — Differentially expressed genes in TEP (NSCLC) vs non-cancer. Reported (Results/Abstract): 111 upregulated + 108 downregulated = 219 DEGs at |log2FC| > 0.58 and adjusted p < 0.05.
- Pipeline: RUVSeq k=5 (empirical genes, upper-quartile norm) → DESeq2 LRT
(glmGamPoi) and limma-voom (TMM), both on the shipped
myDGEListFiltF. - Fully deterministic from the shipped DGEList (no live downloads, no random seed in RUVg/DESeq/limma). This is the headline, low-hanging pipeline output.
OUT OF SCOPE (the hard ~20% — not attempted, with reasons)
- Network controllability node classes (62 indispensable / 127 neutral /
212 dispensable / 86 critical; central subnetwork 188 nodes, 501 edges).
Step7.1.2 does not compute these — it
read_excel(".../NetworkAnalysis.xlsx"), the output of the CytoCtrlAnalyzer Cytoscape GUI app. That xlsx is not shipped in the repo, and the GUI tool is not scriptable here. → not reproducible from shipped artifacts. (Audit note: these counts are not derivable from the deposited data/code alone.) - Platelet network size (401 proteins / 962 interactions). Step7_Interactions
needs (a) a live
OmnipathR::import_post_translational_interactions()snapshot (DB changes over time → not bit-reproducible) and (b) an unshipped proteomics fileplateletProteome(ackr3data).txt. → not reproducible exactly. - 5 key genes (ITGA2B, FLNA, GRB2, FCGR2A, APP), drug targets (fostamatinib …). Downstream of the above unshipped/GUI steps. → out of scope.
- Validation datasets (GSE183635/207586/68086), GSEA, functional enrichment — secondary, not the central claim. → not attempted (80/20).
Compute plan
All compute on «our HPC» (SLURM, account kubisch_std, partition std). Clone repo +
load myDGEListFiltF on «infra»; build conda R/Bioconductor env inside the compute
job; run RUV→DESeq2-LRT + limma. Output: up/down DEG counts → compare to 111/108.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The deposited, derivable result reproduces essentially 1:1: from the repo's shipped filtered DGEList the limma-voom + RUVSeq(k=5) path gives 217 DEGs (110 up/107 down) vs the reported 219 (111/108) — a ~0.9% boundary/version-drift difference — and the 402 NSCLC sample count matches exactly. A robustness caveat (our side, not authors'): the DEG count is strongly method-dependent (DESeq2-LRT yields 593), and the paper reports only the limma figure. The paper's actual central claim — network-controllability mitigation points, 5 key genes, drug repurposing — is not derivable from the deposited artifacts (un-shipped CytoCtrlAnalyzer GUI output, time-varying OmnipathR pull, missing proteomics file), so it is unverified rather than refuted. Net: a solid reproduction of the upstream DE with explainable minor deviation, but the headline novel conclusion remains independently unconfirmed — hence yellow on derivability/core-claim/overall, with no fabrication evidence.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.