Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Network Controllability Reveals Key Mitigation Points for Tumor-Promoting Signaling in Tumor-Educated Platelets.

Int J Mol Sci · 2025
L1 84/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
84/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 63% of all assessed papers rank 392 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH + 1:1 on the headline. Real repo is ozgeosmanoglu/Platelet_NSCLC (the scaffold's plotly/plotly.R code_url was a harvester false-positive, corrected). The repo ships data/myDGEListFiltF (filtered/collapsed/outlier-removed DGEList, 14545x776, sha256 0a291a58...b7f05) = the exact DE input, so the differential-expression result reproduces WITHOUT re-aligning the 826 SRA runs. On «our HPC» (R 4.5.3, limma 3.66.0, RUVSeq 1.44.0, DESeq2 1.50.2) I reran the authors' RUVSeq(k=5)->limma-voom and ->DESeq2-LRT chain on that DGEList. limma-voom: 110 up / 107 down = 217 DEGs vs reported 111/108=219 -> off by 1 gene per direction (~0.9%), a clean within-tolerance match; NSCLC sample count 402 reproduced EXACTLY. Important honest finding: the repo's DESeq2-LRT path gives 593 DEGs, so the paper's reported figure is specifically the limma result (and the downstream network correctly uses limma logFC). NOT ATTEMPTED (declared out of scope, the hard ~20%): network-controllability node classes (62 indispensable/127 neutral/212 dispensable/86 critical), network size (401 nodes/962 edges), the 5 key genes, and drug repurposing -- Step7.1.2 only READS the classification from NetworkAnalysis.xlsx (output of the CytoCtrlAnalyzer Cytoscape GUI app, NOT shipped), and the network build requires a live/time-varying OmnipathR pull plus an unshipped proteomics file; these numbers are not derivable from the deposited artifacts and are flagged for the human reviewer, not asserted.

💻 Code ↗ 🗄 Data: GSE89843

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 84
    assessed: 2026-06-14 ⛓ 4e1ec9c8d2c8
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can integrative transcriptome and network controllability analysis of tumor-educated platelets (TEPs) in NSCLC identify key signaling nodes and FDA-approved drugs that therapeutically disrupt metastasis-promoting platelet–tumor signaling (e.g., ITAM, P2Y12)?

Core claims
  • TEPs in NSCLC show 111 upregulated and 108 downregulated genes versus non-cancer control platelets, enriched in ECM interaction, cytoskeleton, immune signaling, and platelet activation pathways. finding
  • Four complementary strategies identify five high-confidence central TEP signaling genes: ITGA2B, FLNA, GRB2, FCGR2A, and APP, all targetable by FDA-approved drugs. finding
  • Fostamatinib, an SYK inhibitor, is the top candidate drug to selectively disrupt ITAM-mediated platelet activation. finding
  • Network controllability analysis of a TEP-specific signaling network identifies critical and indispensable nodes as key mitigation points for tumor-promoting signaling. method
  • A low-dose combination therapy of fostamatinib, aducanumab, and acetylsalicylic acid (aspirin) may control TEP effects. finding
  • A TEP-specific signaling network was constructed by integrating GSE89843 transcriptome data with directed, signed protein–protein interactions from OmniPath. resource
  • The reconstructed platelet signaling network contains 962 interactions among 401 platelet-specific proteins, with 86 critical nodes and 62 indispensable nodes. finding
  • k-means clustering of DEGs yields nine modules; four upregulated modules represent protein homeostasis, platelet activation, metabolism, and cytoskeletal remodeling. finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq / transcriptome differential expression analysis platelets from NSCLC patients vs non-cancer donors (human) none (disease vs control comparison) differentially expressed genes (up/downregulated) GEO dataset GSE89843
pathway enrichment analysis (KEGG, GO Biological Process, Reactome) DEGs from human TEPs none enriched pathways/processes
k-means / hierarchical clustering of DEGs human platelet samples across NSCLC, healthy, and inflammatory conditions none gene expression modules (9 modules)
gene set enrichment analysis (GSEA) human TEP transcriptome none normalized enrichment score (NES) of Reactome pathways
protein–protein interaction network reconstruction and controllability/topology analysis platelet-specific proteins (human signaling network) in silico node removal/control critical, indispensable, intermittent, redundant nodes; MDS/MSS membership; control capacity OmniPath database
drug target screening TEP signaling subnetworks / module gene products in silico pharmacological targeting FDA-approved druggable targets
Key results
  • 111 genes upregulated in TEPs vs non-cancer platelets 111 genes
  • 108 genes downregulated in TEPs vs non-cancer platelets 108 genes
  • Upregulated genes enriched in ECM–receptor interaction, focal adhesion, cytoskeleton, platelet activation; downregulated in immune signaling, apoptosis, ribosome
  • Platelet signaling network comprises 962 interactions among 401 proteins 962 interactions / 401 proteins
  • 86 critical nodes, 196 intermittent nodes, 119 redundant nodes identified 86 / 196 / 119
  • 62 indispensable nodes, 127 neutral nodes, 212 dispensable nodes identified 62 / 127 / 212
  • Chronic pancreatitis DEG profile most similar to NSCLC among inflammatory conditions
  • Indispensable nodes have a control capacity of zero, unlike critical nodes
Key statistics
  • count 111 upregulated genes (DEGs upregulated in TEPs vs control platelets)
  • count 108 downregulated genes (DEGs downregulated in TEPs vs control platelets)
  • count 962 interactions among 401 proteins (reconstructed platelet signaling network from OmniPath)
  • count 86 critical, 196 intermittent, 119 redundant nodes (node controllability classification)
  • count 62 indispensable, 127 neutral, 212 dispensable nodes (node indispensability classification)
  • other NES of at least 2 (threshold for top upregulated Reactome pathways in GSEA)
  • other 60% healthy, 40% inflammatory conditions (composition of non-cancer donor controls)
  • count 9 modules (k-means clustering of DEGs)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational/bioinformatics study that re-analyzed a publicly available GEO transcriptome dataset (GSE89843) comparing platelet gene expression between NSCLC patients and non-cancer donors. Differential gene expression was used to identify 219 DEGs, which were then subjected to k-means clustering, GSEA, and KEGG/GO/Reactome pathway enrichment. A platelet-specific signaling network was constructed from OmniPath-derived protein–protein interactions, and network controllability analysis (minimum driver node sets, indispensability classification) was used to prioritize therapeutic targets. No primary wet-lab experiments were performed; all statistics are derived from the secondary analysis of the existing dataset.

Replicationbiological Sample sizeSample sizes not stated in the available text; 60% of non-cancer donors described as healthy, 40% with inflammatory conditions; exact per-group n not given GroupsNSCLC patient platelets vs. non-cancer donor platelets (healthy + inflammatory disease subgroups) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Confidence intervalsno Multiplicity correctionnot stated for DEG analysis; NES ≥ 2 threshold applied as inclusion criterion for GSEA pathways; no explicit FDR or family-wise error correction described in available text
Statistical tests used
Test Applied to n Assumptions
Differential gene expression analysis (specific algorithm not stated in the available text) NSCLC platelets vs. non-cancer donor platelets (Figure 1A, Datasheet S1) not stated
K-means clustering Clustering of DEGs into 9 modules to separate NSCLC from inflammatory-condition samples (Figure 2A) 219 DEGs not stated
Gene Set Enrichment Analysis (GSEA) with NES ≥ 2 threshold Top upregulated REACTOME pathways in TEP transcriptome not stated
Over-representation / pathway enrichment analysis (KEGG, Gene Ontology BP/MF/CC, Reactome) Upregulated and downregulated DEG sets (Figure 1C, Figure S1, Datasheets S2–S4) 111 upregulated, 108 downregulated genes not stated
Network controllability analysis (Minimum Driver Set / Minimum Steering Node Set classification; indispensability scoring) Platelet signaling network of 401 nodes / 962 interactions; classification of critical (86), intermittent (196), redundant (119), indispensable (62), neutral (127), dispensable (212) nodes 401 nodes, 962 interactions na
Approaches that could also have been used
  • The DEG analysis method (test/package) is not specified in the available text
    Could also: Standard packages such as DESeq2 (negative binomial Wald test), limma-voom, or edgeR could be used for RNA-seq count data, each with BH-FDR adjusted p-values and log2-fold-change reporting — Naming the specific method and its FDR threshold allows readers to assess the stringency of DEG calling and to reproduce the analysis; each package makes different distributional assumptions suited to different data structures
  • GSEA pathway significance was filtered by a hard NES ≥ 2 threshold
    Could also: Reporting FDR-adjusted p-values (q-values) alongside NES, as is standard in GSEA output (e.g., q < 0.05), would be an additional or alternative criterion — NES alone does not account for sampling variability; an FDR threshold provides a probabilistic statement about false discovery that is interpretable across datasets of different sizes
  • K-means clustering was used to separate DEG modules, with k = 9 chosen
    Could also: Consensus clustering, hierarchical clustering with dynamic tree-cutting (e.g., WGCNA), or silhouette-based k selection could also be used — K-means requires pre-specifying k and is sensitive to initialization; consensus or hierarchical approaches can provide data-driven cluster-number selection and stability estimates
  • Network importance was assessed via controllability (MDS/MSS framework)
    Could also: Classical topological centrality metrics — betweenness centrality, eigenvector centrality, or PageRank — are widely used and computationally simpler alternatives for node prioritization — Controllability analysis captures dynamic reachability properties; centrality measures are more transparent, easier to benchmark, and allow direct comparison across published platelet network studies
  • The study uses a single publicly available dataset (GSE89843) without a held-out validation cohort
    Could also: Cross-validation against an independent TEP or platelet transcriptome dataset (e.g., other GEO series for cancer platelets) could be used to assess reproducibility of the DEG set — A single-dataset analysis cannot distinguish dataset-specific from disease-specific signals; replication in an independent cohort strengthens confidence in the identified gene modules and drug targets
  • DEG results and module memberships are reported as categorical (up/down, module ID) without accompanying effect sizes or adjusted p-values in the main text
    Could also: Volcano plots or MA-plots reporting log2 fold-change and -log10(adjusted p-value) for each gene, together with a fold-change threshold, are standard complements to DEG lists — Effect size (fold-change) combined with statistical significance allows readers to gauge biological magnitude alongside confidence; genes with large fold-changes but high variance are distinguished from those with modest but reliable changes
Software: OmniPath (database for directed protein–protein interactions) · NCBI GEO (dataset GSE89843) · DEG/clustering/GSEA pipeline (specific package not named in available text)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE183635 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE89843 GEO in Introduction (http://purl.org/orb/Introduction)
no other assessed paper uses this yet
PRJNA353588 BioProject in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
R-HSA-1474290 Reactome in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
R-HSA-1592389 Reactome in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
R-HSA-3000178 Reactome in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
R-HSA-381426 Reactome in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
R-HSA-445095 Reactome in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
R-HSA-446728 Reactome in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
R-HSA-5173105 Reactome in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41226816

Paper: Osmanoglu Ö, Özer E, Gupta SK, Heinze KG, Schulze H, Dandekar T. Network Controllability Reveals Key Mitigation Points for Tumor-Promoting Signaling in Tumor-Educated Platelets. Int J Mol Sci 2025. PMID 41226816 · PMCID PMC12609506 · DOI 10.3390/ijms262110780.

Code & data pointers (corrected)

  • Real analysis repo: https://github.com/ozgeosmanoglu/Platelet_NSCLC (commit fd4ee1259ec98c5d44b1db6e0daf8cf908f0fbee, pushed 2024-11-08, public, no LICENSE file). The code_url in the scaffold metadata (plotly/plotly.R) is a harvester false-positive and is WRONG — corrected here.
  • Data: GEO GSE89843 = ENA PRJNA353588 (Best et al. 2017 TEP RNA-seq); 779 samples (402 NSCLC / 377 non-cancer), Illumina HiSeq 2500, Kallisto-quantified.
  • Shipped artifact (key): data/myDGEListFiltF (19 MB) — an R DGEList that is the filtered + technical-replicate-collapsed + outlier-removed count matrix, i.e. the exact input to the RUVSeq + differential-expression step. This lets the DE result be reproduced without re-aligning the 826 SRA runs.

Pipeline map (scripts/)

step script produces
0 Step0_PrepMetaData.R phenodata (class = NSCLC vs Non-cancer)
1 Step1_ImportCounts.R Kallisto→gene counts (tximport)
2 Step2_DataExploration.R filter/collapse/outlier-remove → myDGEListFiltF; RUVSeq k=5 → W_1..W_5
3.1 Step3.1_DGEanalysisDeseq2.R DESeq2 LRT (design ~W_1..W_5+class, reduced ~W_1..W_5), fitType glmGamPoi → DEGs
3.2 Step3.2_DGEanalysisLimma.R limma-voom (TMM, same RUV design) → DEGs
7 Step7_Interactions.R OmnipathR PPI + platelet proteome filter → platelet network
7.1.1 Step7.1.1_BuildGraph_GeneScores.R igraph network + SANTA gene scores
7.1.2 Step7.1.2_NetworkControllability.R reads NetworkAnalysis.xlsx (classes) + drug repurposing
7.1.3 Step7.1.3_NetworkStats.R control-centrality stats

IN SCOPE (reproduce, clear 1:1)

R1 — Differentially expressed genes in TEP (NSCLC) vs non-cancer. Reported (Results/Abstract): 111 upregulated + 108 downregulated = 219 DEGs at |log2FC| > 0.58 and adjusted p < 0.05.

  • Pipeline: RUVSeq k=5 (empirical genes, upper-quartile norm) → DESeq2 LRT (glmGamPoi) and limma-voom (TMM), both on the shipped myDGEListFiltF.
  • Fully deterministic from the shipped DGEList (no live downloads, no random seed in RUVg/DESeq/limma). This is the headline, low-hanging pipeline output.

OUT OF SCOPE (the hard ~20% — not attempted, with reasons)

  • Network controllability node classes (62 indispensable / 127 neutral / 212 dispensable / 86 critical; central subnetwork 188 nodes, 501 edges). Step7.1.2 does not compute these — it read_excel(".../NetworkAnalysis.xlsx"), the output of the CytoCtrlAnalyzer Cytoscape GUI app. That xlsx is not shipped in the repo, and the GUI tool is not scriptable here. → not reproducible from shipped artifacts. (Audit note: these counts are not derivable from the deposited data/code alone.)
  • Platelet network size (401 proteins / 962 interactions). Step7_Interactions needs (a) a live OmnipathR::import_post_translational_interactions() snapshot (DB changes over time → not bit-reproducible) and (b) an unshipped proteomics file plateletProteome(ackr3data).txt. → not reproducible exactly.
  • 5 key genes (ITGA2B, FLNA, GRB2, FCGR2A, APP), drug targets (fostamatinib …). Downstream of the above unshipped/GUI steps. → out of scope.
  • Validation datasets (GSE183635/207586/68086), GSEA, functional enrichment — secondary, not the central claim. → not attempted (80/20).

Compute plan

All compute on «our HPC» (SLURM, account kubisch_std, partition std). Clone repo + load myDGEListFiltF on «infra»; build conda R/Bioconductor env inside the compute job; run RUV→DESeq2-LRT + limma. Output: up/down DEG counts → compare to 111/108.

C1
Reported
219 DEGs (111 up + 108 down), |log2FC|>0.58 & padj<0.05
Reproduced
217 (110 up + 107 down) via limma-voom + RUVSeq k=5
within tolerance
C2
Reported
111 upregulated
Reproduced
110
within tolerance
C3
Reported
108 downregulated
Reproduced
107
within tolerance
C4
Reported
402 NSCLC samples
Reproduced
402
exact
C5
Reported
thresholds |log2FC|>0.58 & adj.p<0.05
Reproduced
identical
exact
C6
Reported
DEG count from limma-voom (network uses limma logFC)
Reproduced
limma matches (217); DESeq2-LRT gives 593 (382/211)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 84/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

The deposited, derivable result reproduces essentially 1:1: from the repo's shipped filtered DGEList the limma-voom + RUVSeq(k=5) path gives 217 DEGs (110 up/107 down) vs the reported 219 (111/108) — a ~0.9% boundary/version-drift difference — and the 402 NSCLC sample count matches exactly. A robustness caveat (our side, not authors'): the DEG count is strongly method-dependent (DESeq2-LRT yields 593), and the paper reports only the limma figure. The paper's actual central claim — network-controllability mitigation points, 5 key genes, drug repurposing — is not derivable from the deposited artifacts (un-shipped CytoCtrlAnalyzer GUI output, time-varying OmnipathR pull, missing proteomics file), so it is unverified rather than refuted. Net: a solid reproduction of the upstream DE with explainable minor deviation, but the headline novel conclusion remains independently unconfirmed — hence yellow on derivability/core-claim/overall, with no fabrication evidence.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

129.9 k
tokens (I/O) · 12.5 M incl. cache
14 min
runtime · 0.06 CPU-h
5 GB
peak RAM
1
HPC jobs
hummel
machine