Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Refining breast cancer biomarker discovery and drug targeting through an advanced data-driven approach.

BMC Bioinformatics · 2024
59/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
59/100
Reproducibility score
0.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 19% of all assessed papers rank 925 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for the DEG pipeline (the only code-free, in-scope part); NOT for the novel classifier (repo gone). Independent re-run on «our HPC» (SLURM «job», EXIT 0) reproduced the experimental design EXACTLY (merged 135 BC/44 normal; GSE45827 144/11; GSE10810 31/27; GSE42568 104/17) and the DEG counts to the same order of magnitude and direction with a faithful RMA->ComBat->limma pipeline (|logFC|>2 & p<0.05): merged 255 vs reported 164 (66up/189down vs 34/130); GSE45827 321 vs 350 (172up/149down vs 208/142) -> PARTIAL. The residual gap is dominated by the paper's under-specified 'unexpressed-probe removal' (paper ~10.6-11.7k genes vs our 20,824); a matched-gene-count sensitivity pulls merged toward the paper (208 vs 164; down 148 vs 130). Outputs are BIT-IDENTICAL to the prior run (sha256 deg_results.json 7ff0adc8...) -> fully deterministic, no fabrication signal: all numbers are plausible reproduction drift from documented under-specification, and sample-design matches to the individual sample. NOT attempted (out of scope): the BGWO_SA_Ens feature-selection + ensemble classifier and ALL downstream results (F1/AUC, 1404/1710 selected genes, 35 superior genes, named biomarkers, drug-target/enrichment/PPI) because the authors' GitHub repo returns 404 (repo_gone, re-verified 2026-06-22). All three GEO datasets (GSE45827/GSE10810/GSE42568, GPL570) are open, complete, and deliver exactly what the paper promises (N matches to the sample). Compute ran entirely on «our HPC» compute nodes (SLURM); a transient shared-account «infra» quota block was waited out, not worked around dishonestly.

💻 Code ↗ 🗄 Data: GSE45827

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 55
    assessed: 2026-06-16 ⛓ 4c43924beefd
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether an integrative machine learning approach—merging breast cancer gene expression datasets and applying a novel hybrid metaheuristic feature selection algorithm (BGWO_SA_Ens, combining binary grey wolf optimization, simulated annealing, and an ensemble classifier)—can more accurately identify breast cancer biomarkers and drug targets than existing methods.

Core claims
  • The BGWO_SA_Ens algorithm (hybrid BGWO + simulated annealing with an ensemble classifier objective function) selects predictive breast cancer biomarker genes with high classification performance method
  • Differential expression analysis of the merged dataset (GSE10810 + GSE42568) identified 164 differentially expressed genes between breast cancer and normal samples finding
  • Differential expression analysis of the separate GSE45827 dataset identified 350 differentially expressed genes finding
  • Intersecting DEGs with BGWO_SA_Ens-selected genes yields 35 'superior genes' consistently significant across both methods finding
  • The 35 superior genes are involved in AMPK, Adipocytokine, and PPAR signaling pathways mechanism
  • Protein-protein interaction network analysis of superior genes highlights subnetworks and central hub nodes finding
  • Drug-gene interaction analysis reveals connections between superior genes and anticancer drugs, informing precision oncology resource
  • This is the first study to apply a combination of hybrid metaheuristic algorithms and ensemble models to gene expression data for breast cancer biomarker discovery method
Experimental setups
Assay System Perturbation Readout Platform
Differential gene expression analysis (limma, RMA-normalized microarray) Merged human breast tissue dataset (GSE10810 + GSE42568) none (breast cancer vs normal tissue comparison) Differentially expressed genes (logFC>2, p<0.05) Affymetrix GPL570 (Human Genome U133 Plus 2.0 Array)
Differential gene expression analysis (limma, RMA-normalized microarray) GSE45827 human breast tissue dataset none (breast cancer vs normal tissue comparison) Differentially expressed genes (logFC>2, p<0.05) Affymetrix GPL570 (Human Genome U133 Plus 2.0 Array)
Feature selection via BGWO_SA_Ens (hybrid metaheuristic + ensemble classifier) Merged breast cancer gene expression dataset (179 samples, 10,629 features) computational feature selection (no biological perturbation) Selected gene subset and classification performance (F1, PR-AUC, ROC-AUC)
Feature selection via BGWO_SA_Ens (hybrid metaheuristic + ensemble classifier) GSE45827 gene expression dataset (155 samples) computational feature selection (no biological perturbation) Selected gene subset and classification performance (F1, PR-AUC, ROC-AUC)
Comparative feature selection (BGWO_Ens, GA_Ens, LASSO, MCFS_IFS, mRMR_IFS) Merged dataset and GSE45827 dataset computational feature selection Comparative classification/selection performance vs BGWO_SA_Ens
Gene Ontology (GO) and KEGG pathway enrichment analysis 35 superior genes (intersection of DEGs and BGWO_SA_Ens selections) none Enriched biological pathways/functions
Protein-protein interaction (PPI) network analysis 35 superior genes none Subnetworks and central/hub nodes
Drug-gene interaction analysis 35 superior genes none Associations between genes and anticancer drugs
Key results
  • Differential expression analysis identified 164 DEGs in the merged dataset
  • Differential expression analysis identified 350 DEGs in the GSE45827 dataset
  • BGWO_SA_Ens selected 1404 genes from >10,000 in the merged dataset with F1=0.981, PR-AUC=0.998, ROC-AUC=0.995 F1: 0.981
  • BGWO_SA_Ens selected 1710 genes in the GSE45827 dataset with F1=0.965, PR-AUC=0.986, ROC-AUC=0.972 F1: 0.965
  • Intersection of DEGs and BGWO_SA_Ens-selected genes yielded 35 superior genes consistently significant across methods
  • Predictive genes identified include TOP2A, AKR1C3, EZH2, MMP1, EDNRB, S100B, and SPP1
  • Superior genes are involved in AMPK, Adipocytokine, and PPAR signaling pathways per enrichment analysis
  • Drug-gene interaction analysis revealed connections between superior genes and anticancer drugs
Key statistics
  • fold_change logFC > 2 (DEG classification threshold criterion)
  • pvalue p < 0.05 (DEG significance threshold)
  • count 164 (DEGs identified in merged dataset)
  • count 350 (DEGs identified in GSE45827 dataset)
  • count 35 (superior genes at intersection of DEGs and BGWO_SA_Ens-selected genes)
  • other F1=0.981, PR-AUC=0.998, ROC-AUC=0.995 (BGWO_SA_Ens classification performance on merged dataset (1404 genes selected))
  • other F1=0.965, PR-AUC=0.986, ROC-AUC=0.972 (BGWO_SA_Ens classification performance on GSE45827 dataset (1710 genes selected))
  • count 179 BC samples / 44 normal samples (merged dataset composition (10,629 features) from GSE10810 (58 samples) and GSE42568 (121 samples))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This computational study merges two public microarray datasets (GSE10810 + GSE42568; n=223) and separately analyzes a third dataset (GSE45827; n=155) to identify breast cancer biomarkers. Differential gene expression (DGE) was performed with the limma package in R, using log-fold change >2 and p<0.05 as thresholds; batch effects were corrected with ComBat (empirical Bayes). A novel hybrid metaheuristic feature selection algorithm (BGWO_SA_Ens) was applied to both datasets, and classification performance was summarized using F1 score, PR-AUC, and ROC-AUC. Candidate biomarkers were further characterized via enrichment analysis, protein-protein interaction networks, and drug-gene interaction databases.

Replicationbiological Sample sizeSample sizes reported per dataset and per group (BC vs normal); no formal power analysis described GroupsBreast cancer samples vs normal samples (binary classification) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionnot stated
Statistical tests used
Test Applied to n Assumptions
limma moderated t-test (empirical Bayes linear model) Differential gene expression: breast cancer vs normal in merged dataset and GSE45827 223 (merged: 179 BC + 44 normal); 155 (GSE45827: 144 BC + 11 normal) not stated
ComBat empirical Bayes batch-effect correction Integration of GSE10810 and GSE42568 prior to DGE analysis 223 not stated
Principal Component Analysis (PCA) Visualization of batch-effect correction efficacy (prcomp in R) 223 na
Ensemble classifier with weighted voting (BGWO_SA_Ens objective function) Feature selection and sample classification in merged dataset and GSE45827 223 (merged); 155 (GSE45827) not stated
Approaches that could also have been used
  • The p-value threshold for DEGs is stated as <0.05 without specifying whether raw or FDR-adjusted values were used
    Could also: Explicitly report Benjamini-Hochberg adjusted p-values (q-values) at a standard FDR threshold (e.g., q<0.05 or q<0.1), which limma computes by default — With >10,000 genes tested simultaneously, applying a multiple-testing correction and reporting it explicitly is a standard step that clarifies the expected false discovery rate among reported DEGs
  • DEG thresholds were set at |logFC|>2 combined with p<0.05
    Could also: A |logFC|>1 (twofold change) threshold is also widely used and would capture a broader set of biologically relevant genes; alternatively, a volcano-plot-based approach with FDR<0.05 only (no logFC filter) is common — The logFC>2 cut-off (fourfold change) is stringent and may exclude moderately but consistently dysregulated genes; reporting sensitivity to threshold choice is informative
  • Machine learning classifier performance (F1, PR-AUC, ROC-AUC) is reported as point estimates without confidence intervals or cross-validation variance
    Could also: Report mean ± SD of performance metrics across repeated k-fold cross-validation folds, or bootstrap 95% confidence intervals for each metric — With moderately sized datasets (n=155–223), single point estimates of AUC/F1 can be sensitive to the specific split; variance estimates help characterize the stability of reported performance
  • Datasets from two sources (GSE10810, GSE42568) were merged with ComBat batch correction and then analyzed as one dataset
    Could also: A formal meta-analysis approach (e.g., combining effect sizes across studies using a random-effects model, as implemented in the metaMA or RankProd packages) would also integrate across studies while explicitly modeling inter-study heterogeneity — Meta-analysis preserves between-study variance as an estimable quantity and produces effect-size summaries with confidence intervals, which can complement or contextualize the merged-dataset approach
  • Classification model validation used a separate held-out GEO dataset (GSE45827) as an external test set
    Could also: Nested cross-validation (inner loop for feature selection / hyperparameter tuning, outer loop for performance estimation) on the merged dataset could also provide a nearly unbiased performance estimate without requiring an external cohort — External validation on GSE45827 is a strong design choice; nested CV would additionally allow quantifying variance in feature selection stability across splits within the merged data
  • Batch effect correction was evaluated visually via PCA of the first two principal components
    Could also: Quantitative measures such as the k-nearest-neighbor batch-effect test (kBET) or PVCA (principal variance component analysis) also assess residual batch structure after correction — Numerical batch-effect metrics complement visual PCA inspection and provide a reproducible, threshold-based confirmation that batch variance has been adequately reduced
Software: R/GEOquery · R/affy (RMA normalization) · R/limma · R/sva (ComBat) · R/ggbiplot · BGWO_SA_Ens (custom algorithm, this study)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
22
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.
Cited by (assessed papers) (1)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-38253993

Paper: Refining breast cancer biomarker discovery and drug targeting through an advanced data-driven approach. BMC Bioinformatics 2024. DOI 10.1186/s12859-024-05657-1. PMCID PMC10810249.

Code availability — REPO GONE (screened 2026-06-15)

Paper's Availability statement: "All codes ... available at https://github.com/mornejd/Refining-Breast-Cancer-Biomarker-Discovery-BGWO_SA_Ens".

  • GitHub page → HTTP 404.
  • GitHub REST API /repos/mornejd/Refining-...{"message":"Not Found","status":"404"}.
  • User mornejd exists but lists only ONE public repo (ICH_MHA, unrelated).
  • GitHub code/repo search for BGWO_SA_Ens and Refining Breast Cancer Biomarker Discovery → 0 results. => The authors' analysis code (the novel BGWO_SA_Ens feature-selection + ensemble classifier) is not publicly available. drop_reason candidate repo_gone for that part.

Data availability — PUBLIC (all GEO, GPL570)

  • GSE10810 (58: 31 BC / 27 normal), GSE42568 (121: 104 BC / 17 normal) → MERGED training set (135 BC / 44 normal (179 total); paper: 10,629 genes after preprocessing).
  • GSE45827 (155: 144 BC / 11 normal; paper: 11,731 genes) → validation set. All Affymetrix HG-U133 Plus 2.0 (GPL570). Raw CEL available as GEO supplementary.

What is IN SCOPE (reproducible from public data + described methods)

The differential-expression (DEG) pipeline is fully specified in Methods and needs no bespoke code:

  • GEOquery raw CEL → RMA (background correct + quantile normalize + summarize).
  • Probe→gene: drop unmapped probes, mean-collapse multi-probe genes.
  • Merge GSE10810+GSE42568, batch-correct with ComBat (sva).
  • limma DEG, BC vs normal; a gene is a DEG if |logFC| > 2 AND adj.p < 0.05.
  • Reported counts (Table 1): merged 164 DEGs (34 up / 130 down); GSE45827 350 DEGs (208 up / 142 down). These are the comparison targets.

What is OUT OF SCOPE / NOT ATTEMPTED (and why)

  • BGWO_SA_Ens feature selection + ensemble classifier (F1, PR-AUC, ROC-AUC, #selected genes 1404/1710, Fset choice) — the novel method lives only in the gone repo; it is a stochastic metaheuristic (Grey-Wolf + Simulated Annealing) whose exact result is not recoverable without the authors' code and seeds. Not attempted; recorded as repo_gone.
  • 35 "superior genes", named biomarkers (TOP2A, EZH2, ...), drug-target / enrichment / PPI downstream — all derive from the classifier output and/or external manual curation. Out of scope.
  • Exact post-filter gene counts (10,629 / 11,731): depend on under-specified "unexpressed-probe removal"; treated as secondary (reported, not graded hard).

Reproduction strategy (80/20)

Reproduce the clearly-specified limma DEG counts (the low-hanging, deterministic pipeline output) on the public CEL data; do NOT chase the stochastic classifier. Partial reproduction is the expected, valid outcome.

Figures / tables: TableTables
S1
Reported
merged training design 135 BC / 44 normal
Reproduced
135 BC / 44 normal (GSE10810 31/27 + GSE42568 104/17)
exact
S2
Reported
GSE45827 validation design 144 BC / 11 normal
Reproduced
144 BC / 11 normal (155/155 CELs matched)
exact
C1
Reported
merged DEGs total 164
Reproduced
255 (full 20824 genes); 208 (filtered to paper gene-count 10629)
partial
C2
Reported
merged upregulated 34
Reproduced
66 full; 60 filtered
partial
C3
Reported
merged downregulated 130
Reproduced
189 full; 148 filtered
partial
C4
Reported
GSE45827 DEGs total 350
Reproduced
321 full; 260 filtered
partial
C5
Reported
GSE45827 upregulated 208
Reproduced
172 full; 172 filtered
partial
C6
Reported
GSE45827 downregulated 142
Reproduced
149 full; 88 filtered
partial
C7
Reported
merged gene count after preprocessing 10629
Reproduced
20824 (no well-specified filter to reach 10629; 'unexpressed-probe removal' under-specified)
partial
C8
Reported
GSE45827 gene count after preprocessing 11731
Reproduced
20824
partial
X1
Reported
BGWO_SA_Ens classifier: F1 0.981 / ROC-AUC 0.995 / PR-AUC 0.998 / 1404 & 1710 selected genes / 35 superior genes / named biomarkers
Reproduced
OUT OF SCOPE - NOT ATTEMPTED (authors' code repo HTTP 404, repo_gone, re-verified 2026-06-22; stochastic Grey-Wolf+SA metaheuristic not recoverable without code/seeds)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

796.6 k
tokens (I/O) · 92.3 M incl. cache
293 min
runtime · 0.15 CPU-h
7.4 GB
peak RAM
3 (2 failed)
HPC jobs
hummel
machine