Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Advances in genomic and pharmacokinetic profiling for clinical stratification of metastatic breast cancer.

Discov Oncol · 2025
L1 59/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
59/100
Reproducibility score
0.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 19% of all assessed papers rank 925 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Multi-phase in-silico breast-cancer drug-discovery paper. The registry 'code' repo (ZaFoniX) is a Tkinter GUI drug-lookup tool needing an access key + proprietary DrugsList.xlsx (not shipped) -- NOT the analysis pipeline; most phases (WGCNA, TCGA mutation/VEP, 3D modelling, docking, ADMET) are web-server/GUI/manual and out of scope. The one clearly-specified pipeline output is Phase-1 DEG identification (GEOquery+limma, stated thresholds, Table 2). Reproduced on «our HPC» for the assigned accession GSE46141. RESULT: probe count is essentially exact (51562 vs 51563, off by 1) -> dataset identity confirmed. DEG counts are only PARTIAL and NOT 1:1: the paper never states the contrast and GSE46141 is 91 metastasis-only FNA samples with no tumour-vs-normal split, so 1295/905/390 are not derivable as printed. Under a reconstructed liver-vs-other contrast at the paper's thresholds we got 772/533/239; the up/down ratio matches well (2.23 vs 2.32) and the achievable DEG range (772-1567) brackets 1295, but an exact reproduction is impossible without the undocumented contrast -> flagged possible-fabrication/underspecification for human review. NOT attempted (per 80/20): WGCNA modules, hub genes, TCGA variants, structure/docking/ADMET, and the other 7 GEO datasets in Table 2.

💻 Code ↗ 🗄 Data: GSE46141

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 59
    assessed: 2026-06-14 ⛓ e5de43f822e9
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Integrative transcriptomic, network, structural, and pharmaco-informatic analysis of metastatic breast cancer can identify key molecular driver genes (e.g., ROR1, ROR2, RPS6, SMAD3, UBC, AKT1, AR, CDH1) and prioritize high-affinity candidate inhibitors against them via computer-aided drug design.

Core claims
  • Eight gene modules linked to metastasis were identified via scored network analysis and validated through pathway databases. finding
  • Key genes AR, AKT1, UBC, CDH1, SMAD3, ROR1, and ROR2 are associated with chemotherapy resistance and poor prognosis in metastatic breast cancer. finding
  • Kinases AKT1, ROR1, ROR2 and non-kinase targets UBC, RPS6, CDH1, AR, SMAD3 are the most promising therapeutic candidates. finding
  • Ellagic Acid and Erioflorin stand out as potent candidate compounds against critical metastatic breast cancer targets, with promising pharmacokinetics and safety. resource
  • A multi-stage bioinformatics and computer-aided drug design pipeline (DEG identification, network analysis, structural modeling, virtual screening, docking, ADMET, NMA dynamics) was applied to MBC. method
  • ROR1/ROR2 act through Wnt and MAPK (p38) signaling to drive proliferation, migration, and metastasis, making them targetable. mechanism
  • Differentially expressed genes were identified across primary breast tumors and metastatic organs (lung, liver, bone, brain) after filtering redundant DEG entries. finding
Experimental setups
Assay System Perturbation Readout Platform
Gene expression profiling (microarray) Human metastatic breast cancer samples from GEO datasets (Affymetrix, Illumina HumanHT-12 V4.0, HumArray3.2 platforms) none Differentially expressed genes (DEGs) Affymetrix / Illumina HumanHT-12 V4.0 Expression BeadChip / HumArray3.2
Bulk RNA-Seq / expression profiling TCGA-BRCA Pan-Cancer Atlas primary breast invasive carcinoma (1084 samples; basal-like 171, normal-like 36) none Gene expression patterns in primary tumors and subtypes
RT-PCR / expression profiling by array Homo sapiens metastatic breast tissue (bone, liver, lung, lymph node, brain) none Gene expression / DEGs GEOquery (R)
Weighted gene co-expression network analysis (WGCNA) GSE40622 dataset (275 samples) none Co-expression gene modules TSUNAMI
Gene regulatory network construction / gene set enrichment Curated gene list of 2344 entries none Co-regulated gene clusters / GRNs GenCLiP 2.0/3.0
Differential expression analysis (limma) GEO transcriptomic datasets none DEGs by p-value, FDR, log2 fold change limma (Bioconductor, R)
In silico protein structure modeling, virtual screening and molecular docking Key MBC target proteins (AKT1, ROR1, ROR2, UBC, RPS6, CDH1, AR, SMAD3) drug/ligand (compound library screening) Protein-ligand binding affinity / interactions trRosetta, MODELLER, I-TASSER, PyRx, AutoDock Vina, MOE
ADMET profiling and normal mode analysis (NMA/MD) Candidate ligand-target complexes drug/ligand Pharmacokinetics, toxicity, structural dynamics/stability SwissADME, ProTox 3.0, PaDEL-Descriptor, iMODS
Key results
  • Eight gene modules linked to metastasis identified via scored network analysis and validated through pathway databases. 8 modules
  • Key genes AR, AKT1, UBC, CDH1, SMAD3, ROR1, ROR2 associated with chemotherapy resistance and poor prognosis.
  • Ellagic Acid and Erioflorin identified as the standout potent candidate compounds against critical MBC targets.
  • Significantly altered genes identified across primary breast tumors and metastatic organs (lung, liver, bone, brain) after filtering redundant DEGs.
  • All screened compounds showed varying strong interacting profiles with the target proteins.
Key statistics
  • count 9 datasets (GSE46141, GSE100534, GSE27567, GSE65517, GSE23988, GSE27447, GSE32394, GSE22093, GSE40622) (GEO transcriptomic datasets acquired)
  • count 1084 (TCGA-BRCA primary breast invasive carcinoma samples)
  • count 171 (basal-like subtype samples in TCGA)
  • count 36 (normal-like subtype samples in TCGA)
  • count 275 (samples in GSE40622 used for WGCNA)
  • count 2344 (curated gene list entries for GRN construction in GenCLiP)
  • pvalue p-value ≤ 0.05, FDR < 0.05, |log2FC| > 1 (DEG selection thresholds (Benjamini–Hochberg FDR correction))
  • count 8 (gene modules linked to metastasis)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This computational study integrated nine GEO gene expression datasets (microarray, RT-PCR, bulk RNA-seq) and TCGA-BRCA (1084 samples) to identify differentially expressed genes in metastatic breast cancer using R/limma with Benjamini-Hochberg FDR correction; DEG thresholds were |log2FC| > 1 and FDR < 0.05. Co-expression modules were derived via WGCNA (GSE40622, n=275), and gene regulatory networks were built with GenCLiP fuzzy c-means clustering on 2344 curated genes. Hub genes were validated through survival and prognostic web tools (KM Plotter, GEPIA2, ROC Plotter), and candidate compounds against eight key targets were evaluated by AutoDock Vina virtual screening, ADMET profiling, and normal mode analysis via iMODS.

Replicationunclear Sample sizeNine GEO datasets integrated; TCGA-BRCA n=1084 primary tumors (basal-like n=171, normal-like n=36); GSE40622 n=275 for WGCNA; individual GEO dataset sizes not fully reported in available text; no formal power calculation described GroupsPrimary breast tumors vs. metastatic samples (lung, liver, bone, brain, lymph node); TCGA molecular subtypes (basal-like vs. normal-like); target proteins vs. screened compound libraries Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
limma moderated t-test (empirical-Bayes linear model) DEG identification in each of nine GEO datasets and TCGA-BRCA comparisons TCGA-BRCA: 1084 primary samples (basal-like n=171, normal-like n=36); GSE40622: 275 samples; individual GEO dataset sizes not stated in available text not stated
Benjamini-Hochberg FDR correction Multiple-testing correction applied to all DEG comparisons na
WGCNA soft-thresholding power selection and hierarchical clustering with dynamic tree cut Co-expression module identification (GSE40622) 275 samples not stated
Fuzzy c-means clustering (GenCLiP membership/centroid algorithm) Gene regulatory network cluster construction from curated gene list 2344 genes not stated
AutoDock Vina binding affinity scoring (kcal/mol) Virtual screening and molecular docking of compound libraries against key target proteins na
Survival analysis via KM Plotter (underlying test not explicitly named; likely log-rank) Prognostic validation of hub genes in breast cancer patient cohorts not stated
Approaches that could also have been used
  • limma (developed for microarray continuous intensities with moderated t-statistics) was applied uniformly across all platforms including bulk RNA-seq datasets
    Could also: DESeq2 or edgeR (negative-binomial models for integer read counts) could also be applied to the RNA-seq datasets; limma-voom, which applies a mean-variance trend weight before limma, is another widely used bridge between the two paradigms — Count-specific models explicitly account for discrete overdispersion in RNA-seq data; the choice of method can affect DEG lists particularly for lowly expressed genes, so platform-matched modeling is a common alternative practice in multi-platform studies
  • Nine GEO datasets spanning different platforms (Affymetrix, Illumina BeadChip, HumArray) were analyzed separately and results merged by comparing gene lists
    Could also: A formal meta-analysis framework — such as the R packages MetaDE, RankProd, or a random-effects model pooling per-gene effect estimates — could also be used to integrate findings across datasets — Formal meta-analysis yields a unified weighted effect estimate with confidence intervals and quantifies between-study heterogeneity (I²), which is especially informative when datasets differ in platform, sample size, and clinical context
  • Multiplicity correction (BH-FDR) was applied to DEG identification, but no correction was described for enrichment analyses run in parallel across multiple databases (KEGG, GO, Reactome, DAVID, ShinyGO, GenCLiP)
    Could also: Applying BH-FDR within each enrichment tool's output and reporting adjusted p-values for pathway terms could also be done; limiting enrichment testing to one or two pre-specified databases is another common approach — Testing many pathway terms across multiple tools simultaneously increases the expected number of false-positive enrichments; reporting tool-level adjusted p-values makes the pathway prioritization more reproducible and comparable to other studies
  • Protein structural dynamics were assessed using normal mode analysis (NMA) via iMODS
    Could also: Classical all-atom molecular dynamics (MD) simulation in explicit solvent (e.g., GROMACS, AMBER, NAMD) with MM-PBSA/GBSA binding free energy estimation could also be used — Full MD captures conformational sampling at physiological temperature over nanosecond-to-microsecond timescales and provides quantitative binding free energy estimates; NMA is computationally faster and well-suited to large-scale screening, while MD offers deeper mechanistic resolution for prioritized candidates
  • Prognostic validation was performed through web portal tools (KM Plotter, GEPIA2, ROC Plotter) without explicitly reporting the underlying test statistics, hazard ratios, or patient cohort sizes
    Could also: Explicitly reporting the log-rank test statistic, p-value, hazard ratio, and 95% confidence interval from a Cox proportional hazards model, along with the number of patients and events per group, could also be done — Hazard ratios with confidence intervals quantify the magnitude and precision of the prognostic association and allow cross-study comparisons; event counts are necessary to assess statistical power and interpret the reliability of the survival estimates
  • Candidate compounds were ranked primarily by single-program AutoDock Vina binding affinity scores
    Could also: Consensus scoring across two or more independent docking programs (e.g., Glide, GOLD, AutoDock Vina) or rescoring with a machine-learning scoring function could also be used to rank candidates — Single-program docking scores reflect that program's specific force field and scoring function assumptions; consensus scoring across programs has been shown to improve enrichment of true actives in prospective virtual screening by reducing program-specific bias
Software: R/limma (Bioconductor) · R/GEOquery (Bioconductor) · R/UMAP · TSUNAMI (WGCNA wrapper) · GenCLiP 2.0 / GenCLiP 3.0 2.0 and 3.0 · AutoDock Vina (managed via PyRx) · iMODS (normal mode analysis) · GEPIA2 · KM Plotter · ROC Plotter · SwissADME / ProTox 3.0 / PaDEL-Descriptor ProTox version 3.0; others not stated · trRosetta / MODELLER / I-TASSER (structure prediction)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
Built on 2 assessed reference(s) · mean reproducibility 94/100
stands on reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (2)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

1H10 PDBe in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
2AX9 PDBe in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
3CQU PDBe in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
3O96 PDBe in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
3Q2V PDBe in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
4EJN PDBe in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
5B83 PDBe in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
5KCV PDBe in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
7APJ PDBe in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0070848 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GO:0071363 Gene Ontology (GO) in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
P0CG48 UniProt in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
P10275 UniProt in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
P12830 UniProt in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
P31749 UniProt in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
P62753 UniProt in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
P84022 UniProt in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
Q01973 UniProt in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
Q01974 UniProt in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41369820

Paper: Attique Z, Azhar HMF, Khan S. Advances in genomic and pharmacokinetic profiling for clinical stratification of metastatic breast cancer. Discov Oncol 2025. DOI 10.1007/s12672-025-04203-6. PMID 41369820 / PMC12799884.

"Code" repo (per registry): https://github.com/ZarlishAttique/ZaFoniX — a Tkinter GUI drug-lookup tool (drug_finder.py) that needs an access key and a proprietary DrugsList.xlsx (neither shipped). It is not the analysis pipeline behind the paper's quantitative results; it corresponds only to the "ZaFoniX" ligand-source mention in Phase 4. → not a reproducible analysis pipeline.

This is a large, multi-phase in-silico drug-discovery study. Most phases use interactive web servers / GUI desktop tools / manual curation with no shipped script, parameters file, or seed → not reproducible as a pipeline.

In scope (pipeline-derived, reproducible)

Result Pipeline Reproducible?
Phase 1 — DEG identification per GEO dataset (Table 2) GEOquery + limma in R; thresholds p ≤ 0.05, FDR < 0.05, |log2FC| > 1 YES (attempted) — clearly specified tools + thresholds + a numeric table to compare
Probe/feature count per dataset (Table 2 "Probes") GEO series-matrix dimensions YES — directly checkable, contrast-independent

We focus on GSE46141 (this RU's assigned accession; also row 1 of Table 2: 51,563 probes → 1,295 sig genes, 905 up / 390 down).

Known limitation surfaced during scoping

GSE46141 = 91 fine-needle-aspirate samples of breast-cancer metastases from different anatomical sites (GPL10379). There is no tumour-vs-normal contrast and the paper does not state which two groups were compared for its DEG counts. The reported 1,295/905/390 are therefore not derivable without an undocumented contrast choice — recorded as a possible-fabrication / underspecification note. We attempt the most defensible reconstructed contrast (liver metastasis vs other sites, matching the dataset's stated focus) to test whether the reported counts are even in the achievable range.

Out of scope (not a runnable pipeline / manual / external)

  • WGCNA "64 modules" (TSUNAMI web app, GSE40622) — web GUI, no params seed.
  • Hub-gene selection (8 genes) — Cytoscape/MCODE/Metascape GUI, no script.
  • TCGA mutation analysis (VEP/SIFT/PolyPhen/Pfam, per-variant calls) — manual, portal-driven; specific variants not script-derived.
  • 3D structure modelling & validation (trRosetta/MODELLER/I-TASSER/MolProbity, Table 5 GDT-HA/MolProbity) — web servers, stochastic, manual.
  • Molecular docking / virtual screening (PyRx/AutoDock Vina/MOE, binding energies) — GUI, manual target prep; not a shipped, parameterised workflow.
  • ADMET / pharmacophore / iMODS NMA — SwissADME/ProTox/LigandScout web servers.
  • ZaFoniX GUI tool — needs access key + proprietary DrugsList.xlsx (not shipped).

Approach

Download GSE46141 on «infra» inside a «our HPC» SLURM job, build a conda R env (bioconductor-geoquery, bioconductor-limma), report the series-matrix dimensions (probe count), and run a limma DE under a documented reconstructed contrast at the paper's thresholds. Compare to Table 2; grade provisionally; a human reviewer decides.

Figures / tables: Table
C1
Reported
51563 probes (Table 2, GSE46141)
Reproduced
51562 features in GEO series matrix
within tolerance
C2
Reported
1295 significant DEGs (Table 2, GSE46141; p<=0.05 & FDR<0.05 & |log2FC|>1)
Reproduced
772 (reconstructed liver-vs-other contrast; achievable range 772-1567 brackets 1295)
partial
C3
Reported
905 up (Table 2, GSE46141)
Reproduced
533 up (up/down ratio 2.23 vs reported 2.32)
partial
C4
Reported
390 down (Table 2, GSE46141)
Reproduced
239 down
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 59/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴

Dataset identity is confirmed: the GSE46141 probe count reproduces almost exactly (51562 vs 51563). The DEG counts (1295/905/390) are not derivable as printed — the paper never states the contrast and the dataset is 91 metastasis-only FNA samples with no tumour-vs-normal split — so this is an authors'-side underspecification defect, not a reproduction error. Under the most defensible reconstructed contrast we get 772/533/239 with a matching up/down ratio (2.23 vs 2.32) and an achievable range (772–1567) that brackets 1295, so the figures are plausibly in-range but not independently verifiable. Severity is moderate (direction/structure hold) and the qualitative DEG claim survives, but the exact values fail derivability → flagged possible-fabrication for human review.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

98.1 k
tokens (I/O) · 6.1 M incl. cache
10 min
runtime · 0.01 CPU-h
1.9 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine