Machine learning and free energy clustering reveal PAH protein binding linked to AD risk.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH? Partially. The repo (github.com/pdpmb/PAHs @1ac00d6) ships only loose R scripts with hardcoded «path» paths; the intermediate matrices they load() (expr_last1_bactch.Rdata, dat_factor*.rda) and the 145-gene panel (network_145_genes.csv) are NOT shipped and NO script regenerates them, so the exact pipeline is not runnable end-to-end. Per the brief (80/20, third-party/own-reimplementation equally valid, honest 1:1), I reconstructed the documented pipeline on the paper's OWN GEO data on «our HPC»: GEOquery -> common-symbol merge of GSE44770/44771/33000 (19,480 genes, 927 AD/CTL samples) -> sva::ComBat batch correction -> limma AD-vs-CTL with the paper's EXACT thresholds (verbatim from shipped 'limma' script) -> XGBoost with the shipped hyper-parameters. RESULT vs paper: (C2) XGBoost AUROC 0.993 reproduces the headline '>0.95' claim (within-tol/concordant). (C1) limma gives 6816 DEGs vs reported 2847 - same direction trend (down>up), biologically correct top hits (CRH/DUSP4/VGF/BDNF), but ~2.4x more, attributable to the unshipped preprocessing (gene MAD-filtering / sample-subsetting / exact batch method) that builds the authors' matrix. (C3) sample composition is exact ground truth. NOT ATTEMPTED (out of scope / last-20%): WGCNA 1164 hub genes & 145-gene PPI network (manual module choice + unshipped gene list + external PAH-target predictor), molecular docking dG (structural/manual), GO/KEGG (qualitative only), BMDL model (toxicology), and the GSE15222/GSE118553 external ML validation (secondary). No evidence of fabrication - numbers are plausible and the pipeline behaves correctly; the gap is a reproducibility limitation (missing preprocessing artifacts), not a contradiction. 1:1 vs different: a faithful re-implementation (not the authors' exact code, which is not runnable) confirming the ML claim and partially confirming the DEG claim.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 78assessed: 2026-06-14 ⛓ 753f04bd9757
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusWhether the average molecular docking binding free energy (ΔG) across multiple clustered core protein targets correlates with quantifiable PAH neurotoxicity parameters (BMDL/LD50) and can thereby prioritize polycyclic aromatic hydrocarbons for Alzheimer's disease-associated neurotoxicity.
- ★ An integrated framework of bioinformatics, machine learning, and ΔG clustering can prioritize PAHs for AD-associated neurotoxicity. method
- ★ PARP1, PTPN1, and ITGA4 are candidate core PAH protein targets enriched in neuroinflammation, microglial activation, lipid metabolism, and atherosclerosis pathways. finding
- ★ Category-average docking ΔG values correlate linearly with literature LD50/BMDL toxicity data, yielding BMDL = 1.723 × ΔG + 22.602. finding
- ★ Multi-target average docking ΔG from clustering better reflects overall PAH toxicity than single-target docking scores. mechanism
- Zebrafish motility assays provide preliminary experimental support for the PAH neurotoxicity predictions. finding
- PARP1, PTPN1, and ITGA4 exhibit lower Gibbs free energies with most of the 16 PAHs, indicating higher probability of spontaneous binding. mechanism
- The pipeline provides a hypothesis-generating resource for predicting PAH neurotoxic potency in AD contexts pending experimental validation. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Target prediction (ChEMBL/STITCH database mining) | human proteins / PAH compounds | none | predicted PAH target proteins | ChEMBL, STITCH |
| Differential expression analysis | human AD patient brain microarray (GEO datasets GSE44771, GSE44770, GSE33000, GSE15222, GSE118553) | AD vs control | differentially expressed genes | GEO microarray |
| WGCNA (weighted gene co-expression network analysis) | human AD microarray datasets | AD vs control | AD-associated hub gene modules | — |
| PPI network / GO-KEGG enrichment | AD-associated DEGs and predicted PAH targets | none | network topology and enriched pathways | STRING, Cytoscape |
| XGBoost feature selection | AD microarray gene expression datasets | none | feature importance / AUROC for disease feature genes | XGBoost |
| Molecular docking | 16 PAHs and 8 core target proteins | ligand binding | binding free energy (ΔG) | AutoDock; PyMOL visualization |
| Molecular dynamics simulation / free energy landscape | PARP1 protein bound to 4 PAHs | PAH binding | RMSD and radius of gyration | — |
| Zebrafish motility/swimming assay | zebrafish | PAH exposure (naphthalene 1342 μM, anthracene 45.0 μM, pyrene 13.7 μM) | swimming distance / trajectory | — |
- ▲ Category-average ΔG correlated linearly with literature LD50/BMDL data ρ=1
- – Empirical relationship derived between BMDL and ΔG BMDL = 1.723 × ΔG + 22.602
- ▼ Zebrafish swimming distance vs average ΔG correlation (preliminary) Spearman ρ=−1.0
- – Gray60 module showed strongest positive correlation with AD; black module strongest negative r=0.46; r=−0.45
- ▲ XGBoost classification accuracy on batch-corrected datasets AUROC >95
- ▲ XGBoost AUROC on GSE15222 and GSE118553 AUROC >85
- – BMDL of benz[a]anthracene and anthracene with corresponding average ΔG for PARP1/PTPN1/ITGA4 BMDL 7.856 (ΔG −9.4) and 10.096 (ΔG −8.1)
- – DEGs identified in AD vs control 1,380 upregulated; 1,467 downregulated
- correlation ρ = 1, p = 0.0417 (linear correlation of category-average ΔG with LD50 data)
- correlation Spearman ρ = −1.0, p = 0.167 (zebrafish swimming distance vs average ΔG)
- other BMDL = 1.723 × ΔG + 22.602 (empirical BMDL∼ΔG relationship, slope 1.723)
- correlation r = 0.46 (Gray60 module positive correlation with AD)
- correlation r = −0.45 (black module negative correlation with AD)
- count 7,153 proteins (GeneCards AD-related proteins above median relevance score 2.87)
- count 1,380 up / 1,467 down (DEGs in AD dataset vs control)
- other AUROC > 95 (XGBoost prediction on three batch-corrected datasets)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This primarily computational study integrates differential expression analysis of merged GEO microarray datasets (batch-corrected) with WGCNA, PPI network analysis, and XGBoost feature selection to identify PAH protein targets relevant to Alzheimer's disease. Molecular docking (AutoDock) produced ΔG values for 16 PAHs against three core proteins; hierarchical clustering grouped PAHs into eight categories, and category-average ΔGs were correlated with literature LD50 and calculated BMDL values using Spearman rank correlation and linear regression. A small zebrafish motility experiment (n = 3 treatment groups) provided preliminary experimental support, reporting Spearman ρ = −1.0, p = 0.167.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential expression analysis with adjusted p value (padj) threshold | GEO microarray datasets (GSE44771, GSE44770, GSE33000) batch-corrected; volcano plot of AD vs. control DEGs | 2,847 DEGs identified at |log2FC| ≥ 0.07 and padj ≤ 0.001 (1,380 up, 1,467 down) | not stated |
| Weighted Gene Co-expression Network Analysis (WGCNA) with soft-threshold β = 5, Pearson-based adjacency | Module-trait correlation to AD phenotype; hub gene selection at r > 0.4 (module-trait) and kME_MM > 0.8 | 1,164 hub genes identified | not stated |
| XGBoost classifier evaluated by AUROC | Feature importance ranking of top-5 downstream genes per PAH target protein in five independent GEO datasets | null | not stated |
| Spearman rank correlation (ρ = 1, p = 0.0417) | Category-average docking ΔG vs. literature LD50 values across PAH categories (Figure 9B) | null | not stated |
| Linear regression (BMDL ~ ΔG); slope 1.723, intercept 22.602 | Category-average ΔG vs. calculated BMDL for four PAHs (Figure 9A) | 4 PAHs with available BMDL data | not stated |
| Spearman rank correlation (ρ = −1.0, p = 0.167) | Average ΔG vs. zebrafish swimming distance across PAH treatment groups (Figure 9D) | n = 3 (PAH-concentration data points) | not stated |
| Benchmark Dose (BMD) modeling; Hill model selected by minimum AIC with visual fit assessment | Dose-response fitting for fluorene and phenanthrene (earthworm digging), benz[a]anthracene and anthracene (hippocampal cell MAO) | null | not stated |
| Leading eigenvector community detection algorithm | Subdivision of 145-gene PPI network into six sub-networks (Cytoscape) | 145 genes | na |
| Hierarchical clustering | Clustering of 16 PAHs × 3-protein ΔG matrix into eight toxicity-similarity categories | 16 PAHs | na |
| GO and KEGG overrepresentation enrichment analysis | Functional annotation of 145 PPI network genes across six sub-networks | 145 genes | not stated |
| Principal component analysis (PCA) | Batch effect assessment and dimensionality reduction of six GEO microarray datasets before and after batch correction | null | na |
-
The BMDL~ΔG linear regression and LD50 Spearman correlation were each calibrated on a very small number of data points (4 PAHs for BMDL; apparent n ≈ 4–5 categories for LD50, giving ρ = 1, p = 0.0417)↳ Could also: Prediction intervals for the regression line, alongside leave-one-out cross-validation, could also be reported to characterize how wide the uncertainty band is for an unobserved PAH — With very small calibration sets, a prediction interval conveys the expected range for a new observation rather than the mean trend alone, and leave-one-out cross-validation assesses whether the linear relationship generalizes outside the training points — both are standard diagnostics for small-sample regression models
-
XGBoost feature importance (split-based gain) was used to rank and select genes across datasets↳ Could also: LASSO regularized logistic regression with cross-validated penalty selection, or Random Forest with permutation-based importance, could also be applied for gene feature selection — LASSO provides a statistically principled sparsity penalty whose λ can be cross-validated, and permutation importance in Random Forest is less sensitive to correlated features than tree-split-based scores; either approach would offer a complementary or confirmatory ranking alongside XGBoost
-
Hierarchical clustering grouped 16 PAHs into eight categories based on ΔG profiles, with the number of clusters and linkage chosen by the analysts↳ Could also: Consensus clustering or gap-statistic-guided k-means could also be applied to select the number of clusters empirically and assess cluster stability across resampling iterations — The choice of cluster count directly determines the category-average ΔG values that enter the regression; stability-based methods provide a data-driven rationale for the number of clusters and a measure of confidence in PAH category membership
-
DEG thresholds were set at a relatively low |log2FC| ≥ 0.07 combined with padj ≤ 0.001↳ Could also: A more stringent fold-change filter (e.g., |log2FC| ≥ 0.5 or ≥ 1.0) alongside the same padj threshold could also be used, as is common in microarray transcriptomic studies — Combining a biologically meaningful effect-size filter with statistical significance is standard practice to reduce the influence of statistically significant but very small expression differences; the threshold choice affects the size of the gene list entering downstream PPI and WGCNA analyses
-
WGCNA module-trait associations were assessed using Pearson correlation with fixed thresholds (r > 0.4 for module-trait, kME_MM > 0.8 for hub genes)↳ Could also: Permutation-based significance testing of module-trait correlations, or bicor (biweight midcorrelation) in place of Pearson, could also be used to set empirically derived thresholds and improve robustness to outlier samples — Permutation testing accounts for the inherent correlation structure among genes and provides a data-adaptive threshold rather than a fixed cutoff; bicor is recommended in the WGCNA literature for data with outliers or non-normal expression distributions
-
Multiple individual BMD candidate models were available; the Hill model was selected based on AIC and graphical visual fit↳ Could also: Model averaging across all adequately fitting BMD models (e.g., using BMDS model-averaging procedures) could also be applied to derive BMD/BMDL estimates that account for model uncertainty — When multiple dose-response models fit the data comparably, model-averaged BMDL estimates incorporate uncertainty about the correct functional form, which is a recommended approach in EPA benchmark dose guidance and typically yields wider, more conservative confidence limits than single-model selection
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
0 downstream papers · 1 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41953002
Paper: Chen C et al. "Machine learning and free energy clustering reveal PAH protein binding linked to AD risk." iScience 2026. DOI 10.1016/j.isci.2026.115311. Code: https://github.com/pdpmb/PAHs (commit pinned at run time) Data: GEO GSE44771 (VC), GSE44770 (PFC), GSE33000 (PFC); + GSE15222, GSE118553 (ML validation).
What the repo actually ships
Loose, un-orchestrated R scripts, one per analysis step:
limma, wgcna1/2/3, xgboost, XG_GSE44771_GSE44770_GSE33000,
benchmarking against similar methods, go_and_kegg, docking. README says
"This code is mainly used for submission requests and will be updated after the
official publication."
Critical reproducibility gaps (recorded, not worked around):
- Every script begins
setwd("«path»")— author's local machine. - All scripts
load()intermediate binaries that are NOT in the repo:expr_last1_bactch.Rdata(the merged, batch-corrected expression matrix +group_factor),dat_factor1..8.rda,xgbfit.RData. - The 145/139-gene panel is read from
«path»— not shipped. - No preprocessing/merge/batch-correction script is shipped — the step that
turns the three raw GEO series into
expr_last1_bactchis absent. - No README run instructions, no environment lock, no entrypoint, no expected-output files.
⇒ The shipped code is not runnable end-to-end as-is. Per the brief (P16 / third-party-tool clause and 80/20), we reproduce the clearly-specified pipeline outputs by re-implementing the documented steps on the paper's own named data, following the shipped scripts' exact statistical choices.
In scope (pipeline-derived, attempted)
| Result | Pipeline | Reported value (location) |
|---|---|---|
| DEG count | limma AD-vs-CTL on merged+batch-corrected GSE44771/44770/33000, ` | log2FC |
| ML classifier | XGBoost (nrounds=100, max_depth=3, eta=0.1, gamma=0.5, 5-fold CV, 70/30 split) on the network gene panel | AUROC > 95 on GSE44771/44770/33000 (Fig 5) |
The limma thresholds and the XGBoost hyper-parameters are taken verbatim from
the shipped limma and xgboost scripts.
Out of scope (not attempted) — with reason
- WGCNA hub-gene set (1,164) / 145-gene PPI network (Fig 4): module choice (gray60, black), soft-power, MAD filtering and the PAH-target merge involve manual curation + an external PAH-target predictor not specified runnably; the 145-gene list itself is not shipped. (last-20%, skipped.)
- Molecular docking ΔG (BaP −10.1, anthracene −8.1, etc., Fig 6): wet-/manual structural docking (AutoDock-class), out of the bioinformatic-pipeline scope (P2).
- GO/KEGG enrichment: reported only qualitatively (term names, no numbers to pin) → no specific reproducible value.
- BMDL linear model (Fig 9): derived from docking + toxicology, not a pipeline.
- GSE15222 / GSE118553 external ML validation (AUROC>85): secondary; attempted only if the primary XGBoost run leaves budget.
Honesty note
Because the merge + batch-correction step and the exact gene panel are not shipped, an exact numeric match is not expected. We reconstruct the standard pipeline (GEOquery → common-symbol merge → ComBat → limma / XGBoost) and report how close the regenerated numbers come — a partial / concordance result is the honest expected outcome, and is explicitly a valid result under the brief.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The public GEO data is identical but the authors' processed matrix and 145-gene panel are not deposited and the repo is not runnable end-to-end, so the reproduction is a faithful re-implementation. The XGBoost AUROC>0.95 claim is confirmed (0.9934) and sample composition is exact, but the C1 DEG count is 2.4x off (6816 vs 2847, Fig 2D), traceable to the unshipped preprocessing/gene-filter step — an authors'-side reproducibility limitation, not a contradiction. Direction and canonical AD top hits (CRH/DUSP4/VGF/BDNF) hold and there is no evidence of fabrication. Overall solid-with-explainable-deviations, with the central conclusion only partly testable since the WGCNA/docking/BMDL arms were out of scope.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.