Autoencoder Networks Decipher the Association between Lung Cancer and Alzheimer's Disease.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the headline result 1:1. The repo (windytian/AE_geneexpression) ships outputs only, but crucially includes the two TRAINED KEGG autoencoders (SofaSofa_modelAD.h5 input=171, SofaSofa_modelCA.h5 input=465) plus all four KEGG expression matrices, so the paper's Figure 4 cross-disease correlation is a DETERMINISTIC load-and-predict (no retraining). On «our HPC» (TF 2.13.1/Keras 2.13.1) we rebuilt the 1-node encoders, applied the script's MinMax norm, and got: AD data Pearson 0.646 / Spearman 0.694 vs reported 0.643 (Pearson essentially EXACT, |delta|=0.003); LC data Pearson 0.363 / Spearman 0.364 vs reported 0.411 (within ~0.05, same positive sign, p<0.001). Both scatterplots visually match Fig 4. Strong consistency check: shipped matrices have exactly 177 AD cases and 251 LC cases, matching the paper's reported disease totals -> the data is genuinely the paper's; no fabrication flag for Fig 4. NOT attempted (the hard ~20%): Fig 3 (266-DEG AE, reported 0.825/0.316) because the repo ships no DEG model and no DEG expression matrices -> it would require re-deriving 266 DEGs from 9 raw GEO series and RETRAINING with unseeded weight init (non-deterministic). Also not re-run: the upstream limma DEG / KEGG-enrichment meta-analysis over the 9 GEO series.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 74assessed: 2026-06-15 ⛓ a8446ac265af
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusBecause conventional bioinformatics methods (e.g., DEG analysis) yield inconsistent conclusions about whether lung cancer and Alzheimer's disease are inversely or positively associated, can autoencoder networks that compress shared dysregulated genes into a one-dimensional representation determine the direction of association between the two diseases?
- ★ Autoencoder networks based on 266 shared DEGs reveal a comorbidity (positive) relationship between Alzheimer's disease and lung cancer. finding
- ★ A novel autoencoder-based bioinformatics method can determine the direction of association between distinct diseases by generating a one-dimensional pseudogene representation of dysregulated genes. method
- ★ Separate AE networks trained per disease on shared DEGs allow a unified estimate of inter-disease association via correlation of encoder outputs. method
- FGF2, SNCA, and LDHA are hub genes located at the centers of the shared-DEG interaction subnetworks. finding
- The HIF-1 signaling pathway is implicated in the comorbidity observed between lung cancer and Alzheimer's disease. mechanism
- Of 266 overlapped DEGs, 36% are concordantly regulated and 64% inconcordantly regulated, explaining inconsistent association inferences from conventional methods. finding
- Python AE modeling code is provided in a public GitHub repository. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Microarray gene expression profiling (Alzheimer's disease cohort) | Human brain tissue (GSE4757, GSE48350, GSE5281) | none (disease vs control) | Gene expression / differentially expressed genes | Affymetrix HG-U133 Plus 2.0 |
| Microarray gene expression profiling (lung cancer cohort) | Human lung tissue (GSE18842, GSE102287, GSE19804, GSE19188, GSE103888, GSE118370) | none (disease vs control) | Gene expression / differentially expressed genes | Affymetrix HG-U133 Plus 2.0 |
| Differential expression analysis (moderated t-tests) | Integrated AD and lung cancer microarray datasets | none | DEGs (FDR<0.05, |log fold change|>0.5) | R limma package |
| Autoencoder deep learning (one-dimensional representation/pseudogene) | 266 shared DEGs across AD and lung cancer cohorts | none | Encoder bottleneck score; Spearman correlation of cross-disease predicted values | Python Keras/TensorFlow |
| Protein-protein interaction network and hub gene analysis | 266 overlapped DEGs | none | Gene-gene interaction network, top 50 hub genes by connectivity | STRING (confidence 0.7), Cytoscape/CytoHubba |
| Pathway enrichment analysis (KEGG and GO) | Overlapped DEGs of AD and lung cancer | none | Enriched KEGG pathways and GO terms (FDR cutoff 0.2) | R clusterProfiler package |
- ▲ Spearman correlation between predicted values of the two AE networks for the Alzheimer's disease test set 0.825 (p<0.001)
- ▲ Spearman correlation between predicted values of the two AE networks for the lung cancer test set 0.316 (p<0.001)
- ▲ AE retrained on KEGG-pathway genes: correlation for the Alzheimer's disease test set 0.643 (p<0.001)
- ▲ AE retrained on KEGG-pathway genes: correlation for the lung cancer test set 0.411 (p<0.001)
- – 266 DEGs overlapped between the two diseases; 21 co-upregulated, 75 co-downregulated, 170 inversely expressed 266 genes
- – Lung cancer DEGs: 1,935 downregulated and 1,353 upregulated between disease and control 3288 DEGs
- – Alzheimer's disease DEGs: 508 upregulated and 263 downregulated between disease and control 771 DEGs
- – 148 of overlapped DEGs directly related and 116 indirectly related to lung cancer (146 direct/118 indirect for AD) per GeneCards
- correlation 0.825 (p<0.001) (Spearman correlation, AD test set, AE networks on 266 DEGs)
- correlation 0.316 (p<0.001) (Spearman correlation, lung cancer test set, AE networks on 266 DEGs)
- correlation 0.643 (p<0.001) (Spearman correlation, AD test set, AE on KEGG pathway genes)
- correlation 0.411 (p<0.001) (Spearman correlation, lung cancer test set, AE on KEGG pathway genes)
- count 251 lung cancer patients and 213 normal controls (Integrated lung cancer cohort)
- count 177 patients and 257 controls (Integrated Alzheimer's disease cohort)
- count 266 overlapped DEGs (36% concordant, 64% inconcordant) (Shared DEGs between the two diseases)
- other 22 seconds per single AE run (Runtime on AMD Ryzen 7 4800U, 16 GB RAM laptop)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper integrates multiple GEO microarray cohorts for lung cancer (n=464) and Alzheimer's disease (n=434) and applies moderated t-tests with Benjamini-Hochberg FDR correction to identify differentially expressed genes (DEGs) in each disease. Overlapping DEGs (n=266) are then used to train separate autoencoder (AE) deep learning networks for each disease, compressing gene expression into a one-dimensional bottleneck representation. The association direction between the two diseases is assessed by computing Spearman's correlation coefficients between the AE-derived scores across test sets, with the direction and magnitude of the correlation taken as the primary result.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Moderated t-test (limma empirical Bayes) | DEG identification: lung cancer patients vs. normal controls in integrated LC cohort; Alzheimer's disease patients vs. controls in integrated AD cohort | LC: 464 total (251 cases, 213 controls); AD: 434 total (177 cases, 257 controls) | not stated |
| Spearman's rank correlation coefficient | Correlation between AE-predicted one-dimensional scores from the lung cancer AE model vs. the Alzheimer's disease AE model, computed on each disease's held-out test set | Test set size not explicitly stated (derived from 40% of combined cohort, split 3:2) | not stated |
-
Spearman correlation between AE bottleneck scores was reported without confidence intervals or bootstrapped uncertainty estimates↳ Could also: Bootstrap resampling (e.g., 1,000 iterations) could be used to produce 95% confidence intervals around each Spearman r, and a Fisher z-transformation test could be used to formally compare the two correlation coefficients (r=0.825 vs. r=0.316) — Confidence intervals would convey the precision of each correlation estimate and clarify whether the difference between the AD and LC test-set correlations is itself statistically distinguishable, which is directly relevant to the paper's central claim about comorbidity
-
The AE hyperparameters were selected via grid search minimizing MSE on a single 3:2 train/validation split↳ Could also: k-fold cross-validation (e.g., 5- or 10-fold) could be used for hyperparameter selection and to estimate out-of-sample MSE with variability — A single train/validation split can yield an optimistic or noisy estimate of generalization error depending on how the split fell; k-fold provides a more stable estimate of MSE across partitions, which is especially relevant when the combined cohort is moderately sized (~400–900 samples)
-
Multiple GEO datasets were combined using fRMA normalization followed by ComBat batch correction, then analyzed as a single pooled cohort↳ Could also: A formal meta-analytic framework (e.g., fixed-effects or random-effects meta-analysis of per-study log-fold changes and standard errors, as implemented in R/metafor or limma's meta-analysis mode) could also be used to synthesize evidence across studies — Meta-analysis explicitly models between-study heterogeneity and produces per-gene heterogeneity statistics (e.g., I²), making it straightforward to assess which DEGs are consistently dysregulated across studies versus driven by one or two datasets
-
No multiplicity correction was applied to the four Spearman correlation tests (two primary and two secondary AE analyses)↳ Could also: A Bonferroni or Holm correction across the four correlated tests, or pre-registration of which comparison is primary, could also be applied — When multiple related tests are reported from the same dataset and all support the same conclusion, acknowledging the family of tests and applying or discussing a correction helps readers calibrate the strength of evidence
-
DEG thresholds were set at FDR<0.05 and |log-fold change|>0.5; pathway enrichment used a loosened FDR<0.2↳ Could also: The pathway enrichment FDR threshold could be kept at 0.05 to match the DEG analysis, or the relaxed threshold could be explicitly justified with a sensitivity analysis comparing results at FDR<0.05 and FDR<0.2 — Reporting results under both thresholds would clarify which enriched pathways are robust to the choice of cutoff versus appear only under the more lenient criterion
-
The association between the two diseases was summarized by a single Spearman correlation per test set, without adjustment for potential covariates (e.g., age, sex, study of origin)↳ Could also: A linear mixed model or partial correlation controlling for study-of-origin (as a random or fixed effect) could also be applied to the AE scores, since samples come from multiple GEO studies with different demographic profiles — Residual between-study variation in the AE scores after ComBat could inflate or deflate the observed correlation; a mixed-model approach would partition this source of variation explicitly
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
- CoINcIDE: A framework for discovery of patient... L1 87/100
- A curated collection of transcriptome datasets... L1 62/100
- An NMF-Based Methodology for Selecting Biomark... L1 84/100
- Unveiling prognostics biomarkers of tyrosine m...⚑ L1 51/100 ⚑
- Meta-analysis of gene expression profiles of l... L1 78/100
- Colorectal Cancer Prediction Based on Weighted...⚑ L1 80/100 ⚑
- Curation of over 10 000 transcriptomic studies... L1 80/100
- Construction and Validation of an Immune Infil...⚑ L1 51/100 ⚑
- Identification of a novel 10 immune-related ge...
- Exploration of the shared diagnostic genes and... L1 76/100
- IRSN-23 gene diagnosis enhances breast cancer... L1 71/100
- Molecular Classification Models for Triple Neg... L1 86/100
- Predicting Bone Metastasis Using Gene Expressi... L1 62/100
- Comprehensive analysis of a novel RNA modifica... L1 71/100
- Discovery and validation of molecular patterns... L1 83/100
- Comparative profiling of skeletal muscle model... L1 64/100
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36518809
Paper: Li J, Gao X, Tang M, Wang C, Liu W, Tian S. Autoencoder Networks Decipher the Association between Lung Cancer and Alzheimer's Disease. Comput Intell Neurosci. 2022. PMID 36518809 / PMC9744611 / DOI 10.1155/2022/2009545.
Repo: https://github.com/windytian/AE_geneexpression (own code, P16 own-tool).
Ships outputs only, no driver script wired to raw data: two Python scripts
saved as .txt (python-DEG.txt, KEGG/python-KEGG.txt), gene lists,
KEGG-filtered+normalized expression matrices, and two trained Keras
autoencoders (KEGG/SofaSofa_modelAD.h5, KEGG/SofaSofa_modelCA.h5).
Pipeline (what produces the reported results)
- Per-disease meta-analysis of GEO microarrays (AD: GSE4757+GSE48350+GSE5281;
LC: GSE18842+GSE102287+GSE118370+GSE19188+GSE19804+GSE103888), RMA/normalize,
limmamoderated t-test DEGs (FDR<0.05, |logFC|>0.5). - 266 overlapping DEGs between LC and AD.
- KEGG enrichment of disease genes -> AD gene set (171) and LC/CA gene set (465).
- Two 1-D-bottleneck autoencoders (128-64-10-1 encoder, mirror decoder, ReLU, dropout 0.2, Adam lr 5e-4, 3:2 split). Cross-apply each disease's encoder to both diseases' samples; scatter the two 1-D embeddings; report Spearman rho.
Reported quantitative results
- Fig 3 (266 DEG AE): AD test rho = 0.825 (p<0.001); LC test rho = 0.316 (p<0.001).
- Fig 4 (KEGG-pathway AE): AD test rho = 0.643 (p<0.001); LC test rho = 0.411 (p<0.001).
In scope (attempted) — DETERMINISTIC, no retraining
Fig 4 (KEGG AE). Repo ships the trained CA (465-input) + AD (171-input)
models AND all four KEGG expression matrices
({AD,CA}KEGGgenes{AD,CA}EXP.txt). Load shipped model -> rebuild 1-node encoder
-> same MinMax norm as python-KEGG.txt -> predict 1-D embeddings -> Spearman of
the two scatters. Inference is deterministic (dropout off), so this reproduces
the exact shipped-model figure, modulo the arbitrary sign of a 1-D bottleneck
(compare |rho|).
Out of scope (not attempted) — and why
- Fig 3 (266-DEG AE): repo ships no DEG-level model and no
LCexp266DEG.txt/ADexp266DEG.txtmatrices -> nothing deterministic to load; would require re-deriving 266 DEGs from 9 raw GEO series + retraining (random init, no seed control on weight init) -> non-deterministic, the hard 20%. - Steps 1-3 (GEO download, RMA, limma DEGs, KEGG enrichment): upstream of the shipped artifacts; the 266/171/465 counts are inputs we take as given. Not re-run (multi-series meta-analysis, out of the 80/20 sweet spot).
- Retraining either AE: weight init is unseeded in the shipped scripts -> not byte-reproducible; deterministic shipped-model inference is the faithful 1:1 check instead.
Compute
All on «our HPC» (SLURM «job»): conda python 3.10 + tensorflow-cpu==2.13.1
(Keras 2, matches the 2022 .h5), scikit-learn, scipy. Tiny job (~min).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The headline Figure 4 cross-disease correlations reproduce cleanly from the shipped trained autoencoders and matrices: AD 0.643 -> Pearson 0.646 (|delta|=0.003) and LC 0.411 -> Spearman 0.364 (|delta|~0.05), both p<0.001 with the same positive direction, and shipped sample counts (177 AD / 251 LC) match the paper exactly, so there is no fabrication concern and the central LC-AD association holds. The only deviations are small and on the technical/our-method side (float ordering, full-data-vs-subset, and a rho-vs-Pearson labeling ambiguity). What lowers the overall grade to yellow is scope/data availability, not an authors' defect: Figure 3's 266-DEG autoencoder (0.825/0.316) and the upstream limma DEG meta-analysis are not derivable from the shared repo because no DEG model or matrices were deposited.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.