Integrative bioinformatics and artificial intelligence analyses of transcriptomics data identified genes associated with major depressive disorders including <i
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH? Partly. autoNeuro is a generic tabular-ML grid-search tool (LR/SVM/RF/XGB/LGBM x SelectKBest/SelectFromModel, 10-fold StratifiedKFold) applied to GSE98793; the repo ships ONLY the ML stage. Its training inputs (per-batch gene-symbol CSVs) are NOT in the repo, and the GEO->matrix preprocessing (MAS5>50 & CV>10% -> 1446 genes; 157-sample subset) cannot be reconstructed from the public GCRMA series matrix. RESULT: running the unmodified autoNeuro grid on a faithful-as-feasible reconstruction (top-1446-variance gene panel, native 96/96 batch split, full 192 samples) reproduces the SAME performance regime as Table 4 - batch2 within tolerance and slightly exceeding the paper (acc 0.81 vs 0.75, AUC 0.80 vs 0.79, F1 0.76 vs 0.70), batch1 and merged lower by 0.06-0.14 (acc 0.70/0.72 vs 0.79/0.80) but every std band overlaps the paper's. No hard mismatch, no fabrication signal; the shortfalls are explained by the two unreconstructable preprocessing steps that would plausibly lift the paper's numbers. 1:1 vs different: DIFFERENT in exact point values, SAME in regime -> partial. NOT attempted: GSEA gene lists, the NRG1/10-gene biomarker panel, transfer-learning external validation, wet-lab qPCR.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 66assessed: 2026-06-20 ⛓ 7ee86da4537f
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-20
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether integrating bioinformatics (GSEA) and machine learning analysis of transcriptomic data from MDD patients versus healthy controls can identify reliable, non-invasive diagnostic biomarkers for major depressive disorder.
- ★ Differentially expressed genes in MDD patients are enriched in immune response, inflammatory response, neurodegeneration, and cerebellar atrophy pathways. finding
- ★ Feature selection combined with ML algorithms produced predictive models distinguishing MDD from healthy controls with ≥75% accuracy. finding
- ★ Integrative bioinformatics and ML analysis identified ten key MDD-related biomarkers: NRG1, CEACAM8, CLEC12B, DEFA4, HP, LCN2, OLFM4, SERPING1, TCN1 and THBS1. finding
- ★ NRG1 was the most robust and reliable biomarker distinguishing MDD patients from healthy controls across independent external datasets of mixed populations. finding
- ★ NRG1 is upregulated in saliva samples of MDD patients compared to healthy controls in an independent Kazakhstan cohort. finding
- ★ NRG1 shows high expression in subcortical limbic brain regions implicated in depression, per functional mapping to human brain regions. finding
- An integrative pipeline combining absolute/normal GSEA, feature selection methods (PCA, SelectKBest, LR, RF), and multiple ML classifiers (LR, RF, XGBoost, SVM, KNN) was developed to identify MDD biomarkers. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| GSEA (gene set enrichment analysis) | whole blood microarray, human MDD patients and healthy controls (GSE98793) | none (case-control) | enriched cellular pathways and differentially expressed leading genes | Affymetrix Human Genome U133-Plus 2.0 gene-chip |
| Machine learning classification (LR, RF, XGBoost, SVM, KNN with feature selection) | whole blood transcriptomic data, human MDD patients and healthy controls (GSE98793) | none (case-control) | classification accuracy, F1 score, ROC-AUC for MDD vs HC prediction | SkLearn v1.2.1, Python v3.9 |
| Differential gene expression analysis (Limma) | transcriptomic datasets from four external cohorts (GSE99725, GSE76826, GSE38206, GSE32280; French/Caucasian, Japanese, Chinese populations) | none (case-control) | differentially expressed genes between MDD and HC | Limma package, R software |
| Functional mapping and annotation of gene expression to brain regions | postmortem human brain tissue (six neurotypical adult brains, ~3700 tissue samples) | none | regional brain expression of biomarker genes, including NRG1 | Allen Human Brain Atlas (AHBA) |
| qRT-PCR | saliva samples from MDD patients and healthy controls, Kazakhstan population | none (case-control) | NRG1 mRNA relative/fold expression change (normalized to 18S rRNA) | QuantStudio3 system with Maxima SYBR Green/ROX qPCR Master Mix |
- – DEGs in MDD patients were significantly enriched in immune response, inflammatory response, neurodegeneration and cerebellar atrophy pathways. p < 0.05
- – ML models predicted MDD status based on MDD-altered genes. ≥75% accuracy
- – Ten key MDD-related biomarkers identified: NRG1, CEACAM8, CLEC12B, DEFA4, HP, LCN2, OLFM4, SERPING1, TCN1, THBS1.
- – NRG1 best distinguished MDD patients from healthy controls across independent external datasets with mixed populations.
- ▲ NRG1 expression was upregulated in saliva of MDD patients compared to healthy controls.
- ▲ NRG1 showed high expression in main subcortical limbic brain regions implicated in depression.
- count 170 MDD patients and 121 healthy controls (total samples mined from publicly available transcriptomic datasets)
- count 128 MDD patients and 64 healthy controls (discovery dataset GSE98793)
- count 42 MDD patients and 57 healthy controls (combined cohort across four external confirmation datasets)
- count 12 MDD patients and 8 healthy controls (saliva sample validation cohort from Kazakhstan)
- pvalue p < 0.05 (significance cut-off for GSEA and DEG identification)
- fold_change FC > 1.5 (upregulated); FC < 0.5 (downregulated) (thresholds defining differentially expressed genes in GSEA)
- other ≥75% accuracy (performance of ML models predicting MDD vs healthy control status)
- count ~25,828 annotated cellular pathways across seven gene sets (MSigDB gene sets used for absolute GSEA analysis)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study combined bioinformatics and machine-learning analyses of microarray transcriptomic data (GSE98793: 128 MDD patients, 64 healthy controls) to identify MDD-associated genes. Differential pathway enrichment was assessed with Gene Set Enrichment Analysis (GSEA, permutation-based, p<0.05, fold-change cutoffs), and differential expression in four independent external cohorts was assessed with the Limma R package (p<0.05). Multiple machine-learning classifiers (logistic regression, random forest, XGBoost, SVM, KNN) with feature selection and 10-fold cross-validation were used to build and evaluate predictive models, reported via accuracy, F1-macro, and ROC-AUC; the top candidate gene (NRG1) was further evaluated by qRT-PCR in an independent saliva cohort.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Gene Set Enrichment Analysis (absolute GSEA followed by normal/standard GSEA, permutation-based) | Identification of enriched/activated pathways and differentially enriched genes, MDD vs healthy controls, discovery dataset | 128 MDD patients, 64 healthy controls | not stated |
| Limma (linear models for microarray data, moderated statistics) | Differential gene expression between MDD patients and healthy controls in each of four external validation datasets | varies by dataset per Table 1 (e.g., 18 vs 15; 12 vs 10; 9 vs 9; 8 vs 8) | not stated |
| ANOVA F-value (used as a feature-scoring/selection method, SelectKBest) | Feature (gene) selection prior to machine-learning classification | discovery dataset (GSE98793) | not stated |
| 10-fold cross-validated classification (logistic regression, random forest, XGBoost, SVM, KNN) evaluated by accuracy, F1-macro, ROC-AUC | Classification of MDD patients vs healthy controls based on selected gene features | discovery dataset merged batches; independent evaluation on 42 MDD/57 HC across four external datasets | na |
-
GSEA pathway significance is described using a p<0.05 threshold across thousands of annotated gene sets.↳ Could also: Reporting the FDR q-value that GSEA computes by default — would explicitly convey the expected proportion of false discoveries when testing thousands of gene sets simultaneously, which a raw p-value threshold alone does not capture
-
Differential expression between MDD and healthy controls in the external datasets was assessed with Limma at p<0.05.↳ Could also: Reporting Benjamini-Hochberg (or similar FDR) adjusted p-values, which Limma calculates by default — would help control the false discovery rate when testing many genes per array, a standard consideration in microarray/transcriptomic differential expression work
-
Classifier performance (accuracy, F1-macro, ROC-AUC) is summarized as point estimates from 10-fold cross-validation.↳ Could also: Reporting confidence intervals or a bootstrap distribution around these cross-validated metrics — would convey the uncertainty/variability of performance estimates, which is often informative when comparing multiple models or datasets of modest size
-
Feature selection (ANOVA F-value, PCA, embedded LR/RF importance) was performed before model training and cross-validation.↳ Could also: Embedding feature selection within each cross-validation fold (nested cross-validation) — would also address the well-known overfitting risk when the number of features greatly exceeds the number of samples, keeping feature selection and performance evaluation fully separated
-
Four independent external transcriptomic datasets from different populations were each analyzed separately with Limma.↳ Could also: A formal meta-analytic approach combining effect sizes or p-values across the four cohorts (fixed- or random-effects meta-analysis) — would allow pooled evidence and a single combined estimate of effect across ethnically diverse cohorts, in addition to per-dataset results
-
NRG1 saliva expression was compared between a small MDD cohort (n=12) and healthy controls (n=8), though the specific statistical test for this comparison is not shown in the provided excerpt.↳ Could also: A nonparametric test such as the Mann-Whitney U test — is often preferred for small sample comparisons where normality of the expression/fold-change data has not been established
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Running the unmodified autoNeuro grid on a faithful-as-feasible reconstruction of GSE98793 reproduces the same performance regime as Table 4 (acc 0.70–0.81, AUC 0.73–0.80, F1 0.67–0.76): batch2 within tolerance and slightly above the paper, batch1/merged lower by 0.06–0.14 but with overlapping std bands and no hard mismatch. The deviations sit on the input/preprocessing side — the unreconstructable MAS5>50 & CV>10% 1446-gene filter and the undocumented 157-sample subset — split between our forced method choices and the authors' underspecification, not a computation error. No fabrication signal; magnitudes are achievable. The biomarker/NRG1 panel that anchors the paper was out of scope, so the central biological claim itself remains untested here.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.