Identification of common genetic characteristics of rheumatoid arthritis and major depressive disorder by bioinformatics analysis and machine learning.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the two headline pipeline outputs 1:1 even though the paper ships NO own code (registry 'code' link slowkow/ggrepel is a generic plotting library; reproduced per P16 with the named third-party tools DESeq2/limma/pROC on the paper's own GEO data). STRONG near-exact hits: (1) RA DEG counts on GSE169082 via DESeq2 = 2166/1444/722 vs reported 2177/1458/719 (~1% off); (2) MDD hub-gene ROC AUCs on GSE38206 = EAF1 0.9105, SDCBP 0.8951, RNF19B 0.9074 vs reported 0.911/0.895/0.907 (match to 3 decimals) - confirms the 3 hub genes EAF1/SDCBP/RNF19B and the MDD diagnostic analysis are genuine. DIVERGENCES, all explainable: MDD DEG counts (1600/1364 or 1040/626 vs 592/441) - GSE38206 has 36 samples (9+9 patients/controls x 0w/8w) and the paper's design (timepoints/paired/probe-collapse) is unspecified, so the exact count is not pinnable (same order of magnitude). RA per-gene AUCs (0.978/0.6122/0.942) are NOT derivable from GSE169082 (n=7 gives degenerate AUC=1.0; 0.6122+CI impossible at that n) - they require the large RA cohort GSE97476 (246+30), so the RA-AUC dataset attribution is ambiguous in the text (flagged, not fabrication). NOT attempted (80/20): WGCNA modules, the stochastic LASSO+RandomForest gene-selection lists (seeds unstated; but its endpoint - the 3-gene AUC - IS verified), PPI/MCODE counts, GO/KEGG terms, nomogram AUC, and GSE97476 RA ROC. No fabrication detected; the two near-exact matches argue the core pipeline is real.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 58assessed: 2026-06-14 ⛓ 7475b262ec8a
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetRheumatoid arthritis (RA) and major depressive disorder (MDD) share overlapping symptoms and comorbidity, so the study tests whether RA and MDD have common genetic characteristics/pathogenesis that could serve as diagnostic biomarkers to distinguish the two conditions.
- ★ EAF1, SDCBP and RNF19B are common genetic characteristics (hub genes) shared by RA and MDD finding
- ★ Monocyte infiltration is a shared immune mechanism connecting RA and MDD finding
- ★ A nomogram built from EAF1, SDCBP and RNF19B has high diagnostic value for both RA and MDD finding
- ★ WGCNA combined with DEG analysis identifies disease-associated gene modules that intersect between RA and MDD datasets method
- ★ LASSO regression and random forest machine learning were combined to filter candidate hub genes for RA and MDD method
- ★ The shared gene sets between RA and MDD are enriched in immune response and inflammation pathways mechanism
- TIMER 2.0 database and Cibersort algorithm were used to correlate hub gene expression with immune cell infiltration method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| WGCNA (weighted gene co-expression network analysis) | PBMC, RA dataset GSE97476 | RA disease vs control | gene co-expression modules correlated with RA trait | GPL10904; R WGCNA package |
| WGCNA | PBMC, MDD dataset GSE38206 | MDD disease vs control | gene co-expression modules correlated with MDD trait | GPL13607; R WGCNA package |
| Differential gene expression (RNA-seq, DESeq2) | PBMC, RA dataset GSE169082 | RA disease vs control | differentially expressed genes (|log2FC|>1, adj.p<0.05) | GPL20795; DESeq2 |
| Differential gene expression (microarray, limma) | PBMC, MDD dataset GSE38206 | MDD disease vs control | differentially expressed genes (|log2FC|>0.4, adj.p<0.05) | GPL13607; limma |
| GO/KEGG functional enrichment analysis | intersected DEG/module gene sets (194, 135, 42 genes) | none | enriched biological pathways/ontologies | ClusterProfiler R package |
| Protein-protein interaction network construction with MCODE clustering | candidate genes from RA/MDD intersections | none | interacting node proteins/subnetworks | STRING database v11.5; Cytoscape |
| Machine learning feature selection (LASSO regression, random forest) | candidate RA and MDD genes (PBMC datasets) | none | ranked candidate hub genes | R packages glmnet and randomForest |
| ROC curve analysis and nomogram construction | PBMC, RA and MDD datasets | disease vs control | diagnostic AUC/95% CI for EAF1, SDCBP, RNF19B | R packages pROC and rms |
- – EAF1, SDCBP and RNF19B identified as the 3-gene intersection of RA (9 candidate genes) and MDD (7 candidate genes) from LASSO/RF machine learning
- ▲ Combined 3-gene nomogram diagnostic AUC for RA and MDD RA AUC 0.994 (95%CI 0.986-1.000); MDD AUC 0.969 (95%CI 0.925-1.000)
- – Individual gene diagnostic AUC in RA EAF1 0.978 (0.961-0.995); SDCBP 0.612 (0.523-0.702); RNF19B 0.942 (0.909-0.975)
- – Individual gene diagnostic AUC in MDD EAF1 0.911 (0.819-1.000); SDCBP 0.895 (0.789-1.000); RNF19B 0.907 (0.804-1.000)
- ▲ EAF1, SDCBP and RNF19B were significantly upregulated in both RA and MDD patients versus controls (Wilcoxon test)
- ▲ Monocyte levels were higher in RA and MDD patients than in controls
- – DEGs identified in RA (GSE169082) and MDD (GSE38206) datasets RA: 2177 DEGs (1458 up, 719 down); MDD: 592 up, 441 down
- – Gene sets from WGCNA/DEG intersections were enriched for immune response and inflammation pathways
- fold_change RA nomogram AUC 0.994 (95%CI 0.986-1.000) (combined 3-gene diagnostic model for RA)
- fold_change MDD nomogram AUC 0.969 (95%CI 0.925-1.000) (combined 3-gene diagnostic model for MDD)
- correlation lightcyan module r=0.60, p=2e-26 (WGCNA module-trait correlation with RA)
- correlation blue module r=-0.69, p=8e-37 (WGCNA module-trait correlation with RA)
- correlation turquoise module r=0.61, p=1e-04 (WGCNA module-trait correlation with MDD)
- correlation pink module r=-0.84, p=8e-10 (WGCNA module-trait correlation with MDD)
- count 2177 DEGs (1458 up, 719 down) (GSE169082 RA DEGs via DESeq2)
- other RA patients have 47% higher risk of depression than controls (meta-analysis of 11 cohort studies cited in introduction)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This bioinformatics study reanalyzed three public GEO PBMC gene-expression datasets (two RA, one MDD) using a multi-step pipeline: differential expression analysis (DESeq2 on GSE169082; limma on GSE38206), weighted gene co-expression network analysis (WGCNA on GSE97476 and GSE38206), PPI network construction with MCODE filtering, and two machine-learning algorithms (LASSO and Random Forest) whose intersecting outputs nominated three hub genes (EAF1, SDCBP, RNF19B). These genes were assessed by Wilcoxon group comparisons and ROC/AUC analysis with 95% CIs, and a nomogram was constructed. Immune cell infiltration was quantified by CIBERSORT and compared between groups via Wilcoxon tests, with linear-fit scatterplots relating hub gene expression to immune cell fractions.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| DESeq2 (negative binomial GLM / Wald test) | Differentially expressed genes in RA dataset GSE169082 | 7 (3 controls, 4 RA) | not stated |
| limma linear model (moderated t-statistic) | Differentially expressed genes in MDD dataset GSE38206 | 36 (18 controls, 18 MDD) | not stated |
| WGCNA Pearson module-trait correlation | Module-trait associations in GSE97476 (RA) and GSE38206 (MDD) | GSE97476: 276 (30 controls, 246 RA); GSE38206: 36 | not stated |
| GO and KEGG enrichment analysis (ClusterProfiler, adjusted p < 0.05) | Functional annotation of gene sets at four filtering stages (194, 135, 42, and 17 genes) | null | na |
| LASSO regression (glmnet) | Candidate hub gene selection for RA and MDD | null | not stated |
| Random Forest importance ranking (randomForest) | Candidate hub gene selection and importance ranking for RA and MDD | null | not stated |
| Wilcoxon rank-sum test | Comparison of 22 CIBERSORT immune cell proportions between patient and control groups in RA and MDD; expression comparison of three hub genes between groups | RA (GSE97476): 276; MDD (GSE38206): 36 | not stated |
| ROC / AUC analysis with 95% CI (pROC) | Diagnostic value of EAF1, SDCBP, RNF19B individually and as a nomogram-combined set for RA and MDD | RA: 276; MDD: 36 | na |
-
Wilcoxon tests were applied to compare proportions of 22 immune cell types between patient and control groups in both RA and MDD without a stated adjustment for the resulting family of approximately 44 simultaneous comparisons↳ Could also: Apply Benjamini-Hochberg FDR adjustment across all Wilcoxon tests within each disease comparison — BH-FDR adjustment would quantify the expected proportion of false discoveries among the significant immune cell findings, which is informative when many simultaneous tests are run on correlated outcomes such as CIBERSORT fractions that sum to one
-
Hub gene selection combined LASSO and Random Forest by taking the intersection of each method's top output with an ad hoc threshold (top-15/top-25 for RA; top-10/top-20 for MDD)↳ Could also: Use stability selection (bootstrap-based selection probabilities) or elastic net regularization (alpha between 0 and 1 in glmnet) to produce formal, threshold-independent selection probabilities — Stability selection provides calibrated false-discovery control for variable selection and reduces sensitivity to the specific top-N cutoff, which is otherwise an arbitrary choice
-
DEG fold-change thresholds differed between datasets (|log2FC| > 1 for GSE169082; |log2FC| > 0.4 for GSE38206 and GSE97476)↳ Could also: Apply a uniform log2FC threshold across all datasets, or combine the two RA datasets (GSE97476 and GSE169082) via a fixed-effects or random-effects meta-analysis of effect sizes before intersection with MDD DEGs — A consistent threshold or meta-analytic combination reduces the risk that the intersected gene list reflects dataset-specific cutoff choices, and pooling the RA datasets would leverage the larger sample size of GSE97476 (n=276) alongside GSE169082
-
The entire pipeline — from DEG and WGCNA discovery through machine-learning selection to ROC evaluation — was performed within the same datasets with no held-out partition or independent validation cohort↳ Could also: Apply k-fold cross-validation within the larger RA dataset (GSE97476, n=276) or validate the three hub genes in an independent publicly available GEO dataset — Internal cross-validation or external replication would provide an estimate of generalization performance less susceptible to optimistic bias, which is a recognized concern when the same data inform both feature selection and performance evaluation
-
Diagnostic performance of the hub genes and nomogram was summarized exclusively by AUC↳ Could also: Supplement AUC with calibration curves and decision curve analysis — AUC measures discrimination but not calibration; calibration curves show whether predicted probabilities correspond to observed event rates, and decision curve analysis evaluates net clinical benefit across a range of decision thresholds, together providing a more complete picture of diagnostic utility
-
Correlations between hub gene expression and immune cell fractions were visualized with linear-fit scatterplots via ggstatsplot without reporting quantified correlation coefficients or their uncertainty in the main text↳ Could also: Report Spearman (or Pearson) correlation coefficients with 95% CIs and FDR-adjusted p-values alongside the scatterplots — Quantified correlation statistics and their uncertainty allow readers to assess the strength and reliability of the gene–immune cell associations independently of the visual scale of the scatterplot
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37415981
Paper: Jiang W, Wang X, Tao D, Zhao X. Identification of common genetic characteristics of rheumatoid arthritis and major depressive disorder by bioinformatics analysis and machine learning. Front Immunol 2023;14:1183115. PMID 37415981 · PMCID PMC10320004 · DOI 10.3389/fimmu.2023.1183115.
Code-availability reality (P16 note)
The registry "Code" link is github.com/slowkow/ggrepel — a generic R
ggplot2 text-label library, NOT the authors' analysis code. The paper ships
no own analysis repository. Per brief rule 2 / P16, this does not down-rank
the paper: we reproduce by applying the standard third-party tools the paper
names (DESeq2/limma, glmnet/randomForest, pROC) to the paper's own GEO data,
following the parameters stated in Methods.
Datasets (all public GEO)
| GEO | role | platform | design | in scope? |
|---|---|---|---|---|
| GSE169082 | RA DEGs (THIS RU's assigned accession) | GPL20795 HiSeq X Ten, RNA-seq raw counts | 3 RA vs 4 control PBMC | YES — primary |
| GSE38206 | MDD DEGs | GPL13607 Agilent, microarray | 18 MDD vs 18 control PBMC | YES |
| GSE97476 | RA WGCNA | GPL10904 | 246 RA vs 30 control | partial (WGCNA) |
Pipeline-derived results (candidate claims)
| id | result | pipeline | grade-ability |
|---|---|---|---|
| C1 | RA DEGs from GSE169082: 2177 total (1458 up, 719 down), |log2FC|>1 & adjP<0.05 | DESeq2 on raw counts | deterministic — primary target |
| C2 | MDD DEGs from GSE38206: 592 up, 441 down, |log2FC|>0.4 & adjP<0.05 | limma on microarray | deterministic |
| C3 | Common up DEGs = 131, common down DEGs = 4 (intersection RA∩MDD) | set intersection on C1∩C2 | derivable (depends on C1,C2 + symbol mapping) |
| C4 | Hub-gene ROC AUC (EAF1/SDCBP/RNF19B) in RA: 0.978 / 0.612 / 0.942 | pROC on GSE169082 expr | deterministic |
| C5 | Hub-gene ROC AUC in MDD: 0.911 / 0.895 / 0.907 | pROC on GSE38206 expr | deterministic |
IN SCOPE (attempted)
- C1 RA DEG counts (GSE169082, DESeq2) — primary, cleanest 1:1.
- C2 MDD DEG counts (GSE38206, limma).
- C4/C5 ROC AUC of the 3 named hub genes EAF1, SDCBP, RNF19B in both datasets.
- C3 up/down DEG intersection (if symbol harmonization is clean).
OUT OF SCOPE (80/20 — not attempted, why)
- WGCNA module assignment (GSE97476, GSE38206): module colors/sizes (19 modules; lightcyan/salmon/purple; β=7/6) are sensitive to soft-threshold, cut height, and merge params not fully pinned → not a clean 1:1; skip per rule 3.
- LASSO + RandomForest feature selection (9 RA / 7 MDD candidates → 3 common): RF is stochastic (seed unstated) and LASSO λ-selection (lambda.min vs 1se) unstated → the specific 9/7/3 gene lists are not deterministically reproducible. We DO test the endpoint (the 3 hub genes' AUC, C4/C5) which is the headline ML claim. The intermediate gene-count path is the optional last 20%.
- PPI/MCODE counts (30 / 17 genes), GO/KEGG term lists, nomogram AUC (0.994/0.969): downstream of the above, depend on STRING version + multivariate model details → not attempted.
- Wet-lab / manual steps: none (fully computational paper); n/a.
Compute plan
All heavy compute on «our HPC» (SLURM, conda-in-job). R env: r-base + bioconductor-deseq2/limma/geoquery + r-proc + r-edger. Data fetched inside the compute node (has internet) onto «infra»; only small result numbers returned.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Solid partial reproduction with explainable deviations. Two headline outputs reproduced near-exactly — RA DEG counts (2166/1444/722 vs 2177/1458/719, ~1%) and the MDD hub-gene ROC AUCs (EAF1/SDCBP/RNF19B matching 0.911/0.895/0.907 to three decimals) — which strongly argue the core EAF1/SDCBP/RNF19B pipeline is genuine and not fabricated. The divergences sit on the input side and on dataset attribution, not in the core statistics: MDD DEG counts (1600/1364 vs 592/441) reflect an unspecified GSE38206 design (timepoints/probe-collapse) that is partly our methodology and partly paper underspecification, and the RA per-gene AUCs (e.g. SDCBP 0.6122) are not derivable from the assigned GSE169082 (n=7 → degenerate AUC=1.0) and require the unstated GSE97476. Severity is moderate and the central diagnostic-gene conclusion holds in limited form (confirmed on the MDD side, unverified on the RA side), warranting an overall yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.