Identification and verification of diagnostic biomarkers in recurrent pregnancy loss via machine learning algorithm and WGCNA.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH only in part. Authors' own repo (Weihouye/bioinformatics-analysis @ e8a6385) is a set of biowolf.cn template R scripts; data is GEO GSE165004 (Agilent GPL16699, 48 of 72 samples = 24 RPL + 24 control). PARTIAL 1:1: the diagnostic HEADLINE reproduces EXACTLY - individual hub-gene ROC AUCs on the test set (WBP11 1.00, ACTR2 0.99, NCSTN 0.99) and hub-gene expression directions (WBP11/ACTR2 down, NCSTN up) match to reported precision; these are deterministic from the expression matrix, so the biomarkers are genuine strong DEGs. DEG count reproduces within ~18% (416 vs 352, same limma + cutoffs). Validation set GSE183555 (RNA-seq, CPM, n=5+5): ACTR2 0.60 exact, WBP11 0.80 vs 0.60, NCSTN 0.64 vs 0.52 - all modest, same weak-validation conclusion. WHAT DID NOT REPRODUCE (different, not 1:1): the WGCNA turquoise module (reported 1990 genes, r=-0.42, p=0.003, 104 DEGs) is NOT regenerable from the deposited code+data under ANY gene-filter (6-way sweep) - the WGCNA variance filter is commented out in WGCNA.R, the disease module is always weaker (|r|~0.31) and differently sized, and no 'turquoise' module forms; and the LASSO/SVM-RFE/RF feature selection (reported 3/6/7 genes overlapping to the 3 hub genes) does not reproduce because the scripts set no random seed and depend on the non-reproducible module (my LASSO->16 genes, RF->2, overlap->0). NOT ATTEMPTED: RT-qPCR (wet-lab), CMap drug screening + pan-cancer Sangerbox/UCSC Xena (external web tools, no shipped code), CIBERSORT immune infiltration (needs LM22 ref.txt not in repo), ANN combined-model ROC (training randomness + depends on non-reproducible selection). Possible-fabrication: LOW concern at biomarker level (AUCs+directions exact); the WGCNA module numbers are not independently verifiable -> most consistent with under-specification rather than fabrication, flagged for human review.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 55assessed: 2026-06-16 ⛓ 76d17d12a196
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe molecular mechanisms underlying recurrent pregnancy loss (RPL) remain unclear, so this study aimed to identify diagnostic biomarkers and elucidate molecular mechanisms of RPL using WGCNA combined with machine learning algorithms applied to GEO gene expression profiles.
- ★ 352 DEGs (198 up-regulated, 154 down-regulated) were identified between RPL and control endometrial samples finding
- ★ WBP11, ACTR2, and NCSTN are hub diagnostic biomarkers for RPL, identified via intersection of LASSO, SVM-RFE, and RandomForest algorithms on WGCNA turquoise module DEGs finding
- ★ Combined ROC analysis of the three hub genes demonstrates strong diagnostic value for RPL finding
- ★ RT-qPCR validation confirmed low expression of WBP11 and ACTR2, and high expression of NCSTN, in RPL decidua samples finding
- ★ Immune cell infiltration analysis via CIBERSORT revealed an imbalance of macrophages in RPL finding
- The three hub genes show aberrant expression and are associated with poor prognosis across multiple malignancies in pan-cancer analysis finding
- Several small-molecule drugs were identified via CMap as potential RPL treatments resource
- ★ DEGs in RPL are enriched in Fc gamma R-mediated phagocytosis, Fc epsilon RI signaling, and lipid/purine metabolism pathways mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| microarray gene expression profiling / differential expression analysis | human endometrial tissue (GSE165004) | none (RPL vs healthy fertile controls) | differentially expressed genes (log2FC, p-value) | Agilent-039494 SurePrint G3 Human GE v2 8x60K microarray (GPL16699) |
| WGCNA co-expression network analysis | human endometrial tissue (GSE165004, 20,552 genes) | none | gene modules correlated with RPL clinical trait | — |
| machine learning feature selection (LASSO, SVM-RFE, RandomForest) | 104 DEGs from WGCNA turquoise module | none | hub gene candidates | R packages glmnet, e1071/kernlab/caret, randomForest |
| artificial neural network (ANN) modeling | hub gene expression scores (GSE165004 and GSE183555 validation set) | none | diagnostic accuracy (ROC/AUC) | neuralnet, NeuralNetTools packages |
| RT-qPCR | human decidua tissue, RPL patients (n=20) vs controls (n=20) | none (disease vs control) | relative mRNA expression of WBP11, ACTR2, NCSTN (2-ΔΔCT) | SYBR Green Master Mix (Qiagen); NanoDrop 2000 |
| immune cell infiltration analysis (CIBERSORT) | human endometrial tissue (GSE165004) | none | proportions of 22 immune cell types; correlation with hub genes | CIBERSORT with LM22 signature matrix |
| pan-cancer differential expression and survival analysis | TCGA, TARGET, GTEx pan-cancer datasets (multiple tumor types) | none | hub gene expression differences and overall survival (Cox regression, Kaplan-Meier) | UCSC Xena via Sangerbox |
| small-molecule drug screening | in silico, CMap database | computational query of DEG signature | norm_cs score of candidate compounds | Connectivity Map (CMap, clue.io) |
- – 352 DEGs identified in RPL vs control 198 up / 154 down
- ▼ Turquoise WGCNA module most correlated with RPL, containing 104 DEGs r=-0.42, p=0.003
- – Three algorithms (LASSO, SVM-RFE, RandomForest) converged on WBP11, ACTR2, NCSTN as hub genes
- – RT-qPCR shows WBP11 and ACTR2 down-regulated, NCSTN up-regulated in RPL decidua
- – Combined hub gene ROC curve shows strong diagnostic performance for RPL
- – Macrophage proportions imbalanced between RPL and control immune infiltration
- – KEGG/GSEA enrichment highlights Fc gamma R-mediated phagocytosis, Fc epsilon RI signaling, and metabolic pathways (purine, glycerophospholipid, linoleic/alpha-linolenic acid metabolism)
- – Hub genes show aberrant expression and poor prognosis association across multiple cancers in pan-cancer analysis
- count 352 DEGs (198 up-regulated, 154 down-regulated) (differential expression analysis of GSE165004)
- correlation r=-0.42 (turquoise module correlation with RPL clinical trait)
- pvalue p=0.003 (significance of turquoise module correlation with RPL)
- count 104 DEGs in turquoise module (64 up, 40 down) (WGCNA module gene composition)
- other soft-thresholding power β=13, R2=0.82 (WGCNA scale-free network construction)
- other optimal LASSO lambda=0.0197 (LASSO regression hub gene selection after 10-fold cross-validation)
- count 3 overlapping hub genes (WBP11, ACTR2, NCSTN) from LASSO, SVM-RFE, RandomForest (Venn diagram intersection of machine learning algorithm outputs)
- count 24 RPL and 24 control samples (GSE165004); 5 RPL and 5 control (GSE183555 validation); 20 RPL and 20 control decidua samples (RT-qPCR) (dataset and cohort sample sizes)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This bioinformatics study analyzed public microarray datasets (GSE165004, GSE183555) combined with RT-qPCR validation. Differentially expressed genes were identified with limma, co-expression modules with WGCNA (Pearson correlation to traits), and candidate hub genes were selected by intersecting three machine-learning feature-selection methods (LASSO, SVM-RFE, Random Forest); diagnostic performance was summarized with ROC/AUC from an artificial neural network model. For patient/clinical and validation comparisons, normality was checked by Shapiro-Wilk, with Student's t-test for normal continuous variables and Wilcoxon tests for gene expression, while pan-cancer analyses used unpaired Wilcoxon tests, Cox regression, and log-rank tests. Results were reported as mean ± SD with p-values, using a p<0.05 significance threshold throughout.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Student's t-test | baseline clinical variables (age, gestational weeks, BMI) between RPL and control, Table 2 | 20 RPL vs 20 control | stated |
| Shapiro-Wilk test | normality assessment of continuous variables prior to choosing t-test vs Wilcoxon | — | stated |
| Wilcoxon test (Mann-Whitney) | mRNA expression levels of individual hub genes between RPL and control (non-normal variables) | — | stated |
| limma (moderated/empirical Bayes) differential expression | DEG identification in GSE165004 (|log2FC|>0.585, p<0.05) | 24 RPL vs 24 control | not stated |
| Pearson correlation | WGCNA module-trait relationships (turquoise module r=-0.42, p=0.003) | 48 samples | not stated |
| LASSO / SVM-RFE / Random Forest feature selection | hub-gene screening from 104 turquoise-module DEGs | — | na |
| ROC/AUC analysis (pROC) | diagnostic value of hub genes and ANN model in test and validation sets | — | na |
| Unpaired Wilcoxon test | pan-cancer differential expression of hub genes across TCGA/TARGET/GTEx | — | not stated |
| Cox proportional hazards regression | association between hub-gene expression and prognosis in each tumor | — | not stated |
| Log-rank test | comparison of prognostic significance / Kaplan-Meier overall survival in pan-cancer | — | na |
-
DEGs were called from limma with |log2FC|>0.585 and a raw p-value<0.05 threshold.↳ Could also: An FDR-adjusted p-value (e.g., Benjamini-Hochberg) could also be applied across the genome-wide tests. — With ~20,000 genes tested simultaneously, an FDR-controlled threshold also conveys how many discoveries are expected to be false and is a common companion to raw-p reporting in microarray studies.
-
Baseline group means were summarized as mean ± SD with exact p-values.↳ Could also: A 95% confidence interval for the between-group difference could also be reported alongside the p-value. — A CI on the difference also communicates the magnitude and precision of any group difference, which complements the significance test.
-
Per-gene RT-qPCR expression was compared with Wilcoxon tests after a Shapiro-Wilk normality check.↳ Could also: Reporting the chosen test's exact n and an effect-size measure (e.g., rank-biserial correlation or fold-change with CI) could also accompany each comparison. — Pairing the test with an explicit effect size and n also helps readers gauge the practical magnitude of expression differences, not just their significance.
-
Hub genes were selected by intersecting LASSO, SVM-RFE, and Random Forest on the same dataset, with ROC/AUC reported.↳ Could also: Cross-validated or bootstrap-resampled performance estimates (and an external test split) could also be reported for the combined model. — Resampling-based performance estimates also help characterize how the diagnostic signature might generalize beyond the training data.
-
For genes with multiple probes, expression was collapsed by averaging probe values.↳ Could also: Selecting the maximum-variance or maximum-intensity probe could also be used for probe-to-gene mapping. — Alternative collapsing rules also handle probes with differing hybridization behavior and are common defaults in tools such as WGCNA's collapseRows.
-
Enrichment analyses (GO/KEGG/GSEA) used a p<0.05 cutoff.↳ Could also: An adjusted q-value (FDR) cutoff could also be used to rank enriched terms. — FDR-based ranking also accounts for the many gene-set tests performed and is a widely used convention for prioritizing enriched pathways.
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
21 downstream papers · 2 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Identification of novel biomarkers and immune infilt... 2023 · 13 cites
- Mechanisms of Bushen Tiaoxue Granules against contro... 2024 · 8 cites
- Identification of m6A Modification Regulated by Dysr... 2023 · 5 cites
- Towards reproducible research in recurrent pregnancy... 2023 · 4 cites
- Meta-analysis of endometrial transcriptome data reve... 2024 · 4 cites
- Shared diagnostic genes and potential mechanisms bet... 2024 · 4 cites
- Endometrial gland specific progestagen-associated en... 2022 · 9 cites
- Machine learning for predictive risk stratification... 2026 · 0 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37691920
Paper: Wei et al. 2023, Front Immunol — "Identification and verification of diagnostic biomarkers in recurrent pregnancy loss via machine learning algorithm and WGCNA." DOI 10.3389/fimmu.2023.1241816.
Code: https://github.com/Weihouye/bioinformatics-analysis (single commit
e8a6385c9730e26d2165affb253c3f626febbf0b, 2023-06-20). The authors' own repo —
a set of biowolf.cn tutorial-template R scripts (R_Agilent.R, WGCNA.R,
Diagnostic.R, lasso.R, SVM-RFE.R, randomForest.R, neuralNet.R, geneScore.R, ROC.R,
GO/KEGG/GSEA/CIBERSORT/vioplot/Lollipop). Scripts are fragments with hard-coded
Windows paths and a few broken lines; parameters are explicit and reproducible.
Data: GEO GSE165004 (GPL16699, Agilent-039494 SurePrint G3 Human GE v2 8x60K). The series has 72 samples = 24 Control + 24 RPL + 24 UIF. The paper uses only the 48 RPL + Control samples (drops the 24 UIF). Validation set GSE183555 (GPL21697, 5 RPL + 5 control).
Pipeline-derived results — IN SCOPE
| group | what | params (from code/methods) |
|---|---|---|
| DEG | 352 / 198up / 154down | limma on Agilent raw → normexp bg + quantile norm; |log2FC|>0.585 & p<0.05 |
| WGCNA | soft power β=13 (R²=0.82); turquoise module 1990 genes; r=-0.42, p=0.003 | minModuleSize=60, deepSplit=2, MEDissThres=0.25 |
| intersect | 104 DEGs in turquoise (64 up/40 down) | DEG ∩ turquoise |
| ML | LASSO λ.min=0.0197→3; SVM-RFE→6; RF (imp>1)→7; overlap→3 (WBP11,ACTR2,NCSTN) | glmnet binom α=1 cv10; e1071/caret rfe; randomForest ntree=1000 |
| ROC | individual AUC WBP11 1.00 / ACTR2 0.99 / NCSTN 0.99 (test); ANN 1.00 (test) | pROC roc() + smooth(); neuralnet hidden=5 |
| ROC-val | validation AUCs on GSE183555: WBP11 0.60, ACTR2 0.60, NCSTN 0.52, ANN 0.74 | second dataset (harder ~20%) |
OUT OF SCOPE (wet-lab / external web tools / non-pipeline)
- RT-qPCR validation of hub-gene expression — wet-lab, not computational.
- CMap small-molecule drug screening — external web service, no shipped code.
- Pan-cancer (Sangerbox/UCSC Xena TCGA/GTEx) — external web tool, different data.
- GO/KEGG/GSEA enrichment — derivable but secondary; attempt only after core.
CIBERSORT immune infiltration — borderline
run.R + CIBERSORT.R ship; needs LM22 ref ("ref.txt", not in repo). Attempt last.
Determinism notes
- DEG, soft-power, module sizes, module-trait r, individual-gene ROC AUC are deterministic given the input matrix → primary 1:1 targets.
- LASSO/SVM-RFE/RF/ANN involve CV randomness with no seed in the code → the gene sets and exact λ may vary run-to-run; the 3-gene overlap is the robust headline to check.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The diagnostic core reproduces exactly — WBP11 (1.00), ACTR2 (0.99), NCSTN (0.99) test-set ROC AUCs and hub-gene expression directions match deterministically, and all 7 candidates are genuine DEGs, so the biomarker claim itself holds and is not fabrication-suspect. However, the WGCNA turquoise module (1990 genes, r=-0.42, p=0.003, 104 DEGs) does not form under any gene-filter and the LASSO/SVM-RFE/RF selection that nominated the 3 hubs collapses to a disjoint set (overlap 3→0). The deviations are on the authors' side via under-specification — a commented-out variance filter and missing random seeds in the deposited template scripts — rather than fabrication or our own methodology, plus minor preprocessing drift (DEG counts ~18% high). Net: a solid partial reproduction whose headline holds but whose stated derivation pathway is not independently reproducible.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.