Integrated Analysis of Multiple Microarrays Based on Raw Data Identified Novel Gene Signatures in Recurrent Implantation Failure.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓The central claim held under reproduction
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced the pipeline-derived results of the RIF microarray meta-analysis from the repo's shipped data.Rda + cytohubba.csv on «our HPC» (R 4.5.3, RobustRankAggreg). 6 of 7 in-scope numeric anchors reproduce exactly or within tolerance: RRA robust-DEG counts 1532 (Score<0.05) and 438 (Score<0.01) are exact; the 18 key hub genes match identically (and contain all 10 diagnostic-model genes); the GSE111974 diagnostic model reproduces accuracy 0.85, sensitivity 0.889 and AUC 0.980 exactly/within-tol. The single mismatch is the reported validation specificity of 100% vs reproduced 81.8% (2 control false-positives) -- which is also arithmetically inconsistent with the paper's own reported accuracy 85% + sensitivity 88.9%, so it is flagged as a probable reporting error of one metric rather than fabrication of the model (the model's three other metrics reproduce bit-for-bit). Out of scope: per-GSE raw->DEG reprocessing (limma code not shipped, drop_reason=no_code) and GO/KEGG enrichment (deferred under 80/20).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 85assessed: 2026-06-14 ⛓ ae6542bf0bbb
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan integrated analysis of multiple endometrial microarray datasets via Robust Rank Aggregation identify robust differentially expressed genes and hub genes associated with recurrent implantation failure (RIF) that serve as diagnostic biomarkers?
- ★ 1532 robust DEGs in RIF endometrium were identified by integrating four GEO microarray datasets using Robust Rank Aggregation finding
- ★ 18 hub genes (HMGCS1, SQLE, ESR1, LAMC1, HOXB4, PIP5K1B, GNG11, GPX3, PAX2, TF, ALDH6A1, IDH1, SALL1, EYA1, TAGLN, TPD52L1, ST6GALNAC1, NNMT) were identified via PPI network and Cytohubba finding
- ★ A 10-hub-gene model (SQLE, LAMC1, HOXB4, PIP5K1B, PAX2, ALDH6A1, SALL1, EYA1, TAGLN, ST6GALNAC1) predicts RIF with 85% accuracy, 100% specificity, 88.9% sensitivity finding
- ★ Robust DEGs are mainly enriched in extracellular matrix remodeling, adhesion, coagulation, and immunity processes finding
- ★ Integrating raw data from multiple microarrays via Robust Rank Aggregation yields more stable/robust biomarkers than single studies method
- ★ 10 of the 18 hub genes were significantly differentially expressed in RIF patients as validated in independent dataset GSE111974 finding
- Github repository with R scripts and Rdata provides a reproducible resource for the RIF meta-analysis resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Gene expression microarray (Affymetrix) | Human endometrium during window of implantation, RIF patients vs fertile controls (GSE26787, 5 control/5 RIF) | none (disease vs control comparison) | differentially expressed genes | Affymetrix GPL570 |
| Gene expression microarray (Illumina) | Human endometrium during WOI, RIF vs fertile controls (GSE92324, 8 control/12 RIF) | none | differentially expressed genes | Illumina GPL10558 |
| Gene expression microarray (Agilent) | Human endometrium during WOI, RIF vs fertile controls (GSE58144, 72 control/43 RIF) | none | differentially expressed genes | Agilent GPL15789 |
| Gene expression microarray (Agilent) | Human endometrium during WOI, RIF vs fertile controls (GSE71331, 5 control/7 RIF, China) | none | differentially expressed genes | Agilent GPL19072 |
| Gene expression microarray (Agilent) used for validation and diagnostic modeling | Human endometrium during WOI, 24 RIF/24 fertile controls (GSE111974) | none | hub gene expression validation and RIF prediction (ROC/AUC) | Agilent GPL17077 |
| Robust Rank Aggregation integrative meta-analysis | Four GEO microarray datasets (91 RIF/114 control total across five sets) | none | robust DEGs (adjusted p<0.05) | RobustRankAggregation R package |
| Protein-protein interaction network analysis with hub gene screening | 438 robust DEGs (p<0.01) | none | PPI network nodes/edges and hub genes | STRING database / Cytoscape 3.8.2 / CytoHubba (MCC, DMNC, EPC, Degree) |
| GO and KEGG pathway enrichment analysis | 1532 robust DEGs | none | enriched biological processes, molecular functions, cell components, pathways | clusterProfiler R package |
- – 1532 robust DEGs identified at adjusted p<0.05; 438 robust DEGs at adjusted p<0.01 1532 (adj p<0.05); 438 (adj p<0.01)
- – 18 hub genes determined from overlap of CytoHubba top-100 (four methods) and top-100 robust DEGs 18 genes
- – 10-hub-gene model predicts RIF on validation set accuracy 85%, specificity 100%, sensitivity 88.9%
- – 10 of 18 hub genes significantly differentially expressed in RIF as validated by GSE111974 10 of 18
- – PPI network comprised 160 nodes and 174 edges 160 nodes, 174 edges
- – DEGs per dataset: GSE26787 121 up/156 down; GSE92324 343 up/179 down; GSE58144 1045 up/1168 down; GSE71331 128 up/46 down see counts
- – Top enriched KEGG pathways included complement and coagulation cascades, NF-kappa B, PI3K-Akt, ECM-receptor interaction, TNF signaling
- count 1532 robust DEGs (adjusted p<0.05) (RRA integrative analysis of four datasets)
- count 438 robust DEGs (adjusted p<0.01) (used for PPI network construction)
- other accuracy 85% (10-hub-gene RIF prediction model on validation set)
- other specificity 100% (10-hub-gene RIF prediction model)
- other sensitivity 88.9% (10-hub-gene RIF prediction model)
- count 91 RIF patients and 114 control patients (total across five included GSE datasets)
- count 160 nodes and 174 edges (PPI network of 438 robust DEGs)
- other ~15% (proportion of IVF-ET patients experiencing RIF)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper integrated raw endometrial microarray data from four GEO datasets (~157 total samples) using Robust Rank Aggregation (RRA) to identify robust differentially expressed genes (DEGs) in recurrent implantation failure (RIF) vs. fertile controls. Individual-dataset DEGs were first screened with limma (p<0.05, |log2FC|>1), ranked, then aggregated via RRA (adjusted p<0.05/0.01). Hub genes were identified by intersecting top-100 nodes from four CytoHubba network-centrality metrics with the top-100 RRA DEGs. An independent dataset (GSE111974, n=48) served for hub-gene validation (limma, p<0.05) and for building a generalized multivariate regression prediction model, evaluated by accuracy, sensitivity, specificity, and AUC.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma moderated t-test (empirical Bayes) | DEG screening in each of the four RRA-input datasets (GSE26787, GSE92324, GSE58144, GSE71331) and in the validation dataset (GSE111974) | 10, 20, 115, 12 samples per RRA dataset; 48 samples in validation dataset | not stated |
| Robust Rank Aggregation (RRA) — order-statistic / beta-distribution-based non-parametric rank aggregation | Integration of ranked DEG lists across four GSE datasets to produce robust DEGs | 4 datasets | na |
| Generalized multivariate regression | Diagnostic prediction model built from differentially expressed hub genes in the GSE111974 training set | ~29 (60% of 48 samples in GSE111974) | not stated |
| ROC / AUC analysis | Assessment of the hub-gene prediction model in the GSE111974 validation set | ~19 (40% of 48 samples in GSE111974) | na |
| Student's t-test | Comparison of normally distributed continuous variables (stated in Statistical Analysis section; specific application not described further in available text) | — | stated |
| Mann-Whitney U test | Comparison of non-normally distributed continuous variables (stated in Statistical Analysis section; specific application not described further in available text) | — | stated |
-
Individual-dataset DEGs were selected using a nominal p<0.05 threshold (no within-dataset multiple-testing correction) before being fed into RRA↳ Could also: Apply Benjamini-Hochberg FDR correction within each dataset (e.g., adjusted p<0.05) prior to rank aggregation — Per-dataset FDR control would reduce false-positive genes entering the ranked lists; RRA is robust to noisy input, but reporting results under both threshold strategies would let readers assess how input stringency affects the final robust DEG set
-
Cross-dataset integration was performed via rank aggregation (RRA), which uses only the rank order of genes without modeling between-study heterogeneity↳ Could also: Use a fixed-effects or random-effects meta-analysis of log-fold changes (e.g., via R packages metafor or GeneMeta) or Fisher's combined p-value method — Effect-size-based meta-analysis explicitly estimates between-study variance (tau²) and produces pooled fold-change estimates with confidence intervals, which can complement rank-based approaches and quantify study heterogeneity
-
Hub genes were identified by intersecting the top 100 nodes from four CytoHubba network-centrality metrics with the top 100 RRA DEGs — a hard-cutoff intersection approach↳ Could also: Use weighted gene co-expression network analysis (WGCNA) to define co-expression modules from the pooled expression data and identify intramodular hub genes by module membership score — WGCNA derives network structure from sample-level expression correlations rather than curated protein interactions, providing a complementary data-driven perspective; combining both approaches can increase confidence in hub-gene calls
-
A generalized multivariate regression model was used for hub-gene-based RIF classification↳ Could also: Apply LASSO-penalized logistic regression, a random forest, or a support vector machine — With a small validation set (~19 samples) and multiple correlated predictors, regularized or ensemble classifiers handle multicollinearity more explicitly; LASSO also performs built-in feature selection and can further narrow the hub-gene panel
-
AUC and diagnostic performance metrics (accuracy, sensitivity, specificity) were reported as point estimates without confidence intervals↳ Could also: Report bootstrapped 95% confidence intervals for AUC and for sensitivity/specificity — With a validation set of approximately 19 samples, point estimates carry substantial uncertainty; confidence intervals would allow readers to gauge the precision and clinical relevance of the reported diagnostic performance
-
The train/test split of GSE111974 was a single random partition (60/40, seed=1234)↳ Could also: Use leave-one-out cross-validation (LOOCV) or repeated k-fold cross-validation across the full GSE111974 dataset — With only 48 samples, a single split yields a very small validation set (~19); cross-validation uses all samples for both training and testing in turn, producing more stable and less partition-dependent performance estimates
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
scope.md — pmid-35197930
Paper: Zeng H, Liu X, Zhang Y et al. (2022) Integrated Analysis of Multiple
Microarrays Based on Raw Data Identified Novel Gene Signatures in Recurrent
Implantation Failure. Front Endocrinol (Lausanne) 13:785462. PMID 35197930 /
PMC8859149 / DOI 10.3389/fendo.2022.785462.
Repo: https://github.com/minizenghong/Genetic-meta-analysis-on-RIF
(HEAD 5b6838c). Single R-Markdown analysis RIFmeta.Rmd + shipped data.Rda
(17 MB: per-GSE DEG tables + GSE111974 expression & targets) + cytohubba.csv
(Cytoscape cytoHubba node scores).
Rule 2 — code provenance
The repo is the authors' own analysis code (author field of RIFmeta.Rmd is
"Hong Zeng", the paper's first author / corresponding author). Per the operator
clarification, a third-party tool on the paper's data would have been equally
valid; here it is the authors' own pipeline.
Pipeline inventory (what produces the reported numbers)
The study integrates 5 endometrial microarray datasets (GSE26787, GSE58144, GSE92324, GSE71331 for discovery; GSE111974 for validation):
- Per-GSE differential expression (raw-data reprocessing → limma) → the
GSE*_res_DEGtables (symbol, logFC, P.Value). Code NOT shipped — only the resulting DEG tables are stored indata.Rda. - Robust Rank Aggregation (
RobustRankAggreg::aggregateRanks, RRA, exact) over the 4 discovery DEG lists (up/down separately) → robust DEGs, filtered at Score < 0.05 and Score < 0.01. IN SCOPE. - PPI / cytoHubba hub-gene selection: intersect top-100 of 4 cytoHubba
methods (MCC, DMNC, Degree, EPC) with top-100 robust DEGs → key hub genes.
cytoHubba scores shipped in
cytohubba.csv; STRING/Cytoscape network build is external GUI work (not scripted). IN SCOPE (re-run the intersection on the shipped scores). - GO/KEGG enrichment (
clusterProfiler+org.Hs.eg.db). Stretch / 20%. - Diagnostic logistic model on GSE111974 (10 hub genes,
set.seed(1234), 60/40 train/validate split, glm + pROC ROC) → accuracy/spec/sens + AUC. IN SCOPE — fully deterministic from shipped data + code.
IN SCOPE (public data, open-source toolchain, deterministic, tractable)
- C1 — RRA robust-DEG counts: Score<0.05 (reported 1532) and Score<0.01 (reported 438). Pure function of shipped DEG tables + RobustRankAggreg.
- C2 — key hub genes: the cytoHubba ∩ robust-DEG intersection (reported 18
hub genes; 10 carried into the model). Pure function of
cytohubba.csv+ shipped DEG tables. - C3 — diagnostic model validation metrics: accuracy (reported 85%),
specificity (100%), sensitivity (88.9%), AUC (0.980) on GSE111974.
Deterministic via
set.seed(1234).
OUT OF SCOPE (recorded, not attempted — with reason)
- Per-GSE raw→DEG reprocessing and the reported individual up/down counts
(e.g. GSE26787 121 up / 156 down): the limma/raw-data step is not shipped as
code (only the result tables live in
data.Rda). The headline "based on raw data" reprocessing cannot be re-run from the repo → drop_reason =no_codefor that sub-step. (The DEG tables themselves are used downstream, in scope.) - GO/KEGG enrichment (
clusterProfiler/org.Hs.eg.db): heavier Bioconductor install; results are qualitative term lists, not a single pinnable number. Deliberately deferred under 80/20 (the heavy ~20%); may be attempted as a stretch if the core anchors land. - Volcano plots, heatmap, violin plots, Venn diagram: figure rendering only, no distinct reported numeric value to check 1:1.
Honest notes on the anchors
- The Rmd's
rownames_to_column("SYMBOL")runs after abind_rowsthat drops rownames, so itsSYMBOLcolumn is integer indices, not gene symbols. The counts (C1) are unaffected (they arenrow()of the filtered frame). For the hub-gene comparison (C2) we use the gene-Namecolumn (the evident intent) and record the quirk in AUDIT.md. - The model
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Reproduced from the authors' own GitHub repo (shipped data.Rda+cytohubba.csv); 6 of 7 in-scope anchors reproduce exactly or within tolerance — RRA robust-DEG counts (1532/438), the identical 18-gene hub set, and the diagnostic model's accuracy (0.85), sensitivity (0.889) and AUC (0.980). The lone discrepancy is validation specificity: reported 100% vs reproduced 81.8% (FP=2), which is on the authors' side — it is not derivable from the data and is arithmetically inconsistent with the paper's own accuracy+sensitivity, pointing to a reporting/transcription error of one figure rather than model fabrication. The central conclusion (gene signatures + diagnostic model) holds fully, so overall this is a strong reproduction with one flagged, explainable metric error (yellow).
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.