Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Integrated Analysis of Multiple Microarrays Based on Raw Data Identified Novel Gene Signatures in Recurrent Implantation Failure.

Front Endocrinol (Lausanne) · 2022
L1 85/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • The central claim held under reproduction
What did not (or only partly)
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced the pipeline-derived results of the RIF microarray meta-analysis from the repo's shipped data.Rda + cytohubba.csv on «our HPC» (R 4.5.3, RobustRankAggreg). 6 of 7 in-scope numeric anchors reproduce exactly or within tolerance: RRA robust-DEG counts 1532 (Score<0.05) and 438 (Score<0.01) are exact; the 18 key hub genes match identically (and contain all 10 diagnostic-model genes); the GSE111974 diagnostic model reproduces accuracy 0.85, sensitivity 0.889 and AUC 0.980 exactly/within-tol. The single mismatch is the reported validation specificity of 100% vs reproduced 81.8% (2 control false-positives) -- which is also arithmetically inconsistent with the paper's own reported accuracy 85% + sensitivity 88.9%, so it is flagged as a probable reporting error of one metric rather than fabrication of the model (the model's three other metrics reproduce bit-for-bit). Out of scope: per-GSE raw->DEG reprocessing (limma code not shipped, drop_reason=no_code) and GO/KEGG enrichment (deferred under 80/20).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 85
    assessed: 2026-06-14 ⛓ ae6542bf0bbb
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can integrated analysis of multiple endometrial microarray datasets via Robust Rank Aggregation identify robust differentially expressed genes and hub genes associated with recurrent implantation failure (RIF) that serve as diagnostic biomarkers?

Core claims
  • 1532 robust DEGs in RIF endometrium were identified by integrating four GEO microarray datasets using Robust Rank Aggregation finding
  • 18 hub genes (HMGCS1, SQLE, ESR1, LAMC1, HOXB4, PIP5K1B, GNG11, GPX3, PAX2, TF, ALDH6A1, IDH1, SALL1, EYA1, TAGLN, TPD52L1, ST6GALNAC1, NNMT) were identified via PPI network and Cytohubba finding
  • A 10-hub-gene model (SQLE, LAMC1, HOXB4, PIP5K1B, PAX2, ALDH6A1, SALL1, EYA1, TAGLN, ST6GALNAC1) predicts RIF with 85% accuracy, 100% specificity, 88.9% sensitivity finding
  • Robust DEGs are mainly enriched in extracellular matrix remodeling, adhesion, coagulation, and immunity processes finding
  • Integrating raw data from multiple microarrays via Robust Rank Aggregation yields more stable/robust biomarkers than single studies method
  • 10 of the 18 hub genes were significantly differentially expressed in RIF patients as validated in independent dataset GSE111974 finding
  • Github repository with R scripts and Rdata provides a reproducible resource for the RIF meta-analysis resource
Experimental setups
Assay System Perturbation Readout Platform
Gene expression microarray (Affymetrix) Human endometrium during window of implantation, RIF patients vs fertile controls (GSE26787, 5 control/5 RIF) none (disease vs control comparison) differentially expressed genes Affymetrix GPL570
Gene expression microarray (Illumina) Human endometrium during WOI, RIF vs fertile controls (GSE92324, 8 control/12 RIF) none differentially expressed genes Illumina GPL10558
Gene expression microarray (Agilent) Human endometrium during WOI, RIF vs fertile controls (GSE58144, 72 control/43 RIF) none differentially expressed genes Agilent GPL15789
Gene expression microarray (Agilent) Human endometrium during WOI, RIF vs fertile controls (GSE71331, 5 control/7 RIF, China) none differentially expressed genes Agilent GPL19072
Gene expression microarray (Agilent) used for validation and diagnostic modeling Human endometrium during WOI, 24 RIF/24 fertile controls (GSE111974) none hub gene expression validation and RIF prediction (ROC/AUC) Agilent GPL17077
Robust Rank Aggregation integrative meta-analysis Four GEO microarray datasets (91 RIF/114 control total across five sets) none robust DEGs (adjusted p<0.05) RobustRankAggregation R package
Protein-protein interaction network analysis with hub gene screening 438 robust DEGs (p<0.01) none PPI network nodes/edges and hub genes STRING database / Cytoscape 3.8.2 / CytoHubba (MCC, DMNC, EPC, Degree)
GO and KEGG pathway enrichment analysis 1532 robust DEGs none enriched biological processes, molecular functions, cell components, pathways clusterProfiler R package
Key results
  • 1532 robust DEGs identified at adjusted p<0.05; 438 robust DEGs at adjusted p<0.01 1532 (adj p<0.05); 438 (adj p<0.01)
  • 18 hub genes determined from overlap of CytoHubba top-100 (four methods) and top-100 robust DEGs 18 genes
  • 10-hub-gene model predicts RIF on validation set accuracy 85%, specificity 100%, sensitivity 88.9%
  • 10 of 18 hub genes significantly differentially expressed in RIF as validated by GSE111974 10 of 18
  • PPI network comprised 160 nodes and 174 edges 160 nodes, 174 edges
  • DEGs per dataset: GSE26787 121 up/156 down; GSE92324 343 up/179 down; GSE58144 1045 up/1168 down; GSE71331 128 up/46 down see counts
  • Top enriched KEGG pathways included complement and coagulation cascades, NF-kappa B, PI3K-Akt, ECM-receptor interaction, TNF signaling
Key statistics
  • count 1532 robust DEGs (adjusted p<0.05) (RRA integrative analysis of four datasets)
  • count 438 robust DEGs (adjusted p<0.01) (used for PPI network construction)
  • other accuracy 85% (10-hub-gene RIF prediction model on validation set)
  • other specificity 100% (10-hub-gene RIF prediction model)
  • other sensitivity 88.9% (10-hub-gene RIF prediction model)
  • count 91 RIF patients and 114 control patients (total across five included GSE datasets)
  • count 160 nodes and 174 edges (PPI network of 438 robust DEGs)
  • other ~15% (proportion of IVF-ET patients experiencing RIF)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper integrated raw endometrial microarray data from four GEO datasets (~157 total samples) using Robust Rank Aggregation (RRA) to identify robust differentially expressed genes (DEGs) in recurrent implantation failure (RIF) vs. fertile controls. Individual-dataset DEGs were first screened with limma (p<0.05, |log2FC|>1), ranked, then aggregated via RRA (adjusted p<0.05/0.01). Hub genes were identified by intersecting top-100 nodes from four CytoHubba network-centrality metrics with the top-100 RRA DEGs. An independent dataset (GSE111974, n=48) served for hub-gene validation (limma, p<0.05) and for building a generalized multivariate regression prediction model, evaluated by accuracy, sensitivity, specificity, and AUC.

Replicationbiological Sample sizeSample sizes stated per dataset in Table 1 (range 10–115 per dataset); no formal a priori power calculation reported GroupsRIF patients vs. fertile controls in endometrial biopsies collected at the window of implantation Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionAdjusted p-value in RRA output (specific correction method not named; Benjamini-Hochberg FDR is the RobustRankAggregation package default); individual-dataset limma steps used unadjusted p<0.05; clusterProfiler enrichment applies FDR correction by default
Statistical tests used
Test Applied to n Assumptions
limma moderated t-test (empirical Bayes) DEG screening in each of the four RRA-input datasets (GSE26787, GSE92324, GSE58144, GSE71331) and in the validation dataset (GSE111974) 10, 20, 115, 12 samples per RRA dataset; 48 samples in validation dataset not stated
Robust Rank Aggregation (RRA) — order-statistic / beta-distribution-based non-parametric rank aggregation Integration of ranked DEG lists across four GSE datasets to produce robust DEGs 4 datasets na
Generalized multivariate regression Diagnostic prediction model built from differentially expressed hub genes in the GSE111974 training set ~29 (60% of 48 samples in GSE111974) not stated
ROC / AUC analysis Assessment of the hub-gene prediction model in the GSE111974 validation set ~19 (40% of 48 samples in GSE111974) na
Student's t-test Comparison of normally distributed continuous variables (stated in Statistical Analysis section; specific application not described further in available text) stated
Mann-Whitney U test Comparison of non-normally distributed continuous variables (stated in Statistical Analysis section; specific application not described further in available text) stated
Approaches that could also have been used
  • Individual-dataset DEGs were selected using a nominal p<0.05 threshold (no within-dataset multiple-testing correction) before being fed into RRA
    Could also: Apply Benjamini-Hochberg FDR correction within each dataset (e.g., adjusted p<0.05) prior to rank aggregation — Per-dataset FDR control would reduce false-positive genes entering the ranked lists; RRA is robust to noisy input, but reporting results under both threshold strategies would let readers assess how input stringency affects the final robust DEG set
  • Cross-dataset integration was performed via rank aggregation (RRA), which uses only the rank order of genes without modeling between-study heterogeneity
    Could also: Use a fixed-effects or random-effects meta-analysis of log-fold changes (e.g., via R packages metafor or GeneMeta) or Fisher's combined p-value method — Effect-size-based meta-analysis explicitly estimates between-study variance (tau²) and produces pooled fold-change estimates with confidence intervals, which can complement rank-based approaches and quantify study heterogeneity
  • Hub genes were identified by intersecting the top 100 nodes from four CytoHubba network-centrality metrics with the top 100 RRA DEGs — a hard-cutoff intersection approach
    Could also: Use weighted gene co-expression network analysis (WGCNA) to define co-expression modules from the pooled expression data and identify intramodular hub genes by module membership score — WGCNA derives network structure from sample-level expression correlations rather than curated protein interactions, providing a complementary data-driven perspective; combining both approaches can increase confidence in hub-gene calls
  • A generalized multivariate regression model was used for hub-gene-based RIF classification
    Could also: Apply LASSO-penalized logistic regression, a random forest, or a support vector machine — With a small validation set (~19 samples) and multiple correlated predictors, regularized or ensemble classifiers handle multicollinearity more explicitly; LASSO also performs built-in feature selection and can further narrow the hub-gene panel
  • AUC and diagnostic performance metrics (accuracy, sensitivity, specificity) were reported as point estimates without confidence intervals
    Could also: Report bootstrapped 95% confidence intervals for AUC and for sensitivity/specificity — With a validation set of approximately 19 samples, point estimates carry substantial uncertainty; confidence intervals would allow readers to gauge the precision and clinical relevance of the reported diagnostic performance
  • The train/test split of GSE111974 was a single random partition (60/40, seed=1234)
    Could also: Use leave-one-out cross-validation (LOOCV) or repeated k-fold cross-validation across the full GSE111974 dataset — With only 48 samples, a single split yields a very small validation set (~19); cross-validation uses all samples for both training and testing in turn, producing more stable and less partition-dependent performance estimates
Software: R (base environment) 4.1.0 · R/limma · R/RobustRankAggregation · R/ArrayQualityMetrics · R/affy and R/vsn (normalization) · R/clusterProfiler (GO/KEGG enrichment) · Cytoscape / CytoHubba plugin Cytoscape 3.8.2; CytoHubba version not stated · STRING database (PPI network)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
20
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.
Cited by (assessed papers) (1)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-35197930

Paper: Zeng H, Liu X, Zhang Y et al. (2022) Integrated Analysis of Multiple Microarrays Based on Raw Data Identified Novel Gene Signatures in Recurrent Implantation Failure. Front Endocrinol (Lausanne) 13:785462. PMID 35197930 / PMC8859149 / DOI 10.3389/fendo.2022.785462. Repo: https://github.com/minizenghong/Genetic-meta-analysis-on-RIF (HEAD 5b6838c). Single R-Markdown analysis RIFmeta.Rmd + shipped data.Rda (17 MB: per-GSE DEG tables + GSE111974 expression & targets) + cytohubba.csv (Cytoscape cytoHubba node scores).

Rule 2 — code provenance

The repo is the authors' own analysis code (author field of RIFmeta.Rmd is "Hong Zeng", the paper's first author / corresponding author). Per the operator clarification, a third-party tool on the paper's data would have been equally valid; here it is the authors' own pipeline.

Pipeline inventory (what produces the reported numbers)

The study integrates 5 endometrial microarray datasets (GSE26787, GSE58144, GSE92324, GSE71331 for discovery; GSE111974 for validation):

  1. Per-GSE differential expression (raw-data reprocessing → limma) → the GSE*_res_DEG tables (symbol, logFC, P.Value). Code NOT shipped — only the resulting DEG tables are stored in data.Rda.
  2. Robust Rank Aggregation (RobustRankAggreg::aggregateRanks, RRA, exact) over the 4 discovery DEG lists (up/down separately) → robust DEGs, filtered at Score < 0.05 and Score < 0.01. IN SCOPE.
  3. PPI / cytoHubba hub-gene selection: intersect top-100 of 4 cytoHubba methods (MCC, DMNC, Degree, EPC) with top-100 robust DEGs → key hub genes. cytoHubba scores shipped in cytohubba.csv; STRING/Cytoscape network build is external GUI work (not scripted). IN SCOPE (re-run the intersection on the shipped scores).
  4. GO/KEGG enrichment (clusterProfiler + org.Hs.eg.db). Stretch / 20%.
  5. Diagnostic logistic model on GSE111974 (10 hub genes, set.seed(1234), 60/40 train/validate split, glm + pROC ROC) → accuracy/spec/sens + AUC. IN SCOPE — fully deterministic from shipped data + code.

IN SCOPE (public data, open-source toolchain, deterministic, tractable)

  • C1 — RRA robust-DEG counts: Score<0.05 (reported 1532) and Score<0.01 (reported 438). Pure function of shipped DEG tables + RobustRankAggreg.
  • C2 — key hub genes: the cytoHubba ∩ robust-DEG intersection (reported 18 hub genes; 10 carried into the model). Pure function of cytohubba.csv + shipped DEG tables.
  • C3 — diagnostic model validation metrics: accuracy (reported 85%), specificity (100%), sensitivity (88.9%), AUC (0.980) on GSE111974. Deterministic via set.seed(1234).

OUT OF SCOPE (recorded, not attempted — with reason)

  • Per-GSE raw→DEG reprocessing and the reported individual up/down counts (e.g. GSE26787 121 up / 156 down): the limma/raw-data step is not shipped as code (only the result tables live in data.Rda). The headline "based on raw data" reprocessing cannot be re-run from the repo → drop_reason = no_code for that sub-step. (The DEG tables themselves are used downstream, in scope.)
  • GO/KEGG enrichment (clusterProfiler/org.Hs.eg.db): heavier Bioconductor install; results are qualitative term lists, not a single pinnable number. Deliberately deferred under 80/20 (the heavy ~20%); may be attempted as a stretch if the core anchors land.
  • Volcano plots, heatmap, violin plots, Venn diagram: figure rendering only, no distinct reported numeric value to check 1:1.

Honest notes on the anchors

  • The Rmd's rownames_to_column("SYMBOL") runs after a bind_rows that drops rownames, so its SYMBOL column is integer indices, not gene symbols. The counts (C1) are unaffected (they are nrow() of the filtered frame). For the hub-gene comparison (C2) we use the gene-Name column (the evident intent) and record the quirk in AUDIT.md.
  • The model
Figures / tables: figureTable
C1a
Reported
1532 robust DEGs (RRA Score<0.05)
Reproduced
1532
exact
C1b
Reported
438 robust DEGs (RRA Score<0.01)
Reproduced
438
exact
C2
Reported
18 key hub genes (HMGCS1,SQLE,ESR1,LAMC1,HOXB4,PIP5K1B,GNG11,GPX3,PAX2,TF,ALDH6A1,IDH1,SALL1,EYA1,TAGLN,TPD52L1,ST6GALNAC1,NNMT)
Reproduced
identical 18/18 set; all 10 model genes present
exact
C3a
Reported
validation accuracy 85%
Reproduced
0.85 (17/20)
exact
C3b
Reported
validation specificity 100%
Reproduced
0.818 (TN9/FP2)
did not match
C3c
Reported
validation sensitivity 88.9%
Reproduced
0.8889 (8/9)
exact
C3d
Reported
validation ROC AUC 0.980
Reproduced
0.9798
within tolerance
X1
Reported
per-GSE raw->DEG up/down counts
Reproduced
not attempted
not assessable
X2
Reported
GO/KEGG enrichment terms
Reproduced
not attempted
not assessable

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

Reproduced from the authors' own GitHub repo (shipped data.Rda+cytohubba.csv); 6 of 7 in-scope anchors reproduce exactly or within tolerance — RRA robust-DEG counts (1532/438), the identical 18-gene hub set, and the diagnostic model's accuracy (0.85), sensitivity (0.889) and AUC (0.980). The lone discrepancy is validation specificity: reported 100% vs reproduced 81.8% (FP=2), which is on the authors' side — it is not derivable from the data and is arithmetically inconsistent with the paper's own accuracy+sensitivity, pointing to a reporting/transcription error of one figure rather than model fabrication. The central conclusion (gene signatures + diagnostic model) holds fully, so overall this is a strong reproduction with one flagged, explainable metric error (yellow).

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

158.9 k
tokens (I/O) · 11.2 M incl. cache
16 min
runtime · 0 CPU-h
0.3 GB
peak RAM
1
HPC jobs
hummel
machine