Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Molecular subtype of recurrent implantation failure reveals distinct endometrial etiology of female infertility.

J Transl Med · 2025
L1 68/100 PQI 92
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
68/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 32% of all assessed papers rank 765 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to PARTIALLY reproduce. The MetaRIF repo ships only the FINAL diagnostic/subtyping classifier (MetaRIF.R + a pretrained .rda with two caret models and the 64+19 ENTREZID signature genes); the upstream meta-analysis (1,776 MetaDE DEGs, ConsensusClusterPlus k=2 subtyping, CIBERSORT/GSVA, KEGG/GSEA, METAFlux, cMAP) is NOT in the repo and depends on an integrated 238-sample matrix incl. restricted in-house CNGB data -> out of scope, not attempted (the hard 20%). I ran the shipped classifier on the «infra» via a «our HPC» SLURM job: it loads cleanly, and on the paper's listed public cohort GSE111974 (24 RIF/24 control, GPL17077) the signature coverage is full (64/64, 19/19) and the model reproduces the described two-subtype behaviour — all 24 controls -> Normal (prob 0.95-0.9998), RIF -> 5 RIF1 + 12 RIF2; AUC 0.9965, accuracy 0.872. This is strong evidence the classifier artifact is genuine and functional (no fabrication indicators). HOWEVER the headline independent-validation number, AUC=0.94 on GSE71331, was NOT reproduced: GSE71331's GEO platform GPL19072 (custom Agilent lncRNA+mRNA array) provides probe sequences only with no gene/ENTREZ mapping, so the classifier (which needs ENTREZID rows) cannot be applied without probe-sequence alignment (deliberately skipped per 80/20). The GSE111974 AUC is in-sample (discovery cohort) and has no specific paper value to grade against. Verdict: classifier exists and works 1:1 in behaviour; the specific reported validation AUC could not be independently confirmed from public data. Not a drop, not a mismatch — a scope/annotation-limited partial. (Note: paper writes platform 'GPL9072'; the real GEO record is GPL19072.)

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 68
    assessed: 2026-06-14 ⛓ 211babfe6678
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether biologically distinct molecular subtypes of endometrial dysfunction exist among recurrent implantation failure (RIF) patients, and whether such subtypes could guide more personalized and effective treatment strategies.

Core claims
  • RIF endometrial samples segregate into two reproducible molecular subtypes: an immune-driven subtype (RIF-I) and a metabolic-driven subtype (RIF-M) finding
  • RIF-I is enriched for immune and inflammatory pathways (e.g., IL-17 and TNF signaling) with increased effector immune cell infiltration finding
  • RIF-M is characterized by dysregulated oxidative phosphorylation, fatty acid metabolism, steroid hormone biosynthesis, and altered circadian gene PER1 expression finding
  • T-bet/GATA3 protein expression ratio by IHC mirrors the RIF-I/RIF-M subtype distribution, higher in RIF-I and lower in RIF-M finding
  • MetaRIF, a machine-learning classifier built from 64 algorithm combinations, accurately distinguishes RIF-I/RIF-M subtypes and outperforms previously published RIF signature models method
  • CMap-based drug prediction identifies sirolimus (rapamycin) as a candidate therapeutic for RIF-I and prostaglandins as a candidate for RIF-M resource
  • 1,776 robust differentially expressed genes were identified between RIF and normal endometrial samples via integrated multi-dataset meta-analysis finding
  • Multi-platform transcriptomic datasets can be harmonized using a random-effects model to reduce batch effects for cross-cohort DEG discovery method
Experimental setups
Assay System Perturbation Readout Platform
Meta-analysis of microarray/RNA-seq transcriptomics (MetaDE, random-effects integration) human endometrial tissue, mid-luteal/mid-secretory phase (GSE111974, GSE58144, GSE106602, GSE71331) none (RIF vs healthy control) differentially expressed genes GPL17077, GPL15789, GPL16791, GPL9072 (Agilent/RNA-seq)
RNA sequencing (MARS-seq) in-house human endometrial biopsy cohort (RIF n=12, tubal-factor control n=21) none (RIF vs normal) gene expression profiles MARS-seq, HISAT alignment to hg38
Immune/cell-type deconvolution (CIBERSORT, GSVA) bulk endometrial transcriptomic data none relative proportions of 14 endometrial cell types and immune infiltration scores CIBERSORT; GSVA
Unsupervised consensus clustering RIF-DEGs in training cohort endometrial samples none subtype assignment (RIF-I vs RIF-M) ConsensusClusterPlus (R)
Functional/pathway enrichment (GSEA/GSVA, KEGG) RIF-I vs RIF-M endometrial samples none pathway/module activity scores, metabolic flux clusterProfiler; GSVA; METAFlux
Immunohistochemistry human endometrial tissue sections none T-bet and GATA3 protein expression/ratio Leica BOND RX; Olympus VS200; HALO Image Analysis System
Machine-learning classifier development (8 algorithms, 64 combinations, LOOCV) RIF-associated gene expression, training and validation cohorts (incl. GSE71331 and in-house dataset) none F-score and AUC for subtype classification Random Forest, Lasso, Naive Bayes, KNN, Nnet, LightGBM, AdaBoost.M1, SVM
Connectivity Map (CMap) drug prediction overlapping DEGs (RIF-I and RIF-M gene sets) as query none connectivity score, candidate small-molecule compounds CLUE (CMap and LINCS Unified Environment), Gene expression (L1000)
Key results
  • 1,776 robust DEGs identified between RIF and normal endometrial samples 1,776 DEGs
  • Two reproducible RIF subtypes (RIF-I, RIF-M) identified by consensus clustering
  • RIF-I enriched for IL-17 and TNF signaling pathways with elevated effector immune cell infiltration p < 0.01
  • RIF-M shows dysregulated oxidative phosphorylation, fatty acid metabolism, steroid hormone biosynthesis, and altered PER1 expression
  • T-bet/GATA3 expression ratio higher in RIF-I and lower in RIF-M by IHC
  • MetaRIF classifier distinguished subtypes in independent validation cohorts AUC: 0.94 and 0.85
  • MetaRIF outperformed previously published RIF signature models AUC: MetaRIF=0.88 vs koot_sig=0.48, Wang_sig=0.54, OSR_score=0.72
  • CMap predicted sirolimus as top candidate for RIF-I and prostaglandins as top candidate for RIF-M
Key statistics
  • count 1,776 (robust DEGs identified between RIF and normal samples)
  • pvalue p < 0.01 (enrichment of IL-17 and TNF signaling pathways in RIF-I)
  • other AUC = 0.94 and 0.85 (MetaRIF classifier performance in independent validation cohorts)
  • other AUC: MetaRIF = 0.88; koot_sig = 0.48; Wang_sig = 0.54; OSR_score = 0.72 (comparison of MetaRIF vs previously published RIF classifiers)
  • count 134 controls and 104 RIF samples (total samples across all datasets used in the study)
  • count 10,383 common genes from 108 controls and 85 RIF samples (integrated cohort used for DEG discovery (GSE111974, GSE58144, GSE106602))
  • pvalue p < 0.05 (significance threshold for DEG identification via MetaDE)
  • count 33 endometrial biopsy samples (12 RIF, 21 control) (independent in-house prospective validation cohort)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study harmonizes four publicly available GEO endometrial transcriptomic datasets and one prospectively collected in-house cohort using a random-effects meta-analytic framework (MetaDE) to identify differentially expressed genes (DEGs) between RIF and normal endometrium. Unsupervised consensus clustering (ConsensusClusterPlus, PAM/Euclidean, 1000 iterations) partitioned RIF samples into two molecular subtypes, which were characterized via GSEA, GSVA pathway enrichment, and CIBERSORT-based immune deconvolution. A binary classifier (MetaRIF) was constructed by evaluating 64 machine learning algorithm combinations under leave-one-out cross-validation and selecting the model with the highest average F-score, then validated in independent cohorts using ROC/AUC; pairwise group comparisons used the Wilcoxon rank-sum test and Spearman correlation throughout, with a universal significance threshold of p < 0.05.

Replicationmixed Sample sizeTraining: 85 RIF and 108 controls from three GEO datasets (no formal power calculation stated); validation: GSE71331 (7 RIF, 5 controls) and prospective in-house cohort (12 RIF, 21 controls) with pre-specified inclusion/exclusion criteria GroupsRIF vs normal endometrial samples; RIF-I subtype vs RIF-M subtype Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Random-effects meta-analysis for differential expression (MetaDE) Identification of DEGs between RIF and normal samples in integrated training datasets (GSE111974, GSE58144, GSE106602); 1,776 robust DEGs reported 85 RIF vs 108 controls, across 10,383 common genes not stated
Wilcoxon rank-sum test Pairwise group comparisons throughout (subtype-associated features, immune scores, gene expression); described generically in the Statistical Analysis section not stated
Spearman rank correlation Correlation between immune infiltration scores or gene expression levels and disease subtypes not stated
Gene Set Enrichment Analysis (GSEA) Biological characterization of RIF-I and RIF-M subtypes; p < 0.01 specifically cited for IL-17 and TNF signaling pathways in RIF-I not stated
Gene Set Variation Analysis (GSVA) Functional module activity comparison between RIF-I and RIF-M; immune cell infiltration scoring per sample not stated
KEGG pathway enrichment (clusterProfiler) Biological function analysis of RIF-related DEGs; p < 0.05 threshold applied not stated
ROC curve / AUC (pROC) Diagnostic performance of MetaRIF classifier reported for training cohort (AUC 0.88), validation cohort GSE71331 (AUC 0.94), and in-house cohort (AUC 0.85); comparison against published models (koot_sig AUC 0.48, Wang_sig AUC 0.54, OSR_score AUC 0.72) na
F-score (harmonic mean of sensitivity and specificity) for classifier selection Selection of optimal model from 64 ML algorithm combinations across LOOCV training and two validation datasets na
Approaches that could also have been used
  • DEGs were identified with a nominal p < 0.05 threshold across ~10,383 simultaneously tested genes, with no stated FDR correction
    Could also: Applying a Benjamini-Hochberg FDR correction (e.g., q < 0.05 or q < 0.10) over all tested genes would also be a standard approach for transcriptomic discovery — With thousands of simultaneous comparisons, FDR control is widely recommended in transcriptomic studies to characterize the expected proportion of false discoveries among reported DEGs; reporting an FDR alongside the nominal p-value allows readers to calibrate confidence in the DEG list independently of its size
  • The optimal MetaRIF classifier was selected from 64 ML algorithm combinations on the basis of average F-score across training (LOOCV) and two validation datasets
    Could also: AUC, balanced accuracy, or Matthews correlation coefficient (MCC) could also serve as the primary model-selection criterion — F-score weights sensitivity and specificity equally and is threshold-dependent; AUC is threshold-independent and summarizes performance across all decision points, while MCC is particularly informative when class sizes differ — either could offer a complementary or more complete characterization of classifier quality
  • Unsupervised subtype discovery used PAM clustering with Euclidean distance within ConsensusClusterPlus
    Could also: Non-negative matrix factorization (NMF) or hierarchical clustering with correlation-based (Pearson or Spearman) distance could also be applied to identify transcriptomic subtypes — Correlation-based distances capture co-expression patterns independently of absolute expression magnitude, which is commonly preferred for RNA-seq/microarray data; NMF provides additive, parts-based decomposition that can align naturally with co-expressed biological modules; cross-method agreement on cluster membership would also strengthen confidence in the identified subtypes
  • Immune cell infiltration was estimated from bulk expression data using CIBERSORT with a custom single-cell-derived endometrial reference
    Could also: xCell, MCP-counter, or TIMER2 could also be used for immune deconvolution, either alone or alongside CIBERSORT — Different deconvolution algorithms make distinct assumptions about reference profiles and mixture models, and can yield divergent estimates for some cell types; reporting concordance across two or more methods is a common approach to assess robustness of infiltration estimates in the same samples
  • Pairwise group comparisons throughout used the Wilcoxon rank-sum test, treating samples from multiple GEO datasets as exchangeable observations
    Could also: Linear mixed-effects models with dataset of origin as a random effect could also be applied when comparing subtype features across samples pooled from multiple cohorts — Because samples originate from multiple platforms and preprocessing pipelines, modeling cohort as a random effect can more explicitly account for inter-dataset variability than the Wilcoxon test, which treats all observations as independently and identically distributed
  • No formal sample size or statistical power calculation is described for the prospectively collected in-house cohort (12 RIF, 21 controls)
    Could also: A prospective power calculation based on expected effect sizes from existing literature or the GEO training data could also be reported prior to collection — Pre-specified power analyses contextualize whether a validation cohort is adequately sized to detect effects of the anticipated magnitude, and are standard practice in prospective clinical sample collection studies; reporting achieved power post-hoc is also an alternative when pre-specification was not performed
Software: R 4.1.1 · MetaDE 2.2.3 · pROC 1.18.5 · HISAT 0.1.6 · ConsensusClusterPlus · clusterProfiler · GSVA · CIBERSORT · METAFlux · Affy · ggplot2

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
11
Impact: medium
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 56/100
partly built on non-reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE106602 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE111974 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE58144 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE71331 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40660214 (MetaRIF)

Paper: Yang et al. 2025, J Transl Med. "Molecular subtype of recurrent implantation failure reveals distinct endometrial etiology of female infertility." DOI 10.1186/s12967-025-06771-1 · PMCID PMC12257665. Repo: https://github.com/YJ-STU/MetaRIF (R, 3 commits). Repo ships: MetaRIF.R (a classifier wrapper) + MetaRIF_classifier_gene_plus.rda (pretrained ensemble models best_model_RIF11/best_model_RIF22, plus the signature-gene tables gene1/gene2 as ENTREZID). README is only a 1-line abstract.

What the shipped code actually does

MetaRIF(data) takes an expression matrix (rows = ENTREZID, cols = samples), double z-score scales each gene across samples, subsets to the two signature-gene sets, runs the two pretrained probability models, and returns a per-sample pred ∈ {Normal, RIF1, RIF2, Uncertain} and prob (Normal_prob, RIF_prob). => The shipped artifact is the diagnosis/subtyping classifier only.

In scope (pipeline-derived, reproducible with shipped code + public data)

  • C1 — Classifier validation AUC on GSE71331. Paper reports the MetaRIF classifier achieves AUC = 0.94 on the independent validation cohort GSE71331 (7 RIF, 5 control; GPL9072 Agilent). This is the single cleanest end-to-end reproducible claim: shipped .rda model applied to a public GEO dataset, AUC of RIF-probability vs true RIF/control labels. PRIMARY TARGET.
  • C2 (secondary) — Classifier behaviour on GSE111974 (the brief's listed accession; 24 RIF, 24 control; GPL17077). No held-out AUC is reported for this cohort specifically (it is part of the discovery/integration set), so this is reported only as a sanity/behaviour data point, not graded against a number.

Out of scope (the hard ~20% — not attempted, with reasons)

  • 1,776 robust DEGs (MetaDE across 5 datasets + in-house MARS-seq). Requires the in-house RNA-seq (CNGB HRA010876) integrated with 5 platforms via MetaDE/ random-effect batch correction. The integration code is NOT in the repo (only the final classifier is). Not 1:1 reproducible from shipped artifacts → skip.
  • ConsensusClusterPlus k=2 subtyping, CIBERSORT/GSVA immune profiling, KEGG/GSEA, METAFlux, cMAP drug prediction. Pipelines described in Methods but their driver code is not shipped; depend on the integrated 238-sample matrix. Out of scope.
  • IHC (T-bet/GATA3), HALO image analysis — wet-lab/manual, non-pipeline.
  • In-house validation AUC = 0.85 — needs restricted CNGB in-house data.

Data

  • GSE71331, GSE111974 — public GEO series matrices, fetched inside the «our HPC» compute job to «infra». Repo cloned on «infra» inside the job.
Figures / tables: Fig.7Fig.3
C1
Reported
AUC = 0.94 (classifier validation on GSE71331)
Reproduced
not reproduced — blocked
m.public.grade.error
C2
Reported
AUC = 0.85 (in-house validation cohort)
Reproduced
not attempted — data in CNGB HRA010876 (restricted)
m.public.grade.error
C3
Reported
two RIF subtypes RIF1/RIF2 + Normal; signatures 64 + 19 ENTREZID genes
Reproduced
reproduced: 24/24 controls->Normal, RIF->5 RIF1 + 12 RIF2 (+6 Normal,1 Uncertain) on GSE111974; full 64/64 & 19/19 signature coverage
within tolerance
C4
Reported
no held-out AUC reported for GSE111974 specifically
Reproduced
AUC=0.9965, accuracy=0.872, control specificity 24/24 (in-sample)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 68/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

180 k
tokens (I/O) · 12 M incl. cache
39 min
runtime · 0.02 CPU-h
1.7 GB
peak RAM
5 (2 failed)
HPC jobs
hummel
machine