Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Disrupted PGR-B and ESR1 signaling underlies defective decidualization linked to severe preeclampsia.

Elife · 2021
L1 78/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 51% of all assessed papers rank 533 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> essentially 1:1. The GEO record (GSE172381) fully specifies the pipeline (edgeR 3.24.3, R 3.5.1, TMM + glmTreat FC-threshold 1.2, FDR<0.05) and ships both the input (raw counts, all 40 samples) and the answer keys (Figure1=593 DEGs, Figure3=120-gene signature, with per-gene stats). Re-running the exact pipeline (built with the EXACT reported tool versions) on the train set (12 control vs 17 sPE) reproduced 580 DEGs sharing 563/593 (94.9%) of the paper's list, and recovered 109/120 (90.8%) of the signature genes. Decisive check: the authors' shipped per-gene logFC reproduces from the shipped counts at Pearson r=1.0000 -> no fabrication signal. edgeR's conservative glmQLFit->glmTreat path matches the paper (the liberal glmFit/LRT path gives 933, too many), pinning the exact method. NOT attempted (optional last 20%): the test-set validation (4 vs 7), upstream FASTQ->counts re-alignment (counts are shipped), and all wet-lab/PGR-B/ESR1 functional results (out of scope). The harvested 'code' link (bedtools2) is a generic third-party tool, not the authors' code and not on the DE path (brief P16): reproduced via the described pipeline on the paper's own data.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 78
    assessed: 2026-06-15 ⛓ 392ca2c74bad
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper tests whether a preconception endometrial transcriptomic signature of defective decidualization exists in women who developed severe preeclampsia (sPE), and whether dysregulated hormonal (ESR1/PGR-B) signaling underlies this defect linking decidualization to sPE.

Core claims
  • A 120-gene endometrial transcriptomic fingerprint encodes defective decidualization associated with severe preeclampsia. finding
  • This 120-gene signature effectively segregates samples into sPE and control groups in an independent validation cohort. resource
  • ESR1 and PGR are highly interconnected hub genes within the defective decidualization interaction network. mechanism
  • ESR1 and PGR-B gene expression and protein abundance are disrupted in sPE decidual endometrium. finding
  • The defective decidualization transcriptomic signature persists for years after the affected pregnancy and may serve as a preconception/early prenatal screening tool for sPE risk. finding
  • 593 genes were initially identified as differentially expressed in sPE versus control endometrium. finding
  • Global RNA-seq on late secretory-phase endometrial biopsies, split into training and test sets, defines and validates the fingerprint. method
  • No significant transcriptomic differences exist between preterm and full-term controls, ruling out gestational-age bias. finding
Experimental setups
Assay System Perturbation Readout Platform
Global bulk RNA-seq (transcriptome sequencing) Human endometrial biopsies, late secretory phase, non-pregnant women (n=40) none (prior sPE history vs control) Genome-wide gene expression / differentially expressed genes Illumina NextSeq 500/550, TruSeq Stranded mRNA kit, 150 bp paired-end
RT-qPCR Human endometrial tissue (prior sPE n=13 vs controls n=9) none (sPE vs control) Relative gene expression of IHH, MSX2, ESR1, PGR isoforms (PGR-A, PGR-B) Kapa SYBR fast qPCR kit, Roche Lightcycler 480; SuperScript VILO cDNA kit
Immunofluorescence of tissue sections Paraffin-embedded human endometrial tissue none (sPE vs control) Protein abundance/localization of progesterone receptor and estrogen receptor alpha (ESR1) EVOS M5000 microscope; anti-PR [YR85] Abcam, anti-ERα Santa Cruz, AlexaFluor 488 secondaries
Protein–protein interaction network analysis Proteins encoded by 120 fingerprint genes (in silico) none Hub genes via MCC/MNC topological analysis STRING, Cytoscape, CytoHubba
Gene Ontology enrichment analysis 120 fingerprint genes (in silico) none Enriched biological processes (FDR<0.05) goana function in edgeR
Key results
  • Defective decidualization fingerprint comprises 120 genes associated with sPE 120 genes
  • 593 genes differentially expressed in sPE versus control endometrium 593 genes
  • 120-gene fingerprint segregated test-set samples into sPE and control groups (PCA and hierarchical clustering)
  • ESR1 and PGR identified as highly interconnected hub genes in the network
  • ESR1 and PGR-B gene expression and protein abundance disrupted in sPE
  • No significant transcriptomic differences between preterm and term controls FDR ≥ 0.05
  • 40 samples produced 56,638 raw genes; 18,301 genes retained after normalization 18,301 genes
Key statistics
  • count 40 non-pregnant women (24 sPE, 16 controls) (Total cohort enrolled for endometrial RNA-seq)
  • count training n=29 (12 control, 17 sPE); test n=11 (4 control, 7 sPE) (70:30 stratified split of samples)
  • count 120 genes (Defective decidualization fingerprint size)
  • count 593 genes (Initial DEGs in sPE vs control)
  • fold_change ≥1.4-fold (FDR<0.05) (Criteria selecting fingerprint genes, sPE vs control training set)
  • other FDR<0.05, fold-change threshold 1.2 (edgeR glmTreat DEG cutoffs)
  • count RIN values 4.9 to 9.2 (RNA integrity of samples used for global RNA-seq)
  • other 8% of first-time pregnancies (Reported prevalence of preeclampsia)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is an RNA-sequencing biomarker-discovery study comparing endometrial transcriptomes of women with prior severe preeclampsia (sPE) versus controls. Samples were split by stratified random sampling into a training set (n=29) for differential expression and fingerprint definition and a test set (n=11) for validation; differential expression was performed with edgeR (TMM normalization, glmTreat) using an FDR<0.05 cutoff and a fold-change threshold, followed by GO enrichment (goana), a STRING/Cytoscape interaction network, RT-qPCR (ΔΔCT) validation, and immunofluorescence. Clinical data were compared with the Wilcoxon test and reported as mean ± SEM.

Replicationbiological Sample sizeTotal n=40 (sPE=24, controls=16); stratified random 70:30 split into training (n=29) and test (n=11); no formal power/sample-size calculation described Groupsprior sPE vs. controls (term and preterm) Pairingunpaired Randomization/blindingstated DispersionSEM Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg false discovery rate (FDR)
Statistical tests used
Test Applied to n Assumptions
edgeR generalized linear model with glmTreat (negative binomial, threshold-based test against a fold-change) Differential expression between sPE and control endometrial samples (RNA-seq, training set) training set n=29 (control n=12, sPE n=17) not stated
Gene Ontology enrichment via goana (edgeR) Biological-process enrichment of the 120-gene fingerprint 120 fingerprint genes na
Wilcoxon test Comparisons of clinical maternal/neonatal data between sPE and controls not stated
Comparative Ct (2−ΔΔCT) method for relative gene expression RT-qPCR of IHH, MSX2, ESR1, PGR isoforms sPE n=13, controls n=9 na
Principal component analysis and unsupervised hierarchical clustering (Canberra distance) Sample segregation on the fingerprint; controls by gestational age na
Approaches that could also have been used
  • Clinical data were summarized as mean ± SEM.
    Could also: The same data could also be summarized with the standard deviation or a 95% confidence interval (and, for the Wilcoxon comparisons, median with IQR). — SD or IQR conveys the spread of the observations themselves rather than the precision of the mean, and a CI conveys estimate uncertainty directly; these are often preferred for small samples and for nonparametric comparisons.
  • Differential expression used edgeR with TMM normalization and the glmTreat threshold test.
    Could also: DESeq2 (Wald or LRT) or limma-voom could also have been applied to the same count data. — Running an alternative or complementary pipeline can provide a concordance check on the gene list, and each tool offers somewhat different dispersion-shrinkage and normalization models that may suit different count distributions.
  • Clinical comparisons used the Wilcoxon test for each variable.
    Could also: Where multiple clinical variables were tested, a multiplicity adjustment across that family (e.g., Benjamini-Hochberg) could also have been reported. — Applying the same FDR control used for the omics data to the clinical comparisons would keep the family-wise/false-discovery handling consistent across all reported tests.
  • The fingerprint was defined and validated using an FDR and fold-change threshold and then confirmed in an independent test set.
    Could also: Cross-validation or resampling (e.g., k-fold or bootstrap) and reporting of classifier metrics with confidence intervals (ROC-AUC, sensitivity/specificity) could also accompany the single train/test split. — Resampling-based estimates and interval reporting can characterize the stability of the signature and the uncertainty of its classification performance, which is informative given the modest sample size.
  • RT-qPCR was analyzed with the comparative 2−ΔΔCT method and group differences shown relative to the control median.
    Could also: Formal statistical tests on the ΔCT values (e.g., Mann-Whitney U or t-test on ΔCT) with reported effect sizes could also be presented. — Performing the test on ΔCT rather than ratio-transformed 2−ΔΔCT values and reporting effect sizes makes the inferential basis for each gene comparison explicit.
Software: edgeR (Bioconductor) 3.24.3 · R 3.5.1 · STAR 2.4.2a · FastQC 0.11.2 · SAMtools 1.1 · HTSeq 0.6.1p1 · BEDtools 2.17.0 · STRING · Cytoscape / cytoHubba

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
36
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

RRID:AB_2534069 RRID in Article (http://semanticscience.org/resource/SIO_001029)
also used by 3 papers:
RRID:SCR_002105 RRID in Article (http://semanticscience.org/resource/SIO_001029)
also used by 1 paper:
RRID:SCR_014583 RRID in Article (http://semanticscience.org/resource/SIO_001029)
also used by 1 paper:
GSE172381 GEO in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2630356 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_627558 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_777452 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_004463 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_005223 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_005514 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_006646 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_012802 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_017677 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
SCR_003032 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34709177

Paper: Garrido-Gomez et al. 2021, eLife 10:e70753. "Disrupted PGR-B and ESR1 signaling underlies defective decidualization linked to severe preeclampsia." PMID 34709177 · PMCID PMC8553341.

Pipeline-derived results (IN SCOPE — attempted)

The study's bioinformatic core is a bulk RNA-seq differential-expression analysis of preconception endometrium (late secretory phase), comparing women with a prior severe-preeclampsia (sPE) pregnancy vs controls, on a 70:30 train/test split. Method fully specified in the GEO record's data_processing:

  • edgeR v3.24.3 (R 3.5.1), TMM normalization, low-count filter.
  • DE via glmTreat with fold-change threshold 1.2, FDR < 0.05.
  • TRAIN set: control n=12 vs sPE n=17 (29 samples) -> DEG list.
  • Figure 1 SourceData: 593 DEGs (FDR<0.05, FC>=1.2).
  • Figure 3 SourceData: 120-gene "defective decidualization signature" = DEGs with FC>=1.4 and an EntrezID.

Reproduced by re-running this exact edgeR pipeline on the authors' shipped raw count matrix (GSE172381_rawData.csv, 56638 genes x 40 samples, gene symbols, hg19/STAR). See claims.tsv C1-C3.

OUT OF SCOPE (not attempted)

  • The test-set validation (4 control vs 7 sPE) of the signature — a targeted confirmatory step, part of the optional last 20%.
  • Wet-lab / in-vitro decidualization, immunostaining, PGR-B / ESR1 protein and functional assays (Figs 2,4-6) — manual/experimental, not pipeline-derived.
  • Upstream read alignment/quantification (FASTQ->counts): the authors ship the count matrix; we reproduce from counts (the documented, deterministic stage). Re-aligning SRP315497 from FASTQ is not needed to check the reported DE values.

Note on the "code" link (P16)

The harvested code URL is github.com/arq5x/bedtools2 — a generic third-party genome-arithmetic tool, not the authors' analysis code (no authors' repo was deposited). Per the brief, applying the described pipeline (edgeR) to the paper's own data is an equally valid reproduction; bedtools is not on the DE path and was not used.

Data & code provenance (pointers)

  • Data: GEO GSE172381 (SRA SRP315497, BioProject PRJNA723182). Raw counts + Figure1/Figure3 SourceData xlsx downloaded inside the «our HPC» job to «infra».
  • Pipeline tool: edgeR 3.24.3 (Bioconductor), built with conda in the compute job.
  • «infra» work dir: «path»
Figures / tables: Fig.1Figure1_SourceData1Fig.3Figure3_SourceData1
C1_DEGs
Reported
593 DEGs (edgeR glmTreat, FDR<0.05, FC>=1.2), sPE vs control, train set
Reproduced
580 DEGs; 563/593 shared (94.9% recall)
within tolerance
C2_signature120
Reported
120-gene defective-decidualization signature (FC>=1.4 + EntrezID)
Reproduced
109/120 signature genes recovered (90.8%)
partial
C3_logFC_reproducibility
Reported
per-gene logFC (Figure1 SourceData)
Reproduced
Pearson r=1.0000, Spearman rho=1.0000 over 562 shared genes
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1

This reproduces essentially 1:1: re-running the exactly-specified pipeline (edgeR 3.24.3 / R 3.5.1, TMM + glmQLFit→glmTreat lfc=log2(1.2), FDR<0.05) on the shipped GSE172381 counts gives 580 DEGs sharing 563/593 (94.9%) and recovers 109/120 signature genes, with per-gene logFC matching at r=1.0000 — a decisive no-fabrication signal. The only deviation sits on the input/preprocessing side (our low-count filter dropping a handful of extreme-FC tail genes), is negligible in magnitude, and is on our-method/technical side, not the authors'. Core conclusion fully holds; rated yellow overall only because the DEG/signature counts are not a perfect match.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

167.6 k
tokens (I/O) · 9.6 M incl. cache
16 min
runtime · 0.01 CPU-h
0.2 GB
peak RAM
1
HPC jobs
hummel
machine