Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

STAT1 and IL-7 as potential diagnostic biomarkers for distinguishing high-grade from low-grade serous ovarian cancer: a mu

Front Immunol · 2026
L1 50/100 PQI 83
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Same input data as the authors
  • Reported values are derivable from the shared data
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the headline single-dataset signals 1:1 in DIRECTION, same-class in magnitude. Reproduced the paper's own 'Independent GSE27651 Analysis' with limma on GEO GSE27651 (22 HGSOC vs 13 LGSOC, exactly the paper's sample counts): STAT1 up in HGSOC (logFC ~+1.4..+2.1 vs reported +1.625), IL-7 down (logFC -2.80 vs reported -2.212). Also installed and ran the literally-named third-party repo xCell (dviraran/xCell, P16-valid) on the same matrix and reproduced the robust POSITIVE STAT1<->M1-macrophage correlation (rho +0.83 single-dataset vs +0.457 paper combined-matrix). None is bit-exact: expected, because we used the GEO series matrix (paper likely re-processed raw CEL with RMA) and, for the deconvolution, a single dataset vs the paper's ComBat-merged multi-cohort matrix. Fabrication concern: none - every reported value follows in sign and order of magnitude from the shipped data. NOT attempted (hard ~20%, out of scope): the 983-DEG count, training/validation AUCs (STAT1 0.908, IL-7 0.842, combined 0.938), LASSO/SVM-RFE feature selection, nomogram/DCA/bootstrap, CIBERSORT LM22 absolute values, external cohorts (GSE14001/73168/146965), and IHC - all depend on merging GSE27651+GSE126132 with ComBat batch correction whose unspecified covariate/reference choices make exact reproduction infeasible without guessing.

💻 Code ↗ 🗄 Data: GSE27651

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-14 ⛓ 1c58ca2c6346
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests whether differentially expressed immune-related genes can serve as reliable diagnostic biomarkers to distinguish high-grade serous ovarian carcinoma (HGSOC) from low-grade serous ovarian carcinoma (LGSOC), proposing STAT1 and IL-7 as candidates.

Core claims
  • STAT1 and IL-7 are differentially expressed immune-related genes that can distinguish HGSOC from LGSOC and may serve as ancillary diagnostic biomarkers. finding
  • STAT1 protein is significantly higher and IL-7 protein significantly lower in HGSOC tissues versus LGSOC, confirmed by IHC. finding
  • 71 differentially expressed immune-related genes (DIRGs) distinguish HGSOC from LGSOC, enriched in cytokine-mediated signaling, cytokine-cytokine receptor interaction, and JAK-STAT pathways. finding
  • A combined LASSO regression and SVM-RFE machine learning pipeline applied to 10 hub DIRGs selects diagnostic biomarkers, refined to STAT1 and IL-7 after external validation. method
  • HGSOC shows higher fractions of naïve B cells, M2 macrophages, and neutrophils and lower resting memory CD4+ T cells and eosinophils relative to LGSOC. finding
  • STAT1 expression is strongly positively correlated with M1 macrophage infiltration. mechanism
  • A nomogram diagnostic model based on STAT1 and IL-7 was developed and calibrated for HGSOC diagnosis. resource
Experimental setups
Assay System Perturbation Readout Platform
transcriptome microarray differential expression analysis (limma) HGSOC and LGSOC tumor tissue (training cohort GSE27651 + GSE126132) none differentially expressed immune-related genes (|log2FC|≥1, adj P<0.05) Affymetrix U133 Plus 2.0 (GPL570); Illumina HT-12 V4.0 (GPL10558)
machine-learning feature selection (LASSO + SVM-RFE) and ROC/nomogram diagnostic modeling training cohort expression data (n=69; 10 hub DIRGs) none selected diagnostic biomarkers and AUC R glmnet, pROC, rms
diagnostic validation via ROC independent merged validation cohort (GSE14001 + GSE73168 + GSE146965; 55 HGSOC, 13 LGSOC) none AUC for STAT1 and IL-7 GPL570; Affymetrix Clariom D (GPL23126)
immunohistochemistry paraffin-embedded ovarian tissue (26 HGSOC, 12 LGSOC, Second Affiliated Hospital of Fujian Medical University) none STAT1 and IL-7 protein staining scores (intensity × percentage) anti-STAT1 and anti-IL-7 antibodies (Affbiotech)
immune cell deconvolution (CIBERSORT, LM22) HGSOC and LGSOC expression matrix none relative abundance of 22 immune cell subtypes and correlation with STAT1/IL-7 CIBERSORT LM22 (547 genes); validated with xCell
functional enrichment (GO/KEGG) and PPI network / hub gene analysis 71 DIRGs none enriched terms and top 10 MCC hub genes clusterProfiler; STRING; Cytoscape 3.10.0 cytoHubba
Key results
  • STAT1 diagnostic AUC in the training group AUC=0.908
  • IL-7 diagnostic AUC in the training group AUC=0.842
  • STAT1 AUC in independent merged validation cohort AUC=0.703 (95% CI 0.517–0.889)
  • IL-7 AUC in independent merged validation cohort AUC=0.706 (95% CI 0.501–0.912)
  • STAT1 protein expression higher and IL-7 lower in HGSOC by IHC P<0.05
  • STAT1 expression positively correlated with M1 macrophages ρ=0.688, q=9.9×10^-8
  • IL-7 expression negatively correlated with neutrophils (not significant after FDR) ρ=−0.372, raw P=0.0048, q=0.100
  • HGSOC showed higher naïve B cells, M2 macrophages, neutrophils and lower resting memory CD4+ T cells and eosinophils all q<0.05
Key statistics
  • count 71 DIRGs (differentially expressed immune-related genes in HGSOC vs LGSOC)
  • other AUC=0.908 (STAT1 training-group diagnostic performance)
  • other AUC=0.842 (IL-7 training-group diagnostic performance)
  • other AUC=0.703 (95% CI 0.517–0.889) (STAT1 validation cohort (55 HGSOC, 13 LGSOC))
  • other AUC=0.706 (95% CI 0.501–0.912) (IL-7 validation cohort)
  • correlation ρ=0.688, q=9.9×10^-8 (STAT1 vs M1 macrophages)
  • correlation ρ=−0.372, raw P=0.0048, q=0.100 (IL-7 vs neutrophils)
  • pvalue P<0.05 (IHC STAT1/IL-7 difference (26 HGSOC, 12 LGSOC))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This multi-cohort microarray study analyzed GEO expression data to identify differentially expressed immune-related genes (DIRGs) between HGSOC and LGSOC, using limma for differential expression and Benjamini-Hochberg FDR correction throughout. Two complementary machine learning methods (LASSO regression and SVM-RFE) were applied to 10 PPI/MCC-ranked hub DIRGs to select STAT1 and IL-7 as final biomarkers, whose diagnostic performance was quantified by ROC/AUC analysis in training and independent validation cohorts. IHC on clinical specimens provided protein-level validation, and CIBERSORT with xCell quantified immune cell infiltration, with group comparisons by Mann-Whitney U and gene-cell correlations by Spearman rank correlation, both FDR-adjusted.

Replicationbiological Sample sizeTraining: n=69 (13 LGSOC, 56 HGSOC; GSE27651+GSE126132); validation: up to n=68 (13 LGSOC, 55 HGSOC; GSE14001+GSE73168+GSE146965); IHC: n=38 (12 LGSOC, 26 HGSOC) GroupsHGSOC vs LGSOC Pairingunpaired Randomization/blindingstated DispersionCI Exact p-valuesyes Effect sizesyes Confidence intervalsyes Multiplicity correctionBenjamini-Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
limma moderated t-test (empirical Bayes) Differential expression analysis: HGSOC vs LGSOC in merged training cohort; threshold |log2FC|≥1, BH-adjusted P<0.05 69 (13 LGSOC, 56 HGSOC; GSE27651 + GSE126132) not stated
LASSO logistic regression (binomial, lambda.min, single run 10-fold CV) Feature selection from 10 hub DIRGs in training cohort 69 not stated
SVM-RFE (linear kernel, C=1, single run 10-fold CV) Feature selection from 10 hub DIRGs in training cohort 69 not stated
ROC curve analysis with AUC (DeLong method and/or bootstrap for 95% CI, pROC package) Diagnostic evaluation of STAT1 and IL-7 individually and combined nomogram in training cohort and merged validation cohort; Youden index for optimal threshold Training: 69; merged validation: 68 (13 LGSOC, 55 HGSOC from GSE14001+GSE73168+GSE146965) not stated
Two-tailed Mann-Whitney U test with Benjamini-Hochberg FDR correction Immune cell fraction comparisons between HGSOC and LGSOC (CIBERSORT LM22 output, 22 cell types) null (training cohort after CIBERSORT p<0.05 filtering; exact post-filter n not stated in text) not stated
Spearman rank correlation with Benjamini-Hochberg FDR correction Correlation between STAT1/IL-7 expression and 22 CIBERSORT immune cell fractions null (not stated in text) not stated
PERMANOVA (adonis2, 999 permutations) Quantification of proportion of variance explained by batch/cohort before and after ComBat correction null not stated
Logistic regression with bootstrap calibration (1,000 replicates) and decision curve analysis (50 bootstraps) Nomogram construction, calibration curve, and DCA for STAT1+IL-7 combined diagnostic model 69 (training cohort) not stated
Learning curve analysis (5-fold CV, 10 repetitions per step, training fraction 30%–100% in 10% increments) Overfitting assessment of four-gene logistic regression model 69 not stated
Pairwise Spearman correlation Co-expression network among 10 hub DIRGs in training cohort 69 not stated
Approaches that could also have been used
  • LASSO and SVM-RFE feature selection each used a single run of 10-fold cross-validation in a training cohort of n=69
    Could also: Repeated cross-validation (e.g., 50–100 repetitions of 5-fold CV) or a bootstrap feature-selection frequency analysis (e.g., 1,000 subsampled runs, retaining features selected in >50% of runs) could also have been applied — With small n, a single CV split can yield variable feature rankings depending on the random partition; repeated CV or bootstrap selection frequencies quantify how consistently each feature is chosen across different data subsets, providing a stability-informed view of which genes are reliably predictive
  • For LASSO, lambda.min (the λ minimizing mean cross-validated deviance) was selected to retain a comprehensive feature set
    Could also: lambda.1se (the largest λ within one standard error of the minimum) could also have been used — lambda.1se applies greater regularization and yields a sparser model; it is often preferred in small-n settings as it reduces the risk of retaining noise features while achieving comparable cross-validated performance, and the authors explicitly note lambda.min was chosen to be inclusive for downstream validation
  • Immune cell fraction comparisons used Mann-Whitney U tests, but group medians and interquartile ranges were not reported in the text
    Could also: Reporting group medians with IQR (or box-and-whisker plots with data overlay) alongside the Mann-Whitney U results could also have been done — Median ± IQR is the natural descriptive complement to a non-parametric rank test; it allows readers to directly assess effect magnitude and the shape of the immune-fraction distributions for each of the 22 cell types, which is not conveyed by p-values alone
  • IHC protein expression was quantified by semi-quantitative multiplicative scoring (intensity × proportion) assessed by two blinded pathologists
    Could also: Automated digital image analysis (e.g., QuPath or HALO) with pixel-level DAB optical-density quantification could also have been applied — Digital pathology produces fully continuous, operator-independent scores and enables spatial analyses (e.g., distance from tumor margin, co-localization) not available from semi-quantitative scoring; it also provides a directly reproducible measure for future multi-site validation studies
  • The validation cohort included GSE146965 (n=40 HGSOC, 0 LGSOC), which contributed only HGSOC samples to the merged-cohort ROC; per-dataset independent ROC was limited to datasets with both subtypes
    Could also: A leave-one-dataset-out (LODO) cross-validation restricted to the three datasets containing both subtypes (GSE27651, GSE14001, GSE73168) could also have been performed as the primary generalizability estimate — LODO CV treats each complete (both-subtype) dataset as an independent external test set in turn, providing a directly interpretable generalization estimate without mixing HGSOC-only cohorts into the validation ROC; the paper already notes this was not fully feasible due to cohort composition, but a LODO estimate on the three balanced datasets would complement the merged-cohort approach
  • Immune cell deconvolution relied on CIBERSORT (LM22) as the primary method, with xCell used as a secondary concordance check for overlapping cell types
    Could also: Additional orthogonal algorithms such as TIMER2, EPIC, or quanTIseq could also have been included in a multi-method consensus analysis — Different deconvolution algorithms use distinct reference signature matrices and mixture-model assumptions; cross-method concordance across three or more algorithms strengthens confidence in infiltration estimates, particularly for rare cell types (e.g., eosinophils) where individual algorithms may have low signal-to-noise ratios in microarray data
Software: R/limma null (R v4.1.3 stated; package version in Supplementary Table S5) · R/sva (ComBat batch correction) · R/glmnet (LASSO) · R/e1071 or kernlab (SVM-RFE) · R/pROC (ROC, AUC, DeLong CI) · R/rms and rmda (nomogram, DCA) · R/clusterProfiler (GO/KEGG enrichment) · R/vegan (PERMANOVA/adonis2) · CIBERSORT (LM22, 1000 permutations) · xCell · Cytoscape with cytoHubba (MCC) 3.10.0 · STRING (PPI network)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE126132 GEO in Abstract (http://purl.org/dc/terms/abstract)
no other assessed paper uses this yet
GSE27651 GEO in Abstract (http://purl.org/dc/terms/abstract)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Reproduction scope — PMID 42058211

Paper: Wu Z et al., STAT1 and IL-7 as potential diagnostic biomarkers for distinguishing high-grade from low-grade serous ovarian cancer, Front Immunol 2026. DOI 10.3389/fimmu.2026.1779912 · PMCID PMC13120972.

Named code artifact: https://github.com/dviraran/xCell (third-party immune deconvolution tool — per P16 this is a fully valid reproduction target: run the tool on the paper's own data).

Primary data of this RU: GEO GSE27651 (GPL570, 49 samples; the paper uses its 13 low-grade + 22 high-grade serous samples; title prefixes LGOSC / HGOSC).

In scope (attempted) — clear, single-dataset, 1:1 comparable

# Reported result Where Pipeline
C1 STAT1 logFC +1.625 (HGSOC vs LGSOC, GSE27651 alone) "Independent GSE27651 Analysis" limma DE on GSE27651
C2 IL-7 (IL7) logFC −2.212 (HGSOC vs LGSOC, GSE27651 alone) "Independent GSE27651 Analysis" limma DE on GSE27651
C3 xCell runs on the paper's matrix; STAT1↔M1-macrophage correlation positive (paper combined-matrix xCell: ρ=0.457) Fig 8C / Suppl. xCell (dviraran/xCell), Spearman

C1+C2 are the headline low-hanging targets: a single dataset, explicit sample grouping, an explicitly named per-dataset reported value → directly checkable. C3 exercises the literally-named repository (xCell) on the paper's data; the paper's exact xCell number is on the combined batch-corrected multi-dataset matrix, so on GSE27651 alone we check direction/sign, not the exact ρ.

Out of scope (NOT attempted) — the hard ~20%, why skipped

  • 983 DEGs / 71 DIRGs, training-set AUCs (STAT1 0.908, IL-7 0.842), combined model 0.938: require merging GSE27651+GSE126132 and ComBat batch correction. ComBat output is highly sensitive to covariate/reference choices not fully specified → not 1:1 reproducible without guessing; deferred.
  • LASSO / SVM-RFE feature selection, nomogram, DCA, bootstrap stability, learning curves: downstream of the combined matrix; same dependency.
  • External validation cohorts (GSE14001/73168/146965), CIBERSORT LM22 absolute values, IHC (wet-lab, out of scope by definition), PPI/cytoHubba.
  • CIBERSORT: requires the LM22 signature + (historically) registration; xCell is the openly-installable named repo, so we use it for the deconvolution check.

Compute

All on «our HPC» («infra») via SLURM; env built with conda inside the compute job (GEOquery + limma + xCell from GitHub). Data stays on «infra»; only small derived values land in the dataset folder.

Figures / tables: Fig 8C
C1
Reported
STAT1 log2FC HGSOC vs LGSOC (GSE27651) = +1.625
Reproduced
+1.386 (mean of 6 probes); main expressed probe 200887_s_at = +2.13; all probes positive (sign exact)
partial
C2
Reported
IL-7 (IL7) log2FC HGSOC vs LGSOC (GSE27651) = -2.212
Reproduced
-2.795 (single probe 206693_at, adj.P=8.2e-07); sign exact, magnitude within ~26%
partial
C3
Reported
xCell STAT1 vs M1-macrophage Spearman rho = +0.457 (combined matrix)
Reproduced
+0.828 (xCell 1.1.0 on GSE27651 alone); positive STAT1-M1 association reproduced (sign exact)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

All three in-scope single-dataset claims reproduced with exact direction and same-class magnitude (STAT1 +1.386/+2.13 vs +1.625; IL-7 -2.795 vs -2.212; STAT1↔M1 rho +0.828 vs +0.457). The deviations are on our side — series-matrix vs raw-CEL/RMA preprocessing, an unspecified probe-collapse rule, and a single-dataset vs the paper's ComBat-merged matrix for C3 — not authors' defects, and every reported value is derivable in sign and order of magnitude from the shared GSE27651 data (no fabrication concern). Confirmation is partial because the paper's headline diagnostic results (AUCs, 983 DEGs, LASSO, nomogram) require the combined matrix and were out of scope. Overall a solid, explainable reproduction → yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

127.3 k
tokens (I/O) · 9.4 M incl. cache
16 min
runtime · 0.03 CPU-h
1.7 GB
peak RAM
1
HPC jobs
hummel
machine