STAT1 and IL-7 as potential diagnostic biomarkers for distinguishing high-grade from low-grade serous ovarian cancer: a mu
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values are derivable from the shared data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the headline single-dataset signals 1:1 in DIRECTION, same-class in magnitude. Reproduced the paper's own 'Independent GSE27651 Analysis' with limma on GEO GSE27651 (22 HGSOC vs 13 LGSOC, exactly the paper's sample counts): STAT1 up in HGSOC (logFC ~+1.4..+2.1 vs reported +1.625), IL-7 down (logFC -2.80 vs reported -2.212). Also installed and ran the literally-named third-party repo xCell (dviraran/xCell, P16-valid) on the same matrix and reproduced the robust POSITIVE STAT1<->M1-macrophage correlation (rho +0.83 single-dataset vs +0.457 paper combined-matrix). None is bit-exact: expected, because we used the GEO series matrix (paper likely re-processed raw CEL with RMA) and, for the deconvolution, a single dataset vs the paper's ComBat-merged multi-cohort matrix. Fabrication concern: none - every reported value follows in sign and order of magnitude from the shipped data. NOT attempted (hard ~20%, out of scope): the 983-DEG count, training/validation AUCs (STAT1 0.908, IL-7 0.842, combined 0.938), LASSO/SVM-RFE feature selection, nomogram/DCA/bootstrap, CIBERSORT LM22 absolute values, external cohorts (GSE14001/73168/146965), and IHC - all depend on merging GSE27651+GSE126132 with ComBat batch correction whose unspecified covariate/reference choices make exact reproduction infeasible without guessing.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-14 ⛓ 1c58ca2c6346
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study tests whether differentially expressed immune-related genes can serve as reliable diagnostic biomarkers to distinguish high-grade serous ovarian carcinoma (HGSOC) from low-grade serous ovarian carcinoma (LGSOC), proposing STAT1 and IL-7 as candidates.
- ★ STAT1 and IL-7 are differentially expressed immune-related genes that can distinguish HGSOC from LGSOC and may serve as ancillary diagnostic biomarkers. finding
- ★ STAT1 protein is significantly higher and IL-7 protein significantly lower in HGSOC tissues versus LGSOC, confirmed by IHC. finding
- ★ 71 differentially expressed immune-related genes (DIRGs) distinguish HGSOC from LGSOC, enriched in cytokine-mediated signaling, cytokine-cytokine receptor interaction, and JAK-STAT pathways. finding
- ★ A combined LASSO regression and SVM-RFE machine learning pipeline applied to 10 hub DIRGs selects diagnostic biomarkers, refined to STAT1 and IL-7 after external validation. method
- ★ HGSOC shows higher fractions of naïve B cells, M2 macrophages, and neutrophils and lower resting memory CD4+ T cells and eosinophils relative to LGSOC. finding
- STAT1 expression is strongly positively correlated with M1 macrophage infiltration. mechanism
- A nomogram diagnostic model based on STAT1 and IL-7 was developed and calibrated for HGSOC diagnosis. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| transcriptome microarray differential expression analysis (limma) | HGSOC and LGSOC tumor tissue (training cohort GSE27651 + GSE126132) | none | differentially expressed immune-related genes (|log2FC|≥1, adj P<0.05) | Affymetrix U133 Plus 2.0 (GPL570); Illumina HT-12 V4.0 (GPL10558) |
| machine-learning feature selection (LASSO + SVM-RFE) and ROC/nomogram diagnostic modeling | training cohort expression data (n=69; 10 hub DIRGs) | none | selected diagnostic biomarkers and AUC | R glmnet, pROC, rms |
| diagnostic validation via ROC | independent merged validation cohort (GSE14001 + GSE73168 + GSE146965; 55 HGSOC, 13 LGSOC) | none | AUC for STAT1 and IL-7 | GPL570; Affymetrix Clariom D (GPL23126) |
| immunohistochemistry | paraffin-embedded ovarian tissue (26 HGSOC, 12 LGSOC, Second Affiliated Hospital of Fujian Medical University) | none | STAT1 and IL-7 protein staining scores (intensity × percentage) | anti-STAT1 and anti-IL-7 antibodies (Affbiotech) |
| immune cell deconvolution (CIBERSORT, LM22) | HGSOC and LGSOC expression matrix | none | relative abundance of 22 immune cell subtypes and correlation with STAT1/IL-7 | CIBERSORT LM22 (547 genes); validated with xCell |
| functional enrichment (GO/KEGG) and PPI network / hub gene analysis | 71 DIRGs | none | enriched terms and top 10 MCC hub genes | clusterProfiler; STRING; Cytoscape 3.10.0 cytoHubba |
- – STAT1 diagnostic AUC in the training group AUC=0.908
- – IL-7 diagnostic AUC in the training group AUC=0.842
- – STAT1 AUC in independent merged validation cohort AUC=0.703 (95% CI 0.517–0.889)
- – IL-7 AUC in independent merged validation cohort AUC=0.706 (95% CI 0.501–0.912)
- – STAT1 protein expression higher and IL-7 lower in HGSOC by IHC P<0.05
- ▲ STAT1 expression positively correlated with M1 macrophages ρ=0.688, q=9.9×10^-8
- ▼ IL-7 expression negatively correlated with neutrophils (not significant after FDR) ρ=−0.372, raw P=0.0048, q=0.100
- – HGSOC showed higher naïve B cells, M2 macrophages, neutrophils and lower resting memory CD4+ T cells and eosinophils all q<0.05
- count 71 DIRGs (differentially expressed immune-related genes in HGSOC vs LGSOC)
- other AUC=0.908 (STAT1 training-group diagnostic performance)
- other AUC=0.842 (IL-7 training-group diagnostic performance)
- other AUC=0.703 (95% CI 0.517–0.889) (STAT1 validation cohort (55 HGSOC, 13 LGSOC))
- other AUC=0.706 (95% CI 0.501–0.912) (IL-7 validation cohort)
- correlation ρ=0.688, q=9.9×10^-8 (STAT1 vs M1 macrophages)
- correlation ρ=−0.372, raw P=0.0048, q=0.100 (IL-7 vs neutrophils)
- pvalue P<0.05 (IHC STAT1/IL-7 difference (26 HGSOC, 12 LGSOC))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This multi-cohort microarray study analyzed GEO expression data to identify differentially expressed immune-related genes (DIRGs) between HGSOC and LGSOC, using limma for differential expression and Benjamini-Hochberg FDR correction throughout. Two complementary machine learning methods (LASSO regression and SVM-RFE) were applied to 10 PPI/MCC-ranked hub DIRGs to select STAT1 and IL-7 as final biomarkers, whose diagnostic performance was quantified by ROC/AUC analysis in training and independent validation cohorts. IHC on clinical specimens provided protein-level validation, and CIBERSORT with xCell quantified immune cell infiltration, with group comparisons by Mann-Whitney U and gene-cell correlations by Spearman rank correlation, both FDR-adjusted.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma moderated t-test (empirical Bayes) | Differential expression analysis: HGSOC vs LGSOC in merged training cohort; threshold |log2FC|≥1, BH-adjusted P<0.05 | 69 (13 LGSOC, 56 HGSOC; GSE27651 + GSE126132) | not stated |
| LASSO logistic regression (binomial, lambda.min, single run 10-fold CV) | Feature selection from 10 hub DIRGs in training cohort | 69 | not stated |
| SVM-RFE (linear kernel, C=1, single run 10-fold CV) | Feature selection from 10 hub DIRGs in training cohort | 69 | not stated |
| ROC curve analysis with AUC (DeLong method and/or bootstrap for 95% CI, pROC package) | Diagnostic evaluation of STAT1 and IL-7 individually and combined nomogram in training cohort and merged validation cohort; Youden index for optimal threshold | Training: 69; merged validation: 68 (13 LGSOC, 55 HGSOC from GSE14001+GSE73168+GSE146965) | not stated |
| Two-tailed Mann-Whitney U test with Benjamini-Hochberg FDR correction | Immune cell fraction comparisons between HGSOC and LGSOC (CIBERSORT LM22 output, 22 cell types) | null (training cohort after CIBERSORT p<0.05 filtering; exact post-filter n not stated in text) | not stated |
| Spearman rank correlation with Benjamini-Hochberg FDR correction | Correlation between STAT1/IL-7 expression and 22 CIBERSORT immune cell fractions | null (not stated in text) | not stated |
| PERMANOVA (adonis2, 999 permutations) | Quantification of proportion of variance explained by batch/cohort before and after ComBat correction | null | not stated |
| Logistic regression with bootstrap calibration (1,000 replicates) and decision curve analysis (50 bootstraps) | Nomogram construction, calibration curve, and DCA for STAT1+IL-7 combined diagnostic model | 69 (training cohort) | not stated |
| Learning curve analysis (5-fold CV, 10 repetitions per step, training fraction 30%–100% in 10% increments) | Overfitting assessment of four-gene logistic regression model | 69 | not stated |
| Pairwise Spearman correlation | Co-expression network among 10 hub DIRGs in training cohort | 69 | not stated |
-
LASSO and SVM-RFE feature selection each used a single run of 10-fold cross-validation in a training cohort of n=69↳ Could also: Repeated cross-validation (e.g., 50–100 repetitions of 5-fold CV) or a bootstrap feature-selection frequency analysis (e.g., 1,000 subsampled runs, retaining features selected in >50% of runs) could also have been applied — With small n, a single CV split can yield variable feature rankings depending on the random partition; repeated CV or bootstrap selection frequencies quantify how consistently each feature is chosen across different data subsets, providing a stability-informed view of which genes are reliably predictive
-
For LASSO, lambda.min (the λ minimizing mean cross-validated deviance) was selected to retain a comprehensive feature set↳ Could also: lambda.1se (the largest λ within one standard error of the minimum) could also have been used — lambda.1se applies greater regularization and yields a sparser model; it is often preferred in small-n settings as it reduces the risk of retaining noise features while achieving comparable cross-validated performance, and the authors explicitly note lambda.min was chosen to be inclusive for downstream validation
-
Immune cell fraction comparisons used Mann-Whitney U tests, but group medians and interquartile ranges were not reported in the text↳ Could also: Reporting group medians with IQR (or box-and-whisker plots with data overlay) alongside the Mann-Whitney U results could also have been done — Median ± IQR is the natural descriptive complement to a non-parametric rank test; it allows readers to directly assess effect magnitude and the shape of the immune-fraction distributions for each of the 22 cell types, which is not conveyed by p-values alone
-
IHC protein expression was quantified by semi-quantitative multiplicative scoring (intensity × proportion) assessed by two blinded pathologists↳ Could also: Automated digital image analysis (e.g., QuPath or HALO) with pixel-level DAB optical-density quantification could also have been applied — Digital pathology produces fully continuous, operator-independent scores and enables spatial analyses (e.g., distance from tumor margin, co-localization) not available from semi-quantitative scoring; it also provides a directly reproducible measure for future multi-site validation studies
-
The validation cohort included GSE146965 (n=40 HGSOC, 0 LGSOC), which contributed only HGSOC samples to the merged-cohort ROC; per-dataset independent ROC was limited to datasets with both subtypes↳ Could also: A leave-one-dataset-out (LODO) cross-validation restricted to the three datasets containing both subtypes (GSE27651, GSE14001, GSE73168) could also have been performed as the primary generalizability estimate — LODO CV treats each complete (both-subtype) dataset as an independent external test set in turn, providing a directly interpretable generalization estimate without mixing HGSOC-only cohorts into the validation ROC; the paper already notes this was not fully feasible due to cohort composition, but a LODO estimate on the three balanced datasets would complement the merged-cohort approach
-
Immune cell deconvolution relied on CIBERSORT (LM22) as the primary method, with xCell used as a secondary concordance check for overlapping cell types↳ Could also: Additional orthogonal algorithms such as TIMER2, EPIC, or quanTIseq could also have been included in a multi-method consensus analysis — Different deconvolution algorithms use distinct reference signature matrices and mixture-model assumptions; cross-method concordance across three or more algorithms strengthens confidence in infiltration estimates, particularly for rare cell types (e.g., eosinophils) where individual algorithms may have low signal-to-noise ratios in microarray data
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Reproduction scope — PMID 42058211
Paper: Wu Z et al., STAT1 and IL-7 as potential diagnostic biomarkers for distinguishing high-grade from low-grade serous ovarian cancer, Front Immunol 2026. DOI 10.3389/fimmu.2026.1779912 · PMCID PMC13120972.
Named code artifact: https://github.com/dviraran/xCell (third-party immune deconvolution tool — per P16 this is a fully valid reproduction target: run the tool on the paper's own data).
Primary data of this RU: GEO GSE27651 (GPL570, 49 samples; the paper uses its 13 low-grade + 22 high-grade serous samples; title prefixes LGOSC / HGOSC).
In scope (attempted) — clear, single-dataset, 1:1 comparable
| # | Reported result | Where | Pipeline |
|---|---|---|---|
| C1 | STAT1 logFC +1.625 (HGSOC vs LGSOC, GSE27651 alone) | "Independent GSE27651 Analysis" | limma DE on GSE27651 |
| C2 | IL-7 (IL7) logFC −2.212 (HGSOC vs LGSOC, GSE27651 alone) | "Independent GSE27651 Analysis" | limma DE on GSE27651 |
| C3 | xCell runs on the paper's matrix; STAT1↔M1-macrophage correlation positive (paper combined-matrix xCell: ρ=0.457) | Fig 8C / Suppl. | xCell (dviraran/xCell), Spearman |
C1+C2 are the headline low-hanging targets: a single dataset, explicit sample grouping, an explicitly named per-dataset reported value → directly checkable. C3 exercises the literally-named repository (xCell) on the paper's data; the paper's exact xCell number is on the combined batch-corrected multi-dataset matrix, so on GSE27651 alone we check direction/sign, not the exact ρ.
Out of scope (NOT attempted) — the hard ~20%, why skipped
- 983 DEGs / 71 DIRGs, training-set AUCs (STAT1 0.908, IL-7 0.842), combined model 0.938: require merging GSE27651+GSE126132 and ComBat batch correction. ComBat output is highly sensitive to covariate/reference choices not fully specified → not 1:1 reproducible without guessing; deferred.
- LASSO / SVM-RFE feature selection, nomogram, DCA, bootstrap stability, learning curves: downstream of the combined matrix; same dependency.
- External validation cohorts (GSE14001/73168/146965), CIBERSORT LM22 absolute values, IHC (wet-lab, out of scope by definition), PPI/cytoHubba.
- CIBERSORT: requires the LM22 signature + (historically) registration; xCell is the openly-installable named repo, so we use it for the deconvolution check.
Compute
All on «our HPC» («infra») via SLURM; env built with conda inside the compute job (GEOquery + limma + xCell from GitHub). Data stays on «infra»; only small derived values land in the dataset folder.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
All three in-scope single-dataset claims reproduced with exact direction and same-class magnitude (STAT1 +1.386/+2.13 vs +1.625; IL-7 -2.795 vs -2.212; STAT1↔M1 rho +0.828 vs +0.457). The deviations are on our side — series-matrix vs raw-CEL/RMA preprocessing, an unspecified probe-collapse rule, and a single-dataset vs the paper's ComBat-merged matrix for C3 — not authors' defects, and every reported value is derivable in sign and order of magnitude from the shared GSE27651 data (no fabrication concern). Confirmation is partial because the paper's headline diagnostic results (AUCs, 983 DEGs, LASSO, nomogram) require the combined matrix and were out of scope. Overall a solid, explainable reproduction → yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.