Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Stemness genes and miR-1247-3p expression associate with clinicopathological parameters and prognosis in lung adenocarcinoma.

PLoS One · 2023
L1 90/100 PQI 97
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
90/100
Reproducibility score
0.9 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 79% of all assessed papers rank 211 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the PRIMARY result 1:1. The repo (Sashoss/LUAD_PLOS_data @ 8b6d7e1) ships data + derived outputs but NO authors' analysis code; per brief P16 we re-ran the paper-described pipeline (DESeq2 for TCGA, Limma for GEO; FC=tumor/normal) on the shipped inputs on «our HPC» («infra», conda DESeq2 1.42.0 / limma 3.58.1 / R 4.3). TCGA primary cohort (gene + miRNA DESeq2) reproduces EXACTLY: Pearson cor of log2FC = 1.000 over the whole transcriptome (34924 genes) and miRNA set (1359); all headline stemness genes (ORC1L 2.741, KIF20A 3.144, DLGAP5 3.671, RAB3B 4.242, CENPI 2.740) match Table 1 to 3 decimals, and the headline miR-1247-3p downregulation is exact (-2.325). The GEO validation cohorts (Limma) reproduce direction and ranking but NOT magnitude 1:1 (GSE31210 cor=0.9975 but 1.45x smaller; GSE40419 cor0.70-0.75, inconsistent) because the authors' microarray preprocessing/probe-collapsing is undocumented (no code shipped) - this is the hard ~20% and was not chased further. NOT attempted (out of scope): univariate/multivariate Cox survival (clinical survival not shipped as a runnable table), TIMER2 immune infiltration (external web tool output), miRWalk target prediction (external web DB output). One audit note: an early fabrication suspicion for ORC1L was OUR gene-ID error (ENSG00000113742=CPEB4 not ORC1; correct ORC1=ENSG00000085840) and was RESOLVED - ORC1L reproduces exactly. Net: the paper's core differential-expression claims are faithfully reproducible from its own shipped data via the stated standard pipeline.

💻 Code ↗ 🗄 Data: GSE40419

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 90
    assessed: 2026-06-15 ⛓ 093e294c34f8
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The high mortality of lung adenocarcinoma (LUAD) is linked to acquisition of progenitor-like cells with stem-like characteristics that help the tumor regulate immune cell infiltration; the study tests whether stemness-related genes and miRNAs in LUAD associate with immune/accessory cell infiltration rates and patient survival.

Core claims
  • Three stem cell-related genes (ORC1L, KIF20A, DLGAP5) are differentially expressed in LUAD and correlate with altered immune infiltration and reduced patient survival. finding
  • hsa-mir-1247-3p, via targeting of tumor suppressor SLC24A4 and oncogenes RAB3B and HJURP, together with MDSC infiltration modulation, primarily regulates LUAD patient survival. finding
  • Increased expression of stemness-related genes in LUAD correlates with reduced immune cell infiltration and reduced patient survivability. finding
  • Stemness genes ORC1L, KIF20A, DLGAP5 positively correlate with Th2 and MDSC infiltration, while ORC1L negatively correlates with HSC infiltration. finding
  • hsa-mir-1247-3p and/or its gene targets are proposed as candidate biomarkers/therapeutic avenue to enhance LUAD patient survival. resource
  • A bioinformatics mining pipeline (differential expression, univariate/multivariate Cox survival, immune deconvolution, miRNA-3'UTR target prediction) was used to identify markers across TCGA and GEO cohorts. method
Experimental setups
Assay System Perturbation Readout Platform
Bulk RNA-seq differential expression (DeSeq2) LUAD tumor (537) vs normal (59) human samples none (tumor vs normal observational) differentially expressed genes (log2 fold change, BH adjusted p-value) TCGA database
miRNA-seq differential expression (DeSeq2) LUAD tumor vs normal human samples none differentially expressed miRNAs (LFC, BH adjusted p) TCGA database
Microarray gene expression differential expression (Limma) validation LUAD tumor vs normal (GSE40419: 87 tumor/77 normal; GSE31210: 226 tumor/20 normal) none differentially expressed genes (LFC, BH adjusted p) GEO microarray datasets
Univariate Cox-proportional hazard survival analysis LUAD patient clinical + gene expression data (high vs low expression groups) none hazard ratio, p-value for gene expression vs survival Timer2.0
Immune/accessory cell infiltration correlation analysis LUAD TCGA tumor samples, 32 immune/accessory cell types none correlation between gene expression and cell infiltration rate; hazard ratio for infiltration vs survival Timer2.0 (Timer, TIDE, MCP-counter, quanTIseq, xCell, CIBERSORT, Epic)
miRNA–3'UTR target binding prediction differentially expressed stemness-related genes (3'UTR regions) none predicted/validated miRNA-gene binding (3p/5p variants) mirWalk and miRTarBase
Multivariate Cox-proportional hazard survival analysis LUAD patients; covariates = gene/miRNA expression, MDSC/Th2/HSC infiltration, age, cancer stage none hazard ratio, p-value for each covariate vs survival R Survival Analysis package (MDSC from TIDE; HSC, Th2 from Timer2.0)
Key results
  • Four stemness genes (ORC1L, KIF20A, DLGAP5, RAB3B) differentially expressed in LUAD with hazard ratio >1 (DLGAP5 HR 1.314, KIF20A HR 1.314, ORC1L HR 1.263, RAB3B HR 1.181) DLGAP5 LFC 3.671 (TCGA DeSeq2); HR 1.314
  • ORC1L, KIF20A, DLGAP5 positively correlate (>0.6) with Th2 and MDSC infiltration; ORC1L negatively correlates (<-0.6) with HSC infiltration correlation >0.6 or <-0.6
  • High MDSC infiltration strongly associated with lower patient survival HR=62.036 (p=0)
  • High Th2 infiltration associated with lower patient survival HR=3.804 (p=0.006)
  • High HSC infiltration associated with higher patient survival HR=0.167 (p=0.046)
  • Five miRNAs (hsa-mir-197-3p, 1247-3p, 122-5p, 215-5p, 192-5p) differentially expressed targeting RAB3B 3'UTR; only hsa-mir-1247-3p met both DeSeq2 and Limma cutoffs hsa-mir-1247-3p DeSeq2 LFC -2.324, Limma LFC -3.151
  • Multivariate analysis: increased hsa-mir-1247-3p expression and increased MDSC infiltration significantly associated with loss of survival; no correlation between MDSC and miR-1247-3p miR-1247-3p HR 0.667653 p=0.02279; MDSC HR 0.655898 p=0.028825; Pearson r=-0.05
  • hsa-mir-1247-3p targets HJURP and RAB3B (LFC >2 across all datasets) and SLC24A4 (LFC <-2 across all datasets) LFC >2 or <-2
Key statistics
  • fold_change hsa-mir-1247-3p DeSeq2 LFC -2.324 (BH p=3.51E-19), Limma LFC -3.151 (BH p=1.63E-11) (miRNA downregulation in LUAD)
  • other Hazard ratio 62.036 (p=0) (MDSC high vs low infiltration univariate survival)
  • other Hazard ratio 3.804 (p=0.006) (Th2 high vs low infiltration survival)
  • other Hazard ratio 0.167 (p=0.046) (HSC high vs low infiltration survival)
  • pvalue HR 0.667653, p=0.02279 (hsa-mir-1247-3p in multivariate Cox model)
  • correlation Pearson = -0.05 (MDSC infiltration vs hsa-mir-1247-3p (no significant correlation))
  • count 537 LUAD tumor and 59 normal samples (TCGA cohort)
  • fold_change DLGAP5 LFC 3.671 (DeSeq2), 5.411/4.184/3.003 (Limma), HR 1.314 p=0.0 (top stemness gene differential expression and survival)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This retrospective bioinformatics study used TCGA RNA-seq data (537 tumor, 59 normal) and two GEO microarray validation datasets to identify stemness-related genes and miRNAs differentially expressed in lung adenocarcinoma. Differential expression was assessed with DESeq2 (RNA-seq) and limma (microarray), both with Benjamini-Hochberg FDR correction at adjusted p ≤ 0.05 and |log2FC| > 2. Univariate and multivariate Cox proportional hazard models were then used to evaluate associations between gene/miRNA expression, computationally predicted immune cell infiltration rates, and patient survival, with results reported as hazard ratios and exact BH-adjusted p-values.

Replicationbiological Sample sizeSample sizes stated per dataset: TCGA (537 tumor, 59 normal), GSE40419 (87 tumor, 77 normal), GSE31210 (226 tumor, 20 normal); no formal power calculation described GroupsLUAD tumor vs. normal for differential expression; high vs. low gene/miRNA expression groups and high vs. low immune cell infiltration groups for survival analysis Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg (BH) FDR
Statistical tests used
Test Applied to n Assumptions
DESeq2 Wald test Differential expression of genes and miRNAs in TCGA LUAD tumor vs. normal samples 537 tumor + 59 normal (TCGA) not stated
limma moderated t-test Differential expression of genes in GEO microarray datasets GSE40419 and GSE31210 used as validation cohorts 87 tumor + 77 normal (GSE40419); 226 tumor + 20 normal (GSE31210) not stated
Univariate Cox proportional hazard model Association between high vs. low gene expression or immune cell infiltration and LUAD patient survival, performed via Timer2.0 not stated
Multivariate Cox proportional hazard model Joint association of miRNA/gene expression, immune infiltration rates, patient age, and cancer stage with patient survival (R Survival package) not stated
Pearson correlation Association between stemness gene expression and computationally predicted infiltration rates of 32 immune/accessory cell types; also reported for MDSC vs. hsa-mir-1247-3p (r = -0.05) not stated
Kaplan-Meier survival curve Visual display of survival probability for high vs. low immune cell infiltration and gene expression groups, obtained from Timer2.0 na
Approaches that could also have been used
  • Patient samples were dichotomized into high and low expression (or infiltration) groups before Cox proportional hazard model analysis
    Could also: Cox regression treating gene/miRNA expression as a continuous variable, or using a restricted cubic spline to model non-linear dose-response — Continuous Cox modeling preserves information across the full expression range, avoids sensitivity to the choice of cutpoint, and is generally more statistically efficient than a median-split dichotomization
  • Pearson correlation was used to assess associations between gene expression and computationally predicted immune cell infiltration rates across 32 cell types
    Could also: Spearman rank correlation — Spearman correlation does not assume linearity or bivariate normality, assumptions that may not hold for RNA-seq-derived values and computationally estimated infiltration scores
  • Hazard ratios from Cox proportional hazard models were reported with p-values but without confidence intervals
    Could also: Report 95% confidence intervals alongside each hazard ratio — Confidence intervals directly convey the precision and uncertainty of each hazard ratio estimate, complementing the p-value and making effect magnitude more interpretable
  • Correlations between gene expression and 32 immune cell types were screened using a |r| > 0.6 threshold with nominal p ≤ 0.05, without an explicit multiplicity correction for the 32 cell-type tests
    Could also: Apply BH-FDR correction across the family of 32 correlation tests per gene — Testing 32 cell types per gene increases the probability of finding at least one spurious correlation at a nominal threshold; a within-family correction would be consistent with the approach applied to the differential expression and survival analysis steps
  • Univariate Cox models were used to pre-screen individual genes and immune cell types for survival associations before entry into the multivariate model
    Could also: LASSO-penalized Cox regression for simultaneous covariate selection in a single multivariate model — LASSO-penalized Cox regression selects predictors jointly, reducing the potential for bias introduced by sequential univariate pre-screening and handling correlated predictors more explicitly
  • DESeq2 was the sole method used for RNA-seq differential expression in the TCGA dataset
    Could also: edgeR (quasi-likelihood F-test) or voom/limma applied to RNA-seq count data as a concordance check — DESeq2, edgeR, and voom/limma are all widely validated for RNA-seq differential expression; reporting genes that are significant across two methods is a common strategy for increasing confidence in findings, analogous to the cross-platform validation the paper applied to the microarray data
Software: DESeq2 · limma · R (survival package) · Timer2.0 · mirWalk

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
2
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE31210 GEO in Methods (http://purl.org/orb/Methods)
also used by 3 papers:
GSE40419 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37948380

Paper: Limbu & McCloskey 2023, PLoS One. "Stemness genes and miR-1247-3p expression associate with clinicopathological parameters and prognosis in LUAD." DOI 10.1371/journal.pone.0294171 · PMCID PMC10637681.

Code/Data artifact: https://github.com/Sashoss/LUAD_PLOS_data @ commit 8b6d7e125fdbb20685a32a9246bc7f2f4846784f (2023-09-01). This is a data + derived-results repository — it ships NO authors' analysis scripts (no .R/.py/.ipynb/README). Per brief rule P16, reproduction = re-run the standard, paper-described pipeline on the shipped inputs and compare to the shipped derived outputs + paper claims. This is fully valid.

Pipeline (from paper Methods)

  • Differential expression:
    • TCGA gene & miRNA counts → DESeq2.
    • GEO microarray/RNA-seq (GSE40419, GSE31210) → Limma.
    • DEG cutoff: Benjamini-Hochberg adj p ≤ 0.05 AND |log2FC| > 2. FC = cancer/normal.
  • Univariate & multivariate Cox proportional-hazard survival analysis (high vs low expr).
  • TIMER2 immune-infiltration correlation; miRWalk target prediction.

Shipped inputs (on «infra», complete, sizes match GitHub)

  • data_gene.txt (99 MB): TCGA STAR raw gene counts, ENSG ids, 596 samples (537 Primary_Tumor + 59 Solid_Tissue_Normal; condition encoded in column-name prefix).
  • data_mirna.txt (2.5 MB): TCGA miRNA raw counts, 565 samples (519 T + 46 N).
  • GSE40419-GPL11154_withGeneName.csv (17 MB) + GSE40419_annotation.csv (87 T / 77 N).
  • GSE31210-GPL57_withGeneName.csv (53 MB) + GSE31210_annotation.csv (226 T / 20 N).
  • Metadata_LUAD_gene.tsv, Metadata_LUAD_mirna.tsv (clinical/sample metadata).

Shipped derived outputs (comparison ground-truth)

  • TCGA_diffExp_gene_DeSeq2.csv (full transcriptome), TCGA_diffExp_mirna_DeSeq2.csv.
  • GSE40419_diffExp_Limma.csv, GSE31210_diffExp_Limma.csv, TCGA_diffExp_gene_Limma.csv (latter three restricted to a ~15-gene stemness panel).
  • multivariate_survival_analysisData_Table1/2.csv, Timer2_TCGA_infiltration_data.csv.

IN SCOPE (attempted — clearly specified pipeline, inputs+outputs shipped)

  1. TCGA DESeq2 gene diff-exp (tumor vs normal) → compare to TCGA_diffExp_gene_DeSeq2.csv.
  2. TCGA DESeq2 miRNA diff-exp → compare to TCGA_diffExp_mirna_DeSeq2.csv (key: hsa-mir-1247).
  3. GSE40419 Limma diff-exp → compare to GSE40419_diffExp_Limma.csv (stemness panel).
  4. GSE31210 Limma diff-exp → compare to GSE31210_diffExp_Limma.csv (stemness panel).

Key reported claims compared: 3 stemness genes ORC1L/KIF20A/DLGAP5 up in LUAD; hsa-mir-1247-3p down; targets SLC24A4 (down), RAB3B/HJURP (up).

OUT OF SCOPE (80/20 — not attempted, reasons)

  • Univariate/multivariate Cox survival (Table 1/2): requires cBioPortal-sourced clinical survival + stage/age not fully shipped in a single runnable table; lower priority than diff-exp; skipped per 80/20.
  • TIMER2 immune infiltration: TIMER2 is an external web tool; shipped CSV is its output, not reproducible without re-querying the web service. Out of scope.
  • miRWalk target prediction: external web DB output (xlsx), not a runnable pipeline.
  • Wet-lab / manual interpretation claims: out of scope by definition.
Figures / tables: Table
TCGA_DESeq2_DLGAP5
Reported
3.671
Reproduced
3.671
exact
TCGA_DESeq2_KIF20A
Reported
3.143
Reproduced
3.144
exact
TCGA_DESeq2_ORC1L
Reported
2.741
Reproduced
2.741
exact
TCGA_DESeq2_RAB3B
Reported
4.242
Reproduced
4.242
exact
TCGA_DESeq2_CENPI
Reported
2.740
Reproduced
2.740
exact
TCGA_DESeq2_hsa-mir-1247
Reported
down (negative LFC)
Reproduced
-2.325
exact
TCGA_DESeq2_gene_transcriptome
Reported
(shipped DESeq2 CSV)
Reproduced
Pearson cor(log2FC)=1.000, n=34924
exact
TCGA_DESeq2_miRNA_full
Reported
(shipped miRNA DESeq2 CSV)
Reproduced
Pearson cor(log2FC)=1.000, n=1359
exact
GSE40419_Limma_KIF20A
Reported
3.912
Reproduced
3.955
within tolerance
GSE40419_Limma_panel
Reported
(shipped Limma CSV)
Reproduced
cor(logFC)=0.70-0.75, n=15
partial
GSE31210_Limma_panel
Reported
(shipped Limma CSV)
Reproduced
cor(logFC)=0.9975, n=15; magnitudes ~1.45x low
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 90/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

The paper's core differential-expression claims reproduce 1:1 from its own shipped TCGA data: whole-transcriptome and whole-miRNA log2FC correlate at 1.000 and every Table 1 headline stemness gene plus miR-1247-3p matches to 3 decimals, so q5/q7 are fully on the authors' side as confirmed. The only deviations sit in the GEO validation cohorts, where Limma magnitudes run ~1.45x low (GSE31210) or inconsistent (GSE40419 RAB3B 3.152→0.494) — direction and ranking still hold (GSE31210 cor=0.9975). This gap is on our-method/availability side, caused by the authors shipping no analysis code and not documenting their microarray preprocessing, not by fabrication (the early ORC1L suspicion was our own gene-ID error, resolved). Overall: a strong primary reproduction with moderate, fully explainable secondary-cohort magnitude deviations → yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

191.1 k
tokens (I/O) · 15.3 M incl. cache
26 min
runtime · 0.1 CPU-h
6.3 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine