Constructing an APOBEC-related gene signature with predictive value in the overall survival and therapeutic sensitivity in lung adenocarcinoma.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for the HEADLINE result and reproduced ~1:1. The repo (github.com/AnxinGu/raw-data @ 8865db3) ships the authors' actual 4317-line R analysis script plus raw TCGA GDC data and figure PDFs, but it is not turnkey: it load()s several unshipped intermediate .RData objects and calls two undefined in-house helper functions (createCoxModel / predictScoreByCoxModel), so an exact end-to-end re-run from the repo alone is impossible. We therefore took the published 4-gene APOBEC signature (LYPD3, ANLN, MUC5B, FOSL1) as given and tested whether its reported TCGA-LUAD performance reproduces on independent public data (UCSC Xena GDC STAR-TPM, log2(TPM+1), n=503, 182 events), using the authors' exact downstream recipe (refit multivariate Cox -> AMrs risk score; timeROC cause=1/marginal at 1/3/5 yr; optimal cutpoint via maxstat for KM). Result: time-ROC AUC reproduced as 0.715 / 0.698 / 0.624 vs reported 0.73 / 0.69 / 0.62 (all within +/-0.015), and KM high-vs-low log-rank p = 1.02e-13 satisfies the reported P<0.0001. All four genes were retained with positive (risk) coefficients, consistent with Fig 7D. Grade 'partial' because the hard ~20% was deliberately NOT attempted: the upstream AMES(MAF)->162 DEGs->72 univariate-prognostic->LASSO chain that derives the 4 genes (needs unshipped .RData + undefined functions + a fixed random seed over a cohort we cannot rebuild bit-for-bit), the GEO validation cohorts (GSE72094/GSE31210/GSE50081 — no exact AUC quoted in text, figure-only), and the wet-lab RT-qPCR + secondary drug/immune/nomogram analyses (Figs 5-12). Note the coefficients were re-estimated on the same TCGA cohort, so this confirms the reported performance numbers but is not an independent test of the LASSO gene-selection. No fabrication indicators: the headline numbers are reproducible within tolerance from the shipped gene list + public data despite a different (STAR vs deprecated htseq) RNA-seq quantification.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 88assessed: 2026-06-15 ⛓ ebaa82ce25e5
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusDoes dysregulation of the APOBEC family and APOBEC-induced mutagenesis (quantified by an APOBEC mutagenesis enrichment score, AMES) relate to genomic instability in lung adenocarcinoma (LUAD), and can an APOBEC-related gene signature predict overall survival and therapeutic sensitivity?
- ★ APOBEC family genes are aberrantly expressed across cancers, with APOBEC3B the most consistently upregulated, and APOBEC3B/APOBEC3A are markedly altered in LUAD finding
- ★ In LUAD, AMES is positively correlated with intratumor heterogeneity (ITH), tumor neoantigen burden (TNB), and tumor mutation burden (TMB), linking APOBEC mutagenesis to genomic instability finding
- ★ High AMES tumors show more DNA damage repair (DDR) pathway mutations and enrichment of cell cycle and DNA repair pathways finding
- ★ A four-gene AMES-related prognostic signature (LYPD3, ANLN, MUC5B, FOSL1) was constructed via univariate Cox and LASSO regression method
- ★ The AMES-related gene signature (AMrs) predicts LUAD overall survival and guides immunotherapy/chemotherapy decisions resource
- ★ Knockdown of LYPD3, ANLN, MUC5B, and FOSL1 reduces lung cancer cell invasion and viability, supporting their pro-tumor role finding
- High-AMrs group shows increased oxidative stress level finding
- The 'APOBEC Cytidine Deaminase (C>T)' mutational signature has the strongest correlation with AMES among NMF-derived signatures mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq / transcriptomic expression analysis | TCGA pan-cancer (10535 samples) and LUAD tumor vs adjacent normal tissues; GTEx normal tissues | none | APOBEC family gene expression levels | UCSC Xena / TCGA, GTEx |
| somatic mutation / mutagenesis analysis (AMES, TMB, ITH, TCW motif) | TCGA-LUAD samples (565 LUAD samples) | none | AMES, TMB, ITH, TNB, mutation counts, CNV, driver genes | mutect2 MAF, maftools, OncodriveCLUST, NMF/COSMIC |
| microarray gene expression (validation) | GEO LUAD cohorts GSE72094, GSE31210, GSE50081 | none | AMrs risk score and overall survival | GEO microarray |
| immune cell infiltration deconvolution (CIBERSORT) | TCGA-LUAD | none | immune cell proportions, GEP/Th1/cytolytic scores correlated with AMrs | CIBERSORT |
| qRT-PCR | BEAS-2B, A549, H2009 cell lines | siRNA knockdown of LYPD3, ANLN, MUC5B, FOSL1 (vs siNC) | target gene mRNA expression (GAPDH reference) | ABI 7500, SYBR Green Master Mix |
| Transwell invasion assay | BEAS-2B, A549, H2009 cell lines | siRNA knockdown of LYPD3, ANLN, MUC5B, FOSL1 | number of invading cells through Matrigel | Matrigel (BD Biosciences), crystal violet |
| CCK-8 cell viability assay | A549, H2009 (and BEAS-2B) cell lines | siRNA knockdown of LYPD3, ANLN, MUC5B, FOSL1 | OD450 cell viability | Cell Counting Kit-8 (Beyotime), microplate reader |
| GSEA pathway enrichment | TCGA-LUAD high vs low AMES groups | none | enriched KEGG/Hallmark/DDR pathways | MSigDB, GSEA |
- ▲ APOBEC3B expression drastically upregulated in most cancer types except READ and ALL; highest expression change among APOBEC genes in LUAD P = 3.5e-09 (APOBEC3B LUAD vs normal)
- ▲ AMES positively correlated with counts of non-synonymous and synonymous mutations in LUAD R=0.15, P<0.001; R=0.17, P<0.001
- ▲ TMB, ITH, and TNB scores increase with increasing AMES; all strongly correlated with APOBEC3B
- ▲ High AMES group has the most DDR pathway mutations; DDR-mut group has significantly higher AMES than DDR-WT P<0.05
- ▲ 'APOBEC Cytidine Deaminase (C>T)' NMF signature most strongly correlated with AMES R=0.94
- ▼ Knockdown of LYPD3, ANLN, MUC5B, FOSL1 decreased invasion ability and viability of lung cancer cells
- – Seven APOBEC family genes differentially expressed between LUAD and paracancerous samples e.g. APOBEC1 P=4.6e-07, APOBEC4 P=1.6e-07
- – KRAS identified as driver gene in all three AMES groups; moderate AMES highest KRAS mutation rate, high AMES lowest
- pvalue P = 3.5e-09 (APOBEC3B differential expression LUAD vs paracancerous)
- correlation R=0.94 (APOBEC Cytidine Deaminase (C>T) signature vs AMES)
- correlation R=0.15, P<0.001 (AMES vs non-synonymous mutation counts)
- correlation R=0.17, P<0.001 (AMES vs synonymous mutation counts)
- pvalue P = 4.6e-07 (APOBEC1 LUAD vs paracancerous)
- pvalue P = 1.6e-07 (APOBEC4 LUAD vs paracancerous)
- count 27 of 565 samples mutated (APOBEC family mutations in LUAD samples)
- count 162 DEGs (95 up, 67 down) (DEGs between high vs low AMES groups)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper combines a multi-stage bioinformatics analysis of TCGA pan-cancer and LUAD data (with three independent GEO validation sets) with in vitro cell-line experiments. APOBEC Mutagenesis Enrichment Score (AMES) was computed and LUAD samples were stratified into three AMES groups; group comparisons relied on Wilcoxon rank-sum, ANOVA, or Kruskal-Wallis tests and Fisher exact tests. A prognostic risk score (AMrs) was derived through sequential univariate Cox regression and LASSO regression on differentially expressed genes, then evaluated by Kaplan-Meier survival analysis, time-dependent ROC curves, and a nomogram with decision curve analysis. Results were reported with exact p-values for selected comparisons and asterisk threshold notation in figures.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Wilcoxon rank-sum test | All two-group comparisons: APOBEC family expression in LUAD vs. adjacent normal tissues (Fig. 1C/E); DDR-mut vs. DDR-WT AMES comparison (Fig. 2K); general two-group comparisons throughout | 565 LUAD samples stated for somatic mutation analyses; paired normal sample count not separately stated | not stated |
| One-way ANOVA | Comparison of TMB, ITH, and TNB scores across low/moderate/high AMES groups (Fig. 2D) | TCGA-LUAD cohort; per-group n not stated | not stated |
| Kruskal-Wallis test | Stated as an alternative to ANOVA for three-or-more-group comparisons (statistical analysis section); specific figure application not specified | null | not stated |
| Fisher exact test | Significantly mutated genes comparing high AMES group against the other two groups (Fig. 2I) | null | na |
| Pearson correlation | AMES vs. synonymous/non-synonymous mutation counts (Fig. 2E); APOBEC family gene expression vs. TMB/TNB/ITH/AMES (Fig. 2F); AMrs vs. immune cell infiltration | null | not stated |
| Univariate Cox proportional hazards regression | Screening DEGs (high vs. low AMES) for overall survival association; genes with P < 0.05 retained | TCGA-LUAD; exact n not stated | not stated |
| LASSO (least absolute shrinkage and selection operator) penalized Cox regression | Feature selection among univariate Cox-significant prognostic genes to construct the AMrs four-gene signature | null | na |
| Kaplan-Meier survival analysis (log-rank test implied) | Overall survival comparison between low-AMrs and high-AMrs groups in TCGA-LUAD training set and three GEO validation sets | TCGA-LUAD training set and GSE72094, GSE31210, GSE50081; individual n not stated in methods | not stated |
| Time-dependent ROC curve analysis (TimeROC R package) | Assessment of AMrs predictive performance for overall survival | null | na |
| limma linear model (differential expression) | Identification of DEGs between low- and high-AMES groups in TCGA-LUAD | null | not stated |
| Gene Set Enrichment Analysis (GSEA) | Enrichment of KEGG and Hallmark pathways in high vs. low AMES groups | null | na |
-
The AMrs cut-off for dichotomizing patients into low/high risk was determined by Survminer, which selects the threshold maximizing the observed survival difference in the training data↳ Could also: A pre-specified cut-off (e.g., median, quartile) or analyzing AMrs as a continuous variable in a multivariable Cox model could also be applied — Data-driven optimal cut-point selection tends to inflate the apparent separation between survival curves; pre-specified or continuous approaches are less susceptible to this and are common recommendations in prognostic biomarker reporting guidelines (e.g., TRIPOD)
-
Multiple simultaneous Wilcoxon rank-sum tests were performed across 11 APOBEC family genes comparing tumor vs. normal tissues without a stated family-wise or FDR correction↳ Could also: Applying Benjamini-Hochberg FDR correction or Bonferroni correction across the 11-gene family of tests could also control the false discovery rate — Correcting for the number of simultaneous tests reduces the probability that any individual significant result is a chance finding; reporting adjusted alongside raw p-values is standard in multi-gene expression analyses
-
Pearson correlation was used to assess relationships between AMES and mutation burden indicators (TMB, ITH, TNB), which are typically right-skewed count variables↳ Could also: Spearman rank correlation could also be applied, as it does not require bivariate normality and is robust to the heavy-tailed distributions typical of mutation burden data — Spearman correlation makes no distributional assumption and is less sensitive to extreme outliers, which are common in TMB/ITH distributions; reporting which assumption was checked or using Spearman as a sensitivity analysis is common practice
-
ANOVA was used to compare TMB, ITH, and TNB across three AMES groups; no post-hoc pairwise test with multiplicity correction is described↳ Could also: Pairwise post-hoc comparisons with correction (e.g., Tukey HSD after ANOVA, or Dunn's test with Bonferroni/BH correction after Kruskal-Wallis) could also identify which specific group pairs drive the omnibus result — An omnibus test confirms that groups are not all equal but does not identify which pairs differ; post-hoc tests with multiplicity correction provide that specificity and are standard when three or more groups are compared
-
Prognostic gene selection used a two-step pipeline: univariate Cox filtering at P < 0.05 followed by LASSO, without stated cross-validation or bootstrap resampling to assess selection stability or optimize the shrinkage parameter↳ Could also: Cross-validated LASSO (e.g., cv.glmnet with minimum-lambda or 1-SE rule), or bootstrap stability selection, could also estimate the optimal penalty and assess which genes are consistently selected across resamples — Univariate pre-filtering before LASSO can introduce selection bias; cross-validation or bootstrap approaches provide an estimate of model generalizability and are recommended in prognostic signature development frameworks
-
The in vitro gene knockdown experiments (RT-qPCR, CCK-8, Transwell) do not report the number of biological or technical replicates, measures of dispersion, or the specific statistical test used for group comparisons↳ Could also: Reporting the number of independent biological replicates (conventionally ≥ 3), a dispersion measure (SD or SEM), and the test applied (e.g., two-tailed Student's t-test or one-way ANOVA with post-hoc correction) would also be standard for cell-based assays — Replicate counts and dispersion allow readers to assess experimental reproducibility and statistical power; this reporting is required by most journals and facilitates meta-analysis or replication
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37954334
Paper: Luo et al. 2023, Heliyon e21336. "Constructing an APOBEC-related gene signature with predictive value in the overall survival and therapeutic sensitivity in lung adenocarcinoma."
Code: https://github.com/AnxinGu/raw-data @ commit
8865db37ffc1c2a3072a7cb670f3721975deb83c (2023-06-29). The repo ships the
authors' actual analysis script (scripts/20220530_LUAD_APOBEC.R, 4317
lines) plus raw TCGA GDC downloads (UUID dirs), processed GEO .RData, figure
PDFs, and gene lists. It is not a turnkey pipeline: it load()s several
intermediate .RData objects that are NOT in the repo (e.g.
origin_datas/TCGA/luad.tcga.exp.RData, origin_datas/GEO/gse72094.exp.RData)
and calls two undefined, in-house helper functions —
createCoxModel() and predictScoreByCoxModel() — that are not defined in
either shipped .R file (sourced from an unshipped private package). So an
exact end-to-end re-run is not possible from the repo alone.
Data
- TCGA-LUAD (training cohort, RNA-seq FPKM + OS clinical) — public via GDC /
UCSC Xena GDC hub. Same source the authors used (
Merge_TCGA-LUAD_FPKM.txt). - GSE72094, GSE31210, GSE50081 (GEO validation cohorts) — public.
The pipeline (from Methods §2.7-2.8 + the R script)
- APOBEC mutation enrichment score (AMES) from TCGA somatic MAF (maftools).
- AMES High vs Low/Moderate → limma DEGs (162: 95↑/67↓).
- Univariate Cox on DEGs → 72 prognostic genes (p<0.05).
- LASSO Cox (
cv.glmnet, family="cox",set.seed(888826666), lambda.min) → λ=0.0659, 4 genes: LYPD3, ANLN, MUC5B, FOSL1. - Multivariate (stepwise) Cox on the 4 genes → coefficients → AMrs = Σ Exp(i)·coef(i) (risk score).
- Evaluate: time-ROC AUC @1/3/5 yr (TCGA Fig 4B: 0.73 / 0.69 / 0.62),
KM by optimal cutpoint (
surv_cutpoint), logrank p<0.0001.
IN SCOPE (attempted — clearly specified, low-hanging pipeline outputs)
- C1 The 4-gene signature (LYPD3, ANLN, MUC5B, FOSL1) — taken as the published signature; confirm it is prognostic on TCGA-LUAD public data.
- C2 time-ROC AUC @1/3/5 yr on TCGA-LUAD (reported 0.73/0.69/0.62, Fig 4B),
computed with the authors' exact recipe (
timeROC, cause=1, marginal, times=1,3,5 yr; AMrs = linear predictor of multivariate Cox of the 4 genes on log2(TPM+1) tumor expression). - C3 KM separation of AMrs high vs low on TCGA-LUAD (reported logrank
p<0.0001, Fig 4C), optimal cutpoint via
surv_cutpoint. - C4 (validation, stretch) repeat KM/AUC on GSE72094 (the RU's named dataset).
OUT OF SCOPE (not attempted, with reason)
- Full upstream (steps 1-4 above): exact AMES/MAF→DEG→univariate→LASSO chain
to re-derive the 4 genes. This is the hard ~20%: depends on the unshipped
intermediate
.RData, undefined helper functions, and a random seed over a cohort we cannot reconstruct bit-for-bit. We instead take the 4 published genes as given and test whether the reported performance reproduces. Re-deriving the exact gene list is explicitly NOT claimed. - Wet-lab RT-qPCR validation (Fig: RT-qPCR.pdf) — experimental, non-pipeline.
- Drug-sensitivity / immune-infiltration / nomogram secondary analyses (Figs 5-12) — pipeline-derived but secondary; not attempted under 80/20.
Honesty note
Reproducing C2/C3 by refitting the 4-gene Cox on the same TCGA cohort is a partial test (it re-estimates coefficients rather than re-deriving the gene set). It directly checks the headline reported AUC/KM numbers (Fig 4B/4C) but does NOT independently validate that LASSO selects these 4 genes. Graded accordingly.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The paper's headline computational claim — the 4-gene APOBEC signature (LYPD3/ANLN/MUC5B/FOSL1) and its TCGA-LUAD prognostic performance — reproduces essentially 1:1: time-ROC AUC 0.715/0.698/0.624 vs reported 0.73/0.69/0.62 (all ≤0.015) and KM log-rank p=1.02e-13 satisfies the reported P<0.0001, with all four genes retained as risk genes. Deviations are minor and sit on our/input side (UCSC Xena STAR-TPM vs the authors' deprecated htseq quantification, self-filtered n=503, coefficients refit on the same cohort), not on the authors' side. No fabrication indicators. It is marked yellow overall rather than green only because the repo was not turnkey (unshipped .RData + undefined helper functions) so the upstream LASSO gene-selection that derives the signature could not be independently tested.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.