Identification of a novel lncRNA prognostic signature and analysis of functional lncRNA AC115619.1 in hepatocellular carcinoma.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduction of Long et al. 2023 (Front Pharmacol), a TCGA-LIHC lncRNA prognostic-signature paper; named code artifact = the third-party tool pRRophetic (github paulgeeleher/pRRophetic, applied to the paper's own data per P16). Described well enough to reproduce without author contact; the core, clearly-specified pipeline outputs reproduce essentially 1:1. C1 (TCGA-LIHC 374 tumour [371 primary + 3 recurrent] + 50 normal) and C2 (GSE138178 49+49) reproduce EXACTLY by counting GDC STAR-Counts and GEO series-matrix samples. C3: all six signature lncRNAs (LINC02428, LINC02163, AC008549.1, AC115619.1, CASC9, LINC02362) are present in TCGA-LIHC and their univariate-Cox OS directions match Figure 2A 6/6 (3 protective + reported risk both confirmed). C4: AC115619.1 OS hazard ratio reproduces within tolerance — HR 0.542 (95% CI 0.38-0.77) vs reported 0.56 (0.39-0.79). C5 (named tool, secondary): pRRophetic 0.5 / cgp2014 predicted IC50 on TCGA-LIHC tumours split by AC115619.1 median reproduces the paper's claimed direction for 5/7 testable drugs (Gemcitabine, Rapamycin, Vinorelbine significantly; Paclitaxel, Sorafenib non-significantly), with Sunitinib significantly opposite and Imatinib essentially equal; 5-Fluorouracil is not in the cgp2014 training panel. Graded partial because the paper prints no numeric IC50 (figure-only), so only direction is checkable, and 2/7 diverge. NOT attempted (hard 20%, see scope.md): exact ROC AUC (1y train 0.711 / test 0.743; depends on the authors' random train/test split + unshipped LASSO coefficients), exact LASSO/Cox coefficients and the risk-score formula (no published values to compare), the ceRNA network, CIBERSORT immune deconvolution, and all wet-lab (43-pair qRT-PCR Fig 9A, CCK-8/EdU). No fabrication concern: every reported value we could pin is independently derivable from the public TCGA-LIHC + GEO data; the small C4 HR difference and the mixed C5 directions are expected from cohort/parameter choices, not evidence of error. Technical notes for auditors: pRRophetic ships no training data on GitHub (got pRRophetic_0.5 with GDSC data from OSF 5xvsg), the compute nodes cannot reach github/osf (staged via front1), and pRRophetic 0.5 needed a one-line patch (class(x)==... -> class(x)[1]==...) for R 4.5 compatibility. All grades provisional; human reviewer decides via AUDIT.md.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 87assessed: 2026-06-14 ⛓ 18624fba26d0
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study aimed to establish a reliable lncRNA-based prognostic signature for hepatocellular carcinoma (HCC) and to identify and characterize the function of a novel biomarker lncRNA, AC115619.1, in HCC progression.
- ★ A six-lncRNA prognostic signature (LINC02428, LINC02163, AC008549.1, AC115619.1, CASC9, LINC02362) predicts overall survival in HCC patients finding
- ★ AC115619.1 expression is lower in HCC tumor tissues than in adjacent normal tissues finding
- ★ AC115619.1 expression is negatively correlated with the m6A regulator RBMX finding
- ★ Overexpression of AC115619.1 inhibits proliferation, migration, and invasion of HCC cell lines finding
- ★ 32 DElncRNAs were identified by integrating differential expression results from GEO microarray and TCGA RNA-seq datasets method
- ★ AC008549.1, AC115619.1, and CASC9 are independent prognostic factors for HCC by multivariate Cox regression finding
- ★ AC115619.1 expression is significantly related to clinicopathologic features and predicted drug sensitivity finding
- A risk score model was constructed as the sum of coefficient-weighted expression values of the six prognostic lncRNAs method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| lncRNA microarray (probe re-annotation, limma/RRA differential expression) | HCC tumor vs adjacent normal tissue (9 GEO datasets, GSE138178 etc.) | none | differentially expressed lncRNAs | various Affymetrix/Agilent GPL platforms (e.g., GPL21827, GPL16956) |
| RNA sequencing (bulk) | TCGA HCC tumor (374 samples) vs normal liver tissue (50 samples) | none | differentially expressed lncRNAs (FPKM) | — |
| Univariate, LASSO, and multivariate Cox regression | TCGA HCC patient cohort (training and test) | none | prognostic lncRNAs and risk score model | R packages (survival, glmnet) |
| Pearson/Spearman correlation analysis | TCGA HCC patients | none | correlation between prognostic lncRNAs and m6A regulator expression | — |
| qRT-PCR | 43 paired HCC tumor and adjacent normal tissue samples (patient cohort) | none | relative AC115619.1 mRNA expression (2^-ΔΔCt, normalized to GAPDH) | ViiA6 RT-PCR system (Applied Biosystems) |
| CCK-8 cell viability assay | SNU-449 and HepG2 HCC cell lines | AC115619.1 overexpression (pcLV3 vector) vs empty vector | cell viability (absorbance at 450 nm) | Cell Counting Kit-8 (Dojindo) |
| EdU incorporation assay | SNU-449 and HepG2 HCC cell lines | AC115619.1 overexpression vs empty vector | proliferation rate (EdU+ cells) | EdU kit (RiboBio), fluorescence microscope |
| Wound healing and Transwell invasion assays | SNU-449 and HepG2 HCC cell lines | AC115619.1 overexpression vs empty vector | cell migration and invasion capacity | Matrigel-coated Transwell inserts (Corning) |
- – 51 upregulated and 58 downregulated lncRNAs identified across 9 GEO microarray datasets |log2FC|>1, p<0.05
- – 322 dysregulated lncRNAs identified in TCGA (35 up, 287 down) |log2FC|>1, p<0.05
- – 32 DElncRNAs common to both GEO and TCGA datasets
- – 12 lncRNAs associated with HCC prognosis; LASSO selected six key prognostic lncRNAs
- – AC008549.1, AC115619.1, and CASC9 identified as independent prognostic factors by multivariate Cox regression p<0.05
- – Six-lncRNA risk model predicted overall survival with good accuracy in training and test cohorts training 1-year AUC=0.711; test 1-year AUC=0.743
- ▼ AC115619.1 expression was lower in HCC tissues than normal tissues
- ▼ AC115619.1 overexpression inhibited proliferation, migration, and invasion of HCC cells
- fold_change |log2FC| > 1, adjusted p < 0.05 (criteria for identifying DElncRNAs in GEO and TCGA)
- other 1-year AUC = 0.711 (ROC curve accuracy of prognostic model, training cohort)
- other 1-year AUC = 0.743 (ROC curve accuracy of prognostic model, test cohort)
- count 142 paired tumor/normal samples (samples pooled from 9 GEO microarray datasets)
- count 374 tumor samples and 50 normal samples (TCGA RNA-seq cohort)
- pvalue p < 0.05 (significance threshold for univariate/multivariate Cox regression of prognostic lncRNAs)
- count 43 paired tumor and adjacent normal tissue samples (local patient cohort used for qRT-PCR validation)
- fold_change |log2FC| > 2, p < 0.01 (criteria for filtering HCC-related differentially expressed mRNAs for ceRNA network)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study used bioinformatics methods applied to nine GEO microarray datasets and TCGA RNA-seq data to identify differentially expressed lncRNAs (limma/eBayes with robust rank aggregation across datasets), then applied a sequential Cox regression pipeline (univariate → LASSO → multivariate stepwise) on TCGA survival data to build a six-lncRNA prognostic risk score, validated via Kaplan–Meier log-rank tests and time-point ROC in randomly split training/test cohorts. In vitro functional experiments (proliferation, migration, invasion) in two HCC cell lines and qRT-PCR in 43 paired clinical specimens were analyzed with Student's t-test or one-way ANOVA, with continuous data summarized as mean ± SD.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma eBayes moderated t-statistic with robust rank aggregation (RRA) meta-analysis | Identification of DElncRNAs across nine GEO microarray datasets | 142 paired samples (nine GSE datasets combined) | not stated |
| limma eBayes moderated t-statistic | Identification of DElncRNAs in TCGA RNA-seq data (|log2FC|>1, adjusted p<0.05) | 374 tumor samples and 50 normal liver samples | not stated |
| Univariate Cox proportional hazards regression | Screening 32 DElncRNAs for association with overall survival in TCGA HCC cohort; also used for risk score vs. clinical characteristics | TCGA HCC patients with survival data (exact n not stated in available text) | not stated |
| LASSO Cox regression | Dimensionality reduction from 12 prognostic lncRNAs to six key lncRNAs | TCGA HCC patients with survival data (exact n not stated) | not stated |
| Multivariate stepwise Cox proportional hazards regression | Identifying independent prognostic lncRNAs and evaluating risk score with clinical covariates; HR and 95% CI reported | TCGA HCC patients with survival data (exact n not stated) | not stated |
| Kaplan–Meier survival analysis with log-rank test | Comparing overall survival between high- and low-risk subgroups (median risk score cutpoint) in training and test cohorts; also used for AC115619.1 expression subgroups | TCGA patients randomly split into training and test cohorts (proportions not stated) | not stated |
| Time-point ROC curve analysis (AUC) | Evaluating predictive accuracy of prognostic model at 1 year in training cohort (AUC=0.711) and test cohort (AUC=0.743) | TCGA training and test cohorts (exact n not stated) | not stated |
| Pearson correlation (methods section) / described as Spearman in abstract | Association between 12 prognostic lncRNAs and m6A-related regulators in TCGA data | TCGA HCC samples (exact n not stated) | not stated |
| Chi-squared (χ²) test | Comparing categorical clinicopathological variables (tumor grade, tumor invasion, TNM stage) across patient subgroups | TCGA patients (exact n not stated) | not stated |
| Student's t-test or one-way ANOVA | Differences between experimental groups in in vitro assays (CCK-8, EdU, wound healing, transwell); also used for qRT-PCR expression comparison in 43 paired clinical samples | 43 paired clinical samples for qRT-PCR; cell line replicate n not stated | not stated |
| CIBERSORT deconvolution algorithm | Estimating proportions of 22 immune cell types in high- vs. low-AC115619.1 patient groups | TCGA HCC samples (exact n not stated) | na |
-
The abstract describes the lncRNA–m6A regulator association analysis as 'Spearman's analysis,' while the methods section describes it as 'Pearson's correlation analysis'↳ Could also: Spearman rank correlation could be used in place of Pearson for this analysis — Spearman correlation makes no assumption of bivariate normality, which is commonly preferred for RNA expression data that may be skewed or contain outliers; the discrepancy between abstract and methods is worth noting as a transparency consideration
-
The TCGA cohort was randomly divided into training and test subsets for model validation, with performance reported at a single 1-year time point AUC↳ Could also: K-fold cross-validation (e.g., 5- or 10-fold) or repeated random splitting combined with time-dependent AUC (e.g., Uno's C-statistic or Harrell's C-index across the full follow-up) could also be used — K-fold cross-validation uses all data for both training and validation across folds, potentially providing more stable performance estimates when total n is limited; time-dependent AUC across the full follow-up summarizes discrimination over the entire observation period rather than a single landmark time
-
Patients were stratified into high- and low-risk groups using the median risk score as a fixed cutpoint↳ Could also: Optimal cutpoint methods (e.g., maximally selected rank statistics, time-dependent ROC-based cutpoint selection) could also be applied — Median splitting is straightforward and reproducible across cohorts; cutpoint optimization methods can identify the threshold with the strongest prognostic separation, though they require correction for the implicit multiple testing inherent in searching over thresholds
-
In vitro group comparisons (proliferation, migration, invasion) were analyzed with Student's t-test or one-way ANOVA without a stated post-hoc correction procedure↳ Could also: One-way ANOVA followed by a named post-hoc test (e.g., Tukey HSD, Dunnett's test vs. control) could also be used when multiple groups are compared — A named post-hoc procedure makes the family-wise error rate control explicit and is commonly expected in reporting multiple group comparisons; Dunnett's test specifically addresses multiple treatment-vs-control comparisons
-
LASSO regularization was used for variable selection in the Cox model prior to multivariate Cox regression↳ Could also: Elastic net Cox regression (combining L1 and L2 penalties) could also be applied for variable selection in this high-dimensional setting — Elastic net can handle correlated predictors (which co-expressed lncRNAs frequently are) better than pure LASSO by allowing groups of correlated variables to enter together rather than arbitrarily selecting one; both approaches regularize the model to reduce overfitting
-
Continuous data from in vitro experiments and clinical comparisons were summarized as mean ± SD↳ Could also: For small cell-line experiment groups where n per condition is not stated, 95% confidence intervals or individual data points overlaid on bar/box plots could also convey spread — When group sizes are small and potentially unequal, 95% CIs directly communicate estimation uncertainty around the mean; many journals now recommend showing individual data points alongside summary statistics to allow readers to assess distributional assumptions
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The core, clearly-specified pipeline reproduces essentially 1:1 from public data: TCGA-LIHC (424) and GSE138178 (98) sample counts are exact, all 6 signature lncRNA Cox directions match Fig 2A 6/6, and AC115619.1 OS HR 0.542 (CI 0.38–0.77) reproduces the reported 0.56 (0.39–0.79) (C4). The central prognostic conclusion is fully confirmed and every pinnable value is independently derivable from the shared GEO/GDC data — no fabrication concern. The only soft spot is the secondary, figure-only pRRophetic drug-sensitivity claim (C5): 5/7 testable drugs match direction but Sunitinib is significantly opposite, Imatinib is ns, and 5-FU is absent from cgp2014 — divergences that sit on our methodology choices (median split, default parameters, panel version), not on the authors' side.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.