Evaluating Distribution and Prognostic Value of New Tumor-Infiltrating Lymphocytes in HCC Based on a scRNA-Seq Study With CIBERSORTx.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡A deviation arose in the data or preprocessing
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the CORE DEG pipeline; the headline immune method (CIBERSORTx) and enrichment/prognostic steps are web-tool/stochastic and out of scope. C1 (exact): the authors' shipped DESeq2 table (All_results.csv, 25,299 genes) filtered at |log2FC|>1 & padj<0.05 reproduces the reported 755 DEGs / 323 up / 432 down EXACTLY; DEG-genesymbol.txt = exactly 755 genes; group.txt = 269 samples (135/134). C2 (within-tol): re-running DESeq2 High-vs-Low from a COMPLETELY INDEPENDENT quantification (recount3 STAR/gencode-v26 raw counts) with the shipped group labels matched 269/269 TCGA-LIHC patients and reproduced the shipped per-gene log2FoldChange at Pearson r=0.932 / Spearman 0.920 across 21,870 genes, recovering 558 of 696 shipped DEGs (Jaccard 0.733). Strong positive evidence the DEG analysis is genuine and reproducible (cross-pipeline, so not bit-identical). NOT attempted: CIBERSORTx signature matrix + cell fractions (Stanford web server, registration-gated; outputs shipped, which enabled the DEG repro); GO/pathway enrichment (PANTHER + Reactome/KOBAS web tools, only plotting code shipped; flagged discrepancy: enrichment input = 651 genes, not 755 DEGs); LASSO Cox prognostic model C-index 0.775/0.787 (stochastic CV, gene-expression model input not cleanly shipped). Documentation gaps: no shipped DESeq2 code (only inputs/outputs); TCGA download script has an undefined-variable bug (matrixl); group.txt High/Low labels swapped; volcano script uses |log2FC|>2 vs the |log2FC|>1 of the 755 headline.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 78assessed: 2026-06-15 ⛓ d19a06e4ba07
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThis study tests whether immune cell subsets newly identified by scRNA-seq (particularly the CD4-GZMA tumor-infiltrating lymphocyte subset) show distinct tumor-infiltration patterns and carry prognostic value in large bulk HCC cohorts when deconvolved with CIBERSORTx.
- ★ CIBERSORTx can combine scRNA-seq-derived signatures with bulk HCC transcriptomes to estimate proportions of 11 TIL subsets and infer cell-type-specific gene expression. method
- ★ TIL subsets are distributed differently between tumor and normal tissue, with CD4-GZMA and CD8-LAYN higher in normal tissue and CD8-LEF1 higher in tumor. finding
- ★ Low proportion of CD4-GZMA cells is significantly associated with poor overall survival in HCC. finding
- ★ An immune risk score model built from CD4-GZMA cell-specific DEGs via LASSO Cox regression predicts HCC patient prognosis and validates across TCGA and ICGC cohorts. resource
- CD4-GZMA cells and their molecular characteristics may yield therapeutic benefit in HCC immunotherapy. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA-seq (Smart-seq2) | HCC patient T cells from peripheral blood, tumor, and adjacent normal tissue (GSE98638) | none | single-cell transcriptome / 11 immune cell subset clusters used to build signature matrix | Illumina HiSeq2500 and HiSeq4000 |
| bulk RNA-seq deconvolution (CIBERSORTx) | TCGA LIHC HCC cohort (269 patients) and ICGC LIRI cohort (229 patients) | none | relative proportions of 11 TIL subsets in tumor vs normal samples | CIBERSORTx online platform (B-mode batch correction, 1000 permutations) |
| microarray/gene expression deconvolution (CIBERSORTx) | GEO HCC datasets GSE76427, GSE64041, GSE36376, GSE14520 | none | immune subset proportions tumor vs normal | — |
| differential gene expression (DESeq2) | TCGA HCC RNA-seq counts (269 patients), CD4-GZMA high vs low groups | CD4-GZMA proportion median split | 755 differentially expressed genes (|log2FC|>1.0, adj.p<0.05) | — |
| CIBERSORTx group-mode cell-type-specific expression imputation | TCGA HCC tumor vs normal classes | none | 384 CD4-GZMA cell-specific DEGs (FDR<0.05, |log2FC|>1.0) | CIBERSORTx group-mode |
| survival analysis (univariable Cox, LASSO Cox, Kaplan-Meier, ROC, nomogram) | TCGA and ICGC HCC cohorts | none | prognostic genes, immune risk score, overall survival prediction | R (glmnet, survminer, survival, rms, survcomp, ROCR, rmda) |
- ▼ Low CD4-GZMA cell proportion associated with poor overall survival (TCGA, validated in two other datasets) log-rank p<0.05
- ▲ High CD4-CTLA4 cell proportion associated with poor overall survival in TCGA log-rank p<0.05
- – CD4-GZMA and CD8-LAYN cells higher in normal tissue; CD8-LEF1 higher in tumor tissue
- – 19 CD4-GZMA cell-specific genes associated with prognosis by univariable Cox regression p<0.05
- – Immune risk score nomogram showed strong predictive accuracy in TCGA cohort C-index 0.775
- – Immune risk score nomogram validated in independent ICGC cohort C-index 0.787
- – Time-dependent ROC for risk score in TCGA at 2, 3, 4, 5 years AUC 0.748, 0.759, 0.752, 0.747
- – Time-dependent ROC for risk score in ICGC at 2, 3, 4, 5 years AUC 0.686, 0.686, 0.748, 0.762
- count 4,070 cells across 11 immune cell subsets (cells retained after filtering for signature matrix)
- count 2,527 genes in signature matrix of 11 cell clusters (CIBERSORTx signature matrix)
- count 755 DEGs (|log2FC|>1.0, adj.p<0.05) (CD4-GZMA high vs low TCGA groups)
- count 384 CD4-GZMA cell-specific DEGs (FDR<0.05, |log2FC|>1.0) (CIBERSORTx group-mode tumor vs normal)
- other C-index 0.775 (nomogram predictive accuracy in TCGA)
- other C-index 0.787 (nomogram predictive accuracy in ICGC)
- fold_change FNDC4 HR 1.3 (95% CI 1.1–1.6), p=0.0041 (univariable Cox, top prognostic gene)
- count 13 genes with non-zero LASSO coefficients in ≥900/1000 evaluations (final immune risk score model genes)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper applied CIBERSORTx to deconvolve bulk RNA-seq profiles from TCGA (n=269), ICGC (n=229), and four GEO cohorts into proportions of 11 scRNA-seq-derived TIL subsets, compared distributions between tumor and normal samples via Wilcoxon tests, and assessed prognostic relevance of each subset using Kaplan-Meier/log-rank analysis. DESeq2 and CIBERSORTx group-mode were used to identify differentially expressed genes in CD4-GZMA-high vs. -low tumors; univariable Cox regression screened 19 prognostic genes, and LASSO Cox regression (10-fold CV, repeated 1,000 times) built a 13-gene immune risk score. Model performance was reported as time-dependent AUC, C-index, calibration curves, and Decision Curve Analysis, with independent validation in the ICGC cohort.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Wilcoxon rank-sum test | Comparison of 11 immune cell subset proportions between tumor and normal samples across 6 datasets (TCGA, ICGC, GSE76427, GSE64041, GSE36376, GSE14520); violin plots in Figure 2 | Varies by dataset: TCGA n=269, ICGC n=229, GSE14520 n=213; exact n for other GEO datasets not stated in text | not stated |
| Log-rank test | Kaplan-Meier overall survival comparisons for each of 6 immune cell subsets (Figure 3) and for high vs. low immune risk score groups (Figures 6A, 7A) in TCGA and ICGC cohorts | TCGA n=269; ICGC n=229; effective n after CIBERSORTx p<0.05 filter not separately stated | not stated |
| Univariable Cox proportional hazards regression (Wald test) | Association of each CD4-GZMA cell-specific DEG with overall survival in TCGA cohort (Table 3) | TCGA n=269 (or CIBERSORTx-filtered subset; exact effective n not stated) | not stated |
| DESeq2 negative binomial GLM with Wald test | Differentially expressed genes between CD4-GZMA high and low proportion groups defined by median fraction in TCGA (Figure 4) | TCGA n=269 (count-level RNA-seq data) | not stated |
| LASSO Cox regression with 10-fold cross-validation (repeated 1,000 times) | Construction of 13-gene immune risk score model in TCGA cohort (Figure 6); genes selected if non-zero coefficient in ≥900/1,000 evaluations | TCGA n=269 (or CIBERSORTx-filtered subset; exact n not restated here) | not stated |
| Time-dependent ROC / AUC (ROCR package) | Discrimination performance of immune risk score at 2, 3, 4, and 5 years in TCGA and ICGC cohorts (Figures 6B, 7B) | TCGA n=269; ICGC n=229 | na |
-
Overall survival groups were defined by dichotomizing CIBERSORTx-estimated cell fraction at the median for Kaplan-Meier analysis↳ Could also: Cox proportional hazards regression treating cell fraction as a continuous predictor — or modeled via restricted cubic splines — could also characterize the association — Continuous modeling retains information across the full range of values and does not require selection of a threshold; it also allows the shape of the dose-response relationship to be visualized and tested
-
Log-rank tests were performed separately for each of the 11 immune cell subsets across multiple datasets without a stated multiplicity correction↳ Could also: A Benjamini-Hochberg FDR or Bonferroni correction applied across the family of survival tests could also be reported alongside the uncorrected p-values — When many parallel log-rank tests share the same outcome data, a familywise or FDR correction quantifies the expected proportion of false discoveries across the full set of comparisons
-
Nineteen DEGs were pre-selected by univariable Cox regression (p<0.05) before being entered into LASSO Cox regression↳ Could also: LASSO Cox regression applied directly to the full set of CD4-GZMA-specific DEGs (FDR<0.05, |log2FC|>1) without prior univariate filtering could also be used — Two-stage filtering (univariate then penalized) can exclude variables that contribute to prognosis only in combination; applying the penalty directly to the DEG set allows the regularization term to perform variable selection in a single stage
-
Samples with a CIBERSORTx deconvolution p-value ≥0.05 were excluded from downstream analyses↳ Could also: Sensitivity analyses retaining all deconvolved samples (or using alternative confidence thresholds) alongside the filtered analysis could also be reported — If excluded samples are systematically different in clinical or expression characteristics, reporting results under multiple thresholds communicates how sensitive conclusions are to the filtering choice
-
Model discrimination was summarized using time-dependent AUC at selected time points (2, 3, 4, and 5 years)↳ Could also: Harrell's concordance index (C-statistic) or Uno's C-statistic could also summarize overall discrimination across the entire follow-up in a single value — A global C-statistic integrates discrimination over all observed event times rather than at pre-specified points, providing a complementary summary particularly when the censoring pattern varies across cohorts
-
Bulk-level DEG analysis (DESeq2 on TCGA counts comparing high vs. low CD4-GZMA fraction groups) and cell-type-specific DEG analysis (CIBERSORTx group-mode) were conducted as parallel, largely separate analyses↳ Could also: A formal overlap analysis reporting concordance of direction and magnitude of log-fold-changes between the two DEG lists could also be presented alongside the Venn diagram — Quantifying the agreement (or divergence) between bulk and cell-type-deconvolved gene expression estimates helps readers assess how much the cell-type-specific deconvolution step adds beyond standard bulk-level differential expression
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
High CD4-CTLA4 TIL proportion is associated with poor overall survival in HCC.other human hcc up 2020×1papers★ This paper is the founder (earliest)
-
Low CD4-GZMA TIL proportion is associated with poor overall survival in HCC (TCGA and independent cohorts).other human hcc down 2020×1papers★ This paper is the founder (earliest)
-
19 CD4-GZMA TIL-specific genes are independently prognostic for overall survival in HCC by univariable Cox regression.other human hcc 2020×1papers★ This paper is the founder (earliest)
-
A CD4-GZMA-specific gene immune risk score nomogram predicts HCC overall survival with C-index 0.775 in the TCGA cohort.other human hcc 2020×1papers★ This paper is the founder (earliest)
-
CD4-GZMA and CD8-LAYN TIL subsets are enriched in adjacent normal liver tissue, while CD8-LEF1 is enriched in HCC tumor tissue.RNA-seq human hcc down 2020×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-33043022
Paper: Li et al. 2020, Front Med 7:451. "Evaluating Distribution and Prognostic Value of New Tumor-Infiltrating Lymphocytes in HCC Based on a scRNA-Seq Study With CIBERSORTx." PMID 33043022 / PMC7527443 / DOI 10.3389/fmed.2020.00451.
Code: https://github.com/szlilixing/Transcriptome-analysis @ commit
1668b28ab1c30cfd550276268d7e4d5b4734ad46 (pushed 2020-06-28). Authors' own R code,
organized as folders each holding a *code.txt script plus the shipped input/output
data file(s) for that figure.
Data: TCGA-LIHC (RNA-seq, n=269 tumor), ICGC LIRI-JP (n=229), GEO GSE76427 / GSE64041 / GSE36376 / GSE14520 (microarray, immune-infiltration validation), GSE98638 (scRNA-seq, 4,070 T cells, signature-matrix source).
Pipeline map (what produces each reported result)
| Result (reported) | Pipeline | Shipped artifact | In scope? |
|---|---|---|---|
| scRNA signature matrix (2,527 genes × 11 clusters) from GSE98638 | CIBERSORTx web "Create Signature Matrix" (B-mode, Smart-seq2) | none (only used downstream) | OUT — CIBERSORTx is a registration-gated web app (cibersortx.stanford.edu), not runnable headlessly; signature matrix not shipped |
| Immune-cell fractions per bulk sample (TCGA/ICGC/GEO) | CIBERSORTx web "Impute Cell Fractions" (1000 perms, B-mode) | Violin plot/LIHC-Fraction.txt (fractions) |
OUT for re-running CIBERSORTx; the fractions themselves are shipped → downstream steps start from them |
| 755 DEGs (323 up / 432 down), |log2FC|>1 & padj<0.05, High vs Low CD4-GZMA group, TCGA-LIHC | DESeq2 on TCGA-LIHC counts, High/Low split | Volcano plot/All_results.csv (full DESeq2 table), Heatmap/ntdeseq_res.txt, Heatmap/group.txt (269 grp labels), Heatmap/DEG-genesymbol.txt |
IN — (a) internal-consistency of shipped table vs reported counts; (b) re-run DESeq2 on TCGA-LIHC counts + shipped group labels |
| Volcano / Heatmap figures | ggplot2 / pheatmap on the DEG table | scripts + data shipped | IN (derived from DEG table; figure-level, lower priority) |
| GO + KEGG/pathway enrichment of the 755 DEGs | clusterProfiler / KOBAS 3.0 | Bubble plot of GO analysis/F3-GO.txt, Bar plot of Pathway analysis/F3-Pathway.txt |
IN (secondary) |
| Univariable Cox (Table 3, 19 genes w/ HR,p) + LASSO Cox 13-gene model, C-index 0.775 (TCGA) / 0.787 (ICGC), time-dependent AUCs | survival/glmnet on clinical+expression | Construction of clinical model code.txt, Kaplan-Meier analysis/clinical-info-KM.txt |
IN (tertiary; LASSO has CV randomness → approximate) |
| GEO GSE76427 download → probe→symbol matrices | GEOquery + illuminaHumanv4.db | Date processing scripts |
partial scope — deterministic but not a headline number; low priority |
| Kaplan-Meier per immune cell type (CD4-GZMA, CD4-CTLA4, …) | survminer on fractions + clinical | KM-survminer code.txt, clinical-info-KM.txt |
IN (secondary) |
Out-of-scope, with reason
- CIBERSORTx signature-matrix construction and cell-fraction imputation — the
central method — runs only on the Stanford CIBERSORTx web server (account/token
gated, no headless CLI without a Docker token + the scRNA signature). Not attempted.
Its outputs (cell fractions) are shipped (
LIHC-Fraction.txt), so downstream pipeline steps can be reproduced starting from the shipped fractions/groups. - scRNA-seq clustering into 11 subsets (from the source scRNA paper, Zheng et al. GSE98638) — upstream, not re-derived here.
Reproduction priority (80/20)
- C1 (cheap): internal consistency — does shipped
All_results.csvfiltered at |log2FC|>1 & padj<0.05 yield 755 DEGs / 323 up / 432 down? Dogroup.txt(135 High/134 Low) andDEG-genesymbol.txt(755) match? - C2 (core): re-run DESeq2 on TCGA-LIHC raw counts + shipped High/Low group
labels; compare log2FC/padj and DEG counts to
All_results.csv. - C3 (secondary): GO/KEGG enrichment on the 755-gene list vs shipped F3-GO / F3-Pathwa
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The core DEG pipeline reproduces cleanly: the reported 755/323/432 DEGs derive exactly from the shipped DESeq2 table, and a completely independent recount3 quantification concords strongly (log2FC Pearson r=0.932, 558/696 shipped DEGs recovered) with 269/269 barcodes matched — strong evidence the DEG analysis is genuine, not fabricated. Deviations are on our/technical side (the 937-vs-755 count gap is purely because recount3 tests ~2x more genes) plus authors'-side documentation issues (no DESeq2 code, 651≠755 enrichment input, |log2FC|>2-vs->1 threshold mismatch, swapped labels). However, the paper's central claims — CIBERSORTx immune-cell distribution and the prognostic C-index 0.775/0.787 — were untestable (registration-gated web server / stochastic LASSO with un-shipped input), so the headline conclusion is confirmed only in limited part, not refuted. Overall yellow: solid, reproducible DEG core with explainable deviations but unverified central methodology.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.