CCL8 as a promising prognostic factor in diffuse large B-cell lymphoma via M2 macrophage interactions: A bioinformatic analysis of the tumor microenviron
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the central pipeline result 1:1. Reran repo Step01 (ESTIMATE scores via the estimate R package) + Step02 (median-split Kaplan-Meier overall survival) = published Fig 2A, on the primary dataset GSE10846 (GPL570, 420 arrays) on «our HPC». All three KM survival directions and p-values match the published figure within tolerance (stromal p<0.001 vs 0.00025; immune 0.094 vs 0.087; ESTIMATE 0.005 vs 0.0084), confirming the headline claim 'higher stromal score -> favorable prognosis'. Caveat: Fig 2A pools 443 samples (414 GSE10846 + 29 TCGA); we used GSE10846 only (414) -> the small p shifts are consistent with the 29 omitted TCGA cases. Reconstruction gap: the central input merge.txt and the probe->symbol collapse are not shipped/documented; used standard max-mean collapse (ESTIMATE is robust to this). NOT attempted (80/20): Step03 Fig2B-E boxplots; Step04 CIBERSORT (repo omits ICI12.CIBERSORT.R + LM22 ref.txt); Steps 05-24 downstream ICI clustering/gene-cluster/ICI-score/multiGSEA/PPI/COX hub-gene/Venn/hub-gene survival; Figure 6 wet-lab q-PCR + immunofluorescence (out of scope). No fabrication detected.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 81assessed: 2026-06-15 ⛓ 70c45c21003b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a stromal/immune score-based hub gene from the diffuse large B-cell lymphoma (DLBCL) tumor microenvironment serve as a prognostic biomarker? The paper hypothesizes that CCL8, interacting with M2 macrophages and immune checkpoints, is associated with patient survival in DLBCL.
- ★ CCL8 is a promising independent prognostic factor in DLBCL associated with M2 macrophage interactions finding
- ★ Higher stromal score is associated with favorable prognosis in DLBCL finding
- ★ Five immune-related hub genes (C1QB, CCL8, CD3G, CD163, LILRB2) are associated with overall survival and clinical stage in DLBCL finding
- ★ Abundant M2 macrophages were found in the high-CCL8 expression group, and CCL8 is enriched in immune-related processes and secretory granule functions mechanism
- ★ Patients in ICI B and gene B clusters had better outcomes with higher PD-L1 and CTLA4 expression finding
- ★ ESTIMATE algorithm applied to TCGA and GEO transcriptomic data to derive immune/stromal scores and define ICI clusters and hub genes method
- A reproducible analysis pipeline and source code for CCL8-DLBCL is provided as a resource on GitHub resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Bulk transcriptomic profiling (microarray/RNA-seq) with ESTIMATE, CIBERSORT, ICI clustering and Cox regression | 443-449 DLBCL human samples (TCGA n=29, GSE10846 n=420) | none | Immune/stromal/ESTIMATE scores, immune cell infiltration, hub gene expression, overall survival | Affymetrix GPL570; R 3.6.0 (estimate, limma, ConsensusClusterPlus, survival) |
| Quantitative real-time PCR (qPCR) | FFPE human DLBCL tissues (8 DLBCL, 4 normal lymph node) | none | mRNA expression normalized to GAPDH | StepOnePlus (Thermo Fisher); Absolute qPCR SYBR Green Master Mix (Vazyme); High Pure FFPET RNA Kit (Magen) |
| Tissue immunofluorescent staining | Paraffin-embedded human DLBCL patient tissues | none | CCL8 and CD163 protein localization/expression | Anti-CCL8 (Abcam ab155967), anti-CD163 (Proteintech 16646-1-AP), DAPI (Roche) |
| Validation in independent GEO cohorts | Human DLBCL lymph node (GSE136971 n=221, GSE10524 n=40, GSE64555 n=40, GSE114175 n=52) | none | CCL8 prognostic value / survival | GPL570, GPL24975 |
- ▲ Higher stromal score and ESTIMATE score groups showed significantly higher survival probability P<0.01 (stromal), P=0.005 (ESTIMATE)
- – Immune score had limited correlation with survival overall but positively correlated within the first decade P=0.094
- ▼ Lower expression of CD163, CCL8, LILRB2, and C1QB showed significantly higher OS rates CD163 P<0.001, CCL8 P=0.002, LILRB2 P=0.031, C1QB P=0.031
- ▲ CD3G expression negatively correlated with overall survival time P=0.047
- ▲ Four of five hub genes (except CD3G) significantly increased in clinical stage 4 vs stage 1 P=0.002, P=0.028, P=0.012, P=0.0085
- ▲ ICI B group had longer median survival than ICI A group ICI A 7.49 years vs ICI B 17.60 years
- ▼ Patients with low ICI scores exhibited prognosis advantage across datasets GSE10846 P<0.01, TCGA P=0.042, combined P<0.01
- ▼ LDH ratios elevated with lower immune, stromal, and ESTIMATE scores immune P<0.001, stromal P=0.012, ESTIMATE P<0.001
- count 449 DLBCL cases (TCGA n=29, GSE10846 n=420); 443 with survival data (Total samples analyzed)
- pvalue P<0.01 stromal, P=0.005 ESTIMATE (Survival by score groups)
- pvalue CCL8 P=0.002 (CCL8 low expression higher OS in GSE10846)
- other ICI A 7.49 years vs ICI B 17.60 years median survival (Median survival between ICI clusters)
- count 157 genes positively correlated (signature A), 7 genes negatively correlated (signature B) (ICI-related DEGs by logFC>2, P<0.05)
- other HR>1 and P<0.01 for risky OS-related genes (Univariate Cox regression of hub genes)
- count 63 genes at high confidence (0.9) in PPI; top 30 hub genes (STRING/Cytoscape PPI network)
- other survival time 0 to 21.78 years (Survival analysis range)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This bioinformatic study applied the ESTIMATE algorithm to transcriptomic profiles of 443 DLBCL samples (TCGA n=29, GSE10846 n=420) to derive immune and stromal scores, then used consensus clustering (ConsensusClusterPlus) to identify immune cell infiltration (ICI) subgroups and ICI-related gene clusters. Differential expression was assessed via limma, hub genes were selected by intersecting PPI network top nodes with univariate Cox regression results, and survival associations were evaluated using Kaplan–Meier curves with log-rank tests and uni-/multivariate Cox proportional hazards models. CCL8 and four co-identified hub genes were validated in FFPE tissue from 8 DLBCL and 4 normal lymph node samples by qPCR and immunofluorescence.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Unpaired t-test (stated in Methods) / ANOVA (stated in Figure 2 caption) | Comparison of immune, stromal, and ESTIMATE scores across clinical subgroups (LDH ratio, extranodal sites, age, clinical stage) in GSE10846 | 412 (GSE10846 after excluding 8 cases with missing clinical data) | not stated |
| Log-rank test | Kaplan–Meier overall survival curve comparisons: immune/stromal/ESTIMATE score groups; ICI A vs. B subgroups; ICI score groups; five hub gene high vs. low expression groups | 443 (immune/stromal/ESTIMATE KM); 312 (hub gene KM, GSE10846 excluding 108 with missing DEG data) | not stated |
| Limma moderated t-statistic (DEG analysis) | Screening differentially expressed genes between ICI subgroups; logFC > 1 (general DEG screen) or logFC > 2 (ICI signature filter), adjusted P < 0.05 | 312 (samples with complete DEG expression data) | not stated |
| Univariate Cox proportional hazards regression | Association of ICI-related DEG hub genes with overall survival in GSE10846 | 312 (GSE10846, excluding 108 cases with missing DEG data) | not stated |
| Multivariate Cox proportional hazards regression | CCL8 as an independent prognostic factor in GSE10846, with verification in independent GEO cohorts (GSE136971, GSE10524, GSE64555, GSE114175) | null | not stated |
| GSEA permutation test (100 phenotype permutations); FDR < 0.05 or nominal P < 0.05 | Pathway enrichment analysis (KEGG gene sets) between high- and low-ICI score subgroups | null | not stated |
| Hypergeometric enrichment test via clusterProfiler (GO and KEGG); P < 0.05 and q < 1 | Functional annotation of CCL8-related DEGs and ICI hub genes | null | not stated |
-
Hub gene expression groups were defined by splitting at the median, creating binary high/low groups for Kaplan–Meier analysis↳ Could also: A continuous Cox proportional hazards model, or a data-driven optimal cutpoint method (e.g., maxstat or the surv_cutpoint function in survminer), could also have been used — Median dichotomization discards within-group variation and the choice of cutpoint can influence the magnitude of apparent survival differences; continuous or data-driven cutpoint approaches retain distributional information and make the cutpoint selection process more explicit and reproducible
-
Clinical characteristic comparisons used unpaired t-tests (Methods text) or ANOVA (Figure 2 caption) across multiple score types and subgroups without a described multiplicity correction↳ Could also: A single omnibus ANOVA followed by a post-hoc correction (e.g., Tukey HSD) for pairwise contrasts, or a Benjamini-Hochberg adjustment applied across all comparisons, could also have been used — When multiple clinical variables and multiple score types are compared simultaneously, characterizing the overall type I error rate with a family-wise or false-discovery-rate correction is a standard approach that makes the scope of inference explicit
-
Univariate Cox regression was used to pre-screen genes for OS association, and candidates passing a P < 0.01 threshold were carried into further selection steps before multivariate modeling↳ Could also: Penalized regression approaches such as LASSO Cox or elastic net Cox could also have been applied to select among candidate predictors simultaneously in a single modeling step — Univariate pre-filtering followed by multivariate modeling can introduce optimistic bias in coefficient estimates; LASSO and elastic net perform variable selection and shrinkage estimation jointly and are commonly used for high-dimensional survival data
-
Immune and stromal cell content was estimated using the ESTIMATE algorithm, a signature-scoring method based on 141 predefined genes↳ Could also: Other immune deconvolution or scoring approaches such as CIBERSORT (absolute mode), xCell, MCP-counter, or TIMER2 could also have been used to quantify tumor-infiltrating immune cell populations — Different methods make different assumptions (linear mixture models vs. gene-set enrichment) and vary in sensitivity for specific cell types; applying a second orthogonal method is a common approach to assess the robustness of immune infiltration estimates
-
The optimal number of consensus clusters (k) was chosen as k = 2 based on visual inspection of the consensus CDF and delta area plots↳ Could also: Formal cluster-validity indices such as the gap statistic, average silhouette width, or the cophenetic correlation coefficient could also have been used alongside or instead of visual inspection — Quantitative indices provide a more reproducible and less subjective basis for selecting k, and can be reported numerically to allow readers to assess stability of the chosen partition
-
Laboratory validation of CCL8 was performed by qPCR and immunofluorescence on FFPE samples from 8 DLBCL cases and 4 normal lymph node controls↳ Could also: Validation of CCL8 expression in one of the existing independent GEO microarray cohorts (e.g., GSE136971 n=221, GSE10524 n=40) at the transcriptomic level could also have provided an additional layer of independent confirmation — With n = 8 DLBCL and n = 4 controls the laboratory validation step has limited statistical power; a transcriptomic validation in an available independent cohort would complement the bench data and allow estimation of the CCL8 expression difference with greater precision
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
High CCL8 expression associates with significantly worse overall survival in DLBCL patientsmicroarray human dlbcl up 2022×1papers★ This paper is the founder (earliest)
-
CD3G expression negatively correlates with overall survival time in DLBCL patientsmicroarray human dlbcl up 2022×1papers★ This paper is the founder (earliest)
-
Higher ESTIMATE and stromal scores associate with significantly improved overall survival in DLBCLmicroarray human dlbcl up 2022×1papers★ This paper is the founder (earliest)
-
Low immune cell infiltration score associates with improved prognosis in DLBCL across multiple independent cohortsmicroarray human dlbcl down 2022×1papers★ This paper is the founder (earliest)
-
Immune score shows limited overall survival association in DLBCL (P=0.094) but positively correlates with survival within the first decade post-diagnosismicroarray human dlbcl mixed 2022×1papers★ This paper is the founder (earliest)
-
Elevated LDH ratio associates with lower immune, stromal, and ESTIMATE scores in DLBCL, marking immune-cold tumorsmicroarray human dlbcl up 2022×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36072582 (CCL8 as prognostic factor in DLBCL via M2 macrophages)
Front Immunol 2022; DOI 10.3389/fimmu.2022.950213. Repo: github.com/sherrylou92/CCL8-DLBCL
(commit pushed 2022-07-03; default branch main). Bioinformatic analysis of public
expression data — the repo ships 24 R scripts ("Step01"–"Step24"), each mapped to a
published figure panel. NOTE: the repo's figure labels use a draft numbering that differs
from the published numbering (e.g. repo "Fig 1A ESTIMATE survival" = published Figure 2A).
Data
- GSE10846 (GPL570, Affymetrix HG-U133 Plus 2.0, DLBCL lymph node, n=420 used; Lenz/Staudt). Series-matrix processed values via GEOquery (compute node has internet).
- TCGA-DLBC (n=29/443 combined) — also used by authors for some panels.
- The central pipeline input
merge.txt(gene-symbol × sample expression matrix) is NOT shipped in the repo, nor is the probe→symbol collapse documented → minor reconstruction gap.
IN SCOPE (deterministic, pipeline-derived, clearly specified) — attempted
| Repo step | Published fig | Pipeline | Why in scope |
|---|---|---|---|
| Step01 ESTIMATE score | (input to Fig 2) | estimate R pkg: filterCommonGenes + estimateScore on GSE10846 |
Fully deterministic given expression matrix |
| Step02 ESTIMATE survival | Fig 2A | median-split KM + log-rank on stromal/immune/ESTIMATE scores | Deterministic; headline claim "higher stromal score → favorable prognosis" |
Primary verifiable claim: direction + significance of the stromal-score OS association (paper text: "A higher stromal score was associated with favorable prognosis in DLBCL").
OUT OF SCOPE (not attempted) — and why
- Step04 CIBERSORT — repo references
ICI12.CIBERSORT.R+ref.txt(LM22) which are NOT in the repo; perm-based p-values are non-deterministic. (Possible later add.) - Steps 05–18 (ICI consensus clustering, gene clusters, ICI score, multiGSEA, PPI, COX) — multi-stage, several rely on Step04 output and unshipped intermediates; the hard ~20%, skipped.
- Steps 19–21 (Venn, hub-gene survival, gene-stage) — depend on COX hub-gene selection.
- Figure 6 (q-PCR, immunofluorescence TIFs) — wet-lab, out of scope by definition.
- TCGA-only panels — secondary cohort; GSE10846 is the primary, cleaner target.
Reproduction approach
All compute on «our HPC» (SLURM «job»). Build R env (r-base 4.3 + GEOquery + survival)
with conda inside the job; install estimate from R-Forge (not on conda). Download GSE10846,
collapse probes→gene symbols (max-mean per symbol), run ESTIMATE, then median-split survival
using GSE10846's own follow-up time + final status fields. Compare to the paper's qualitative
claim and (where readable) the Fig 2A log-rank p-values.
80/20: reproduce the clean ESTIMATE+survival core; do NOT chase the downstream clustering/COX 20%.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The paper's headline claim — higher stromal score -> favorable OS in DLBCL — reproduces 1:1 on public GSE10846: High-score group better OS (Cox HR 0.560), and all three Fig 2A KM p-values match within tolerance (stromal p<0.001 vs 0.00025; immune 0.094 vs 0.087; ESTIMATE 0.005 vs 0.0084). The only deviations are on our/data side: we used GSE10846 only (n=414) vs the paper's combined 443 (GEO+TCGA) because the exact sample list was not deposited, and the merge.txt probe->symbol collapse was undocumented and had to be reconstructed. Both are minor, explainable input-side choices that do not change magnitude, direction, or significance — no fabrication and the core conclusion holds.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.