T Cells With Activated STAT4 Drive the High-Risk Rejection State to Renal Allograft Failure After Kidney Transplantation.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL, leaning POSITIVE for the in-scope scRNA-seq arm. DESCRIBED WELL ENOUGH: Methods specify a concrete Scanpy pipeline; CODE is P16-valid (cited PLOGS renamed to ZhangHongbo-Lab/DEAPLOG, a general scanpy single-cell DEG package); DATA is public+complete (GSE145927 7x 10x matrices + GSE109564 DGE, all downloaded and parsed). I re-ran a Methods-faithful pipeline (merge->filter min_genes=500/min_cells=3->[scrublet skipped: skimage absent]->normalize 1e4+log1p->cell-cycle regress->2000 HVG->bbknn->PCA40->UMAP->Leiden res=1.0->canonical marker-panel annotation->STAT4 per cell type) on «our HPC» «job» (COMPLETED, 13m40s). RESULTS vs paper: C1 cell types 10 vs 11 reported (within tolerance; resolution unspecified; only difference is M1/M2 macrophage not split) -> within-tol. C2 identities 10 of 11 recovered (all epithelial PT/LOH_AL/LOH_DL/PC/IC, stromal MyoFB, endothelial Endo, lymphoid T/B, macrophage; M1 not separately resolved) -> partial. C3 the paper's CENTRAL claim STAT4 enriched in T cells reproduces EXACTLY: STAT4 is the #1-ranked cell type for both mean expression (0.225, ~2.7x the 2nd) and fraction positive (13.2%). Grades provisional, for human sign-off. NOT ATTEMPTED (hard ~20%, scope.md): bulk-microarray arm (6 rejection states from ~2,611 unspecified samples, WGCNA 18 modules, IRegulon, Scissor, CIBERSORT, NicheNet, LASSO 7/8-gene survival) — hinges on an un-pinnable accession list and interactive R/Java-GUI tools. Note on path: two earlier run attempts failed for fixable reasons (home-quota-broken conda pkgs cache -> base-python numba bug; missing bioconda channel + dense-load OOM) and the account briefly hit an all-filesystem storage-quota wall (resolved centrally in ~5 min); the successful run reuses the proven pmid-38879692 scanpy env + bbknn via pip and sparse loading. Large output adata_processed.h5ad (2.9 GB, sha256 505ce2c263c0546fae29b4fc766998c93713101ac310d8123dcf39bfb9c700c0) kept on «infra».
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-15 ⛓ de55f2c10cef
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetTraditional histology-based Banff classification of kidney allograft rejection fails to capture molecular heterogeneity, so the study tests whether unsupervised transcriptomic reclassification (UMAP/Leiden) can reveal a distinct high-risk rejection state and its underlying cellular and molecular (T-cell/STAT4) drivers of allograft failure.
- ★ Unsupervised UMAP/Leiden clustering of 2,611 microarray datasets reveals 6 rejection states that diverge from traditional Banff clinical classification finding
- ★ The inflammatory state 1 (Infla1) cluster represents a high-risk (HR) state prone to allograft loss finding
- ★ T cells are specifically recruited and enriched in driving the HR rejection state, more so than other immune cell types finding
- ★ STAT4, regulated upstream by PTPN6, is a core transcription factor in T cells linked to poor allograft function and prognosis mechanism
- ★ Higher STAT4 expression is associated with poorer renal allograft survival in an external validation dataset (GSE21374) finding
- ★ WGCNA identifies the MEblack gene co-expression module as most correlated with the HR state, with hub genes regulated by STAT4 mechanism
- ★ Traditional clinical diagnostic categories do not correspond well to transcriptomic heterogeneity of rejection status finding
- PLOGS, a custom Python package, was developed to identify cell type/cluster-specific gene markers resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Microarray gene expression profiling with UMAP/Leiden reclassification | human kidney allograft biopsies (2,611 samples, GEO) | none (observational across rejection diagnoses) | unsupervised rejection state classification and signature genes | microarray, Scanpy (UMAP/Leiden), bowtie2/bedtools/limma preprocessing |
| single-cell RNA sequencing | kidney biopsies from 3 transplant patients (GSE145927, GSE109564; acute ABMR and acute Mix rejection) | acute ABMR / acute Mix rejection (disease state, no experimental perturbation) | cell type identity and rejection-state association per cell | Scanpy v1.4.4, scrublet, bbknn |
| Scissor cell-bulk association analysis | scRNA-seq cells mapped to bulk microarray rejection states | none | cell subpopulations most correlated with each rejection state (23,082 positively relevant cells) | Scissor |
| CIBERSORT cell-type deconvolution | 2,611 microarray datasets | none | relative proportion of cell types per rejection state | CIBERSORT, R package pheatmap |
| Weighted gene co-expression network analysis (WGCNA) | microarray gene expression data across rejection states | none | gene modules correlated with rejection states; hub gene identification | WGCNA R package, power=10, Cytoscape visualization |
| Transcription factor prediction (IRegulon) | MEblack co-expression module hub genes | none | predicted transcription factors regulating hub genes (STAT4) | IRegulon |
| Upstream ligand-signaling network prediction (NicheNet) | T cells from scRNA-seq data | none | predicted ligand signaling paths driving STAT4 expression | NicheNet |
| Survival analysis and LASSO regression prognostic modeling | external validation microarray dataset GSE21374 (human renal allograft) | none (stratified by STAT4/gene expression level) | allograft survival outcome; diagnostic model of STAT4-regulated genes | R glmnet, R v4.1.0 |
- – PCA showed allograft samples with different clinical diagnoses were mixed/randomly distributed, indicating discordance between clinical diagnosis and transcriptomic heterogeneity
- – Unsupervised clustering of 2,611 samples yielded 6 distinct rejection states (Infla1, Infla2, Prog1, Prog2, STA, Fib), each with specific signature genes and GO functions
- ▲ More than 70% of samples in Infla1 showed apparent rejection phenotypes, the severest rejection status 70%
- ▲ All four high-risk (HR) transcript sets showed the highest expression scores in Infla1
- ▲ Scissor-based mapping of scRNA-seq to bulk states showed M1/M2 macrophages and T cells strikingly aggregated in the HR state, with T cells showing the most significant enrichment 23,082 positively relevant cells
- ▲ WGCNA module MEblack showed the biggest correlation with the HR state, and most of its hub genes were regulated by STAT4
- ▼ Patients with higher STAT4 expression in GSE21374 showed poorer allograft survival
- ▲ STAT4 was strikingly expressed in the HR state and significantly enriched specifically in T cells
- count 2,611 (human microarray datasets from kidney allograft biopsies used for reclassification)
- count 6 (number of newly defined rejection states identified by UMAP/Leiden)
- count 23,082 (positively relevant cells identified by Scissor linking scRNA-seq cells to rejection states)
- count 11 (main cell types identified from scRNA-seq clustering of 3 patient samples)
- count 18 (gene co-expression modules generated by WGCNA)
- other >70% (proportion of Infla1 samples showing apparent rejection phenotypes)
- pvalue p < 0.001 (significance threshold (***) for correlation between gene modules and rejection states (Figure 3A))
- other ratio ≥ 0.5, q-value ≤ 1e-30 (thresholds used by PLOGS get_DEG_single function to define cell/cluster-specific marker genes)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study applied PCA, UMAP, and Leiden unsupervised clustering to 2,611 publicly available bulk microarray datasets from renal transplant biopsies to derive six molecular rejection states, then used scRNA-seq data and the Scissor tool to link single-cell populations to those states. Downstream analyses included WGCNA co-expression network construction with IRegulon transcription-factor prediction, CIBERSORT cell-type deconvolution, NicheNet upstream signaling inference, Kaplan-Meier survival analysis for STAT4 stratification, and LASSO regression to build a prognostic gene model. Results were reported primarily as UMAP embeddings, heatmaps, enrichment scores, and Pearson correlation coefficients, with statistical significance for DEGs conveyed via q-value thresholds and for WGCNA module-trait correlations via asterisk-level p-value annotations.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Leiden graph-based clustering | Reclassification of 2,611 bulk microarray samples into six rejection states; clustering of scRNA-seq data into 11 cell-type populations | 2,611 bulk samples; scRNA-seq cells from 3 patients (23,082 Scissor-positive cells retained) | not stated |
| WGCNA Pearson module-trait correlation (Figures 3A, Supplementary Figure 5C) | Association between 18 co-expression gene modules and the six newly defined rejection states | 2,611 bulk microarray samples | not stated |
| LASSO (L1-penalized) logistic/Cox regression via glmnet | Selection of critical prognostic genes from STAT4-regulated gene set and construction of a diagnostic model | GSE21374 training set (60% split; exact n not stated in provided text) | not stated |
| Kaplan-Meier survival analysis (log-rank test not explicitly named) | Allograft survival stratified by high vs. low STAT4 expression in external dataset GSE21374 (Figure 3D) | — | not stated |
| CIBERSORT linear support-vector-regression deconvolution | Estimation of relative cell-type proportions across six rejection states from 2,611 bulk microarray samples (Supplementary Figure 4C) | 2,611 bulk microarray samples | not stated |
| Scissor phenotype-cell Pearson correlation | Identification of scRNA-seq cells positively correlated with each of the six bulk-defined rejection states (Figure 2A, Supplementary Figure 4A) | scRNA-seq cells from 3 patients; 23,082 positively correlated cells retained | not stated |
| DEG identification via PLOGS get_DEG_single (q-value ≤ 1e-30, ratio ≥ 0.5) | Cluster marker genes for both scRNA-seq cell-type clusters and bulk rejection-state clusters | 2,611 bulk samples; scRNA-seq cells from 3 patients | not stated |
| Gene Ontology enrichment analysis | Functional annotation of rejection-state signature genes (Figure 1C) and STAT4 hub-gene targets (Figure 3C) | — | not stated |
-
WGCNA module-trait Pearson correlations across 18 modules and 6 states (108 simultaneous tests) were reported with *** p < 0.001 threshold annotations and no stated multiple-testing correction↳ Could also: Benjamini-Hochberg FDR correction or Bonferroni adjustment applied across the full 108 module-trait correlation matrix could also be reported — When many module-trait pairs are tested simultaneously, unadjusted p-values increase the expected number of false discoveries; FDR-adjusted q-values are routinely reported alongside WGCNA correlation heatmaps in the literature and help readers calibrate confidence in individual module-trait associations
-
A single random 6:4 train/validation split of GSE21374 was used to develop and evaluate the LASSO prognostic model↳ Could also: Repeated k-fold cross-validation (e.g., 10-fold repeated 10 times) within the dataset, combined with validation in a fully independent external cohort, could also estimate model generalisability — A single random split can yield optimistic or pessimistic performance estimates depending on the particular draw; repeated cross-validation reduces this variance and is a standard benchmark approach for LASSO-based biomarker models in transplant genomics
-
Allograft survival was compared between high and low STAT4 expression groups using a binary split and Kaplan-Meier curves↳ Could also: A Cox proportional-hazards model treating STAT4 expression as a continuous predictor, adjusted for clinical covariates such as donor age, cold ischemia time, and HLA mismatch count, could also be applied — Cox regression avoids the need for an arbitrary dichotomisation threshold, estimates hazard ratios with 95% confidence intervals, and can account for potential confounders, which provides additional context for interpreting the prognostic association
-
Batch effects across 2,611 heterogeneous GEO microarray datasets were corrected with ComBat (bulk) and bbknn (scRNA-seq) before unsupervised clustering↳ Could also: Harmony, fastMNN, or Seurat CCA integration could also be applied for the cross-dataset bulk or single-cell integration; for bulk data, limma's removeBatchEffect is a widely used ComBat alternative — Different batch-correction algorithms make different assumptions about batch structure and balance biological versus technical variance differently; reporting cluster stability across two or more correction methods can help confirm robustness of the six-state partition
-
CIBERSORT with a custom single-cell-derived marker matrix was used to estimate cell-type proportions from bulk microarray data↳ Could also: Because the authors generated their own matched scRNA-seq atlas, methods such as MuSiC, BayesPrism, or EPIC that directly use single-cell reference profiles as input could also be applied — Matched reference-based deconvolution methods leverage study-specific cell states (including the novel T-cell subpopulation identified here) rather than a generic signature matrix, which can improve proportion estimates when reference and bulk data are from the same tissue context
-
Leiden clustering was used to define six rejection states from bulk microarray data, with the number of clusters determined by algorithm resolution rather than a formal stability criterion↳ Could also: Consensus clustering (e.g., NMF or k-means with silhouette width or gap-statistic selection, as implemented in ConsensusClusterPlus) could also be used to select and validate the number of clusters — Stability-based methods provide quantitative metrics (cophenetic correlation, silhouette score, proportion of ambiguous clustering) that help justify the chosen cluster count and assess the reproducibility of the partition across subsamples
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35844542
Paper: Chen et al. 2022, T Cells With Activated STAT4 Drive the High-Risk Rejection State to Renal Allograft Failure After Kidney Transplantation. Front Immunol 13:895762. PMID 35844542 / PMC9283858 / DOI 10.3389/fimmu.2022.895762.
Code: https://github.com/ZhangHongbo-Lab/PLOGS — renamed to DEAPLOG
(https://github.com/ZhangHongbo-Lab/DEAPLOG, repo id 214124277). DEAPLOG is the
authors' general-purpose single-cell DEG / pseudotime Python package
(pip install deaplog, v0.0.5). It is not a paper-specific reproduction repo:
the repo ships only benchmark notebooks (DEAPLOG vs Seurat/monocle on simulated
data) — no kidney-transplant analysis script. The Methods cite the old PLOGS
function get_DEG_single; the shipped deaplog package exposes
get_DEG_uniq / get_DEG_multi instead (renamed). Marker/annotation here uses
canonical cell-type markers (robust, method-independent).
Data (public):
GSE145927(Malone et al. 2020, JASN; PMID 32669324) — scRNA-seq, 5 kidney transplant biopsies, 81,139 cells (10x). Suppl:GSE145927_RAW.tar(per-sample 10x matrices) +GSE145927_MGI.metadata.csv.gz(cell metadata).GSE109564(Wu et al. 2018, JASN; PMID 30093597) — scRNA-seq, single ABMR biopsy, 4,487 cells (Drop-seq). Suppl:GSE109564_Kidney.biopsy.dge.txt.gz.- (bulk:
GSE21374+ 2,611-sample microarray compendium from Suppl. Table 1.)
This is a P16-valid reproduction: apply an existing tool (the lab's own scanpy-based pipeline + DEAPLOG) to the paper's public data, per the described parameters. The code need not be a turnkey repro repo.
IN SCOPE — scRNA-seq pipeline outputs (low-hanging, clearly specified)
The Methods §"Single-Cell Transcriptomic Sequencing Data Preprocessing" specify a concrete Scanpy pipeline on GSE145927 + GSE109564: merge → filter (cells <500 genes; genes in <3 cells) → scrublet doublets → normalize total 1e4 + log1p → regress_out cell cycle → HVG → bbknn batch correction → PCA (40 PCs) → UMAP → Leiden clustering → marker genes (PLOGS get_DEG, ratio≥0.5, q≤1e-30).
| id | claim | reported | paper location |
|---|---|---|---|
| C1 | # main cell types from unsupervised clustering of scRNA-seq | 11 main cell types | Results §"T Cells Are Recruited…", Fig 2A |
| C2 | identity of cell types | T cells, B cells, M1+M2 macrophages, + kidney epithelial/stromal: PT, LOH_AL, LOH_DL, PC, IC, Endo, MyoFB | Fig 2A legend |
| C3 | STAT4 expression is specifically enriched in T cells among cell types | "STAT4 … significantly enriched in T cells" | Results §"STAT4 Is Essential…", Fig 3F |
These are directly regenerable from public 10x data with a standard Scanpy run.
OUT OF SCOPE — the hard ~20% (not attempted; documented why)
| result | pipeline | why dropped |
|---|---|---|
| 6 rejection states from 2,611 microarray samples (Fig 1B) | bowtie2 re-annotation + limma + ComBat + Scanpy reclustering | requires the exact 2,611-sample accession list from Suppl. Table 1 (not text-mined); Leiden resolution unspecified; very heavy multi-GSE assembly |
| WGCNA 18 modules, MEblack↔HR, IRegulon STAT4 hubs (Fig 3A–C) | WGCNA + IRegulon (R/Java GUI) | depends on the microarray reclassification above; IRegulon is interactive Cytoscape plugin |
| Scissor 23,082 positively-relevant cells (Fig 2A right) | Scissor (bulk↔sc integration) | depends on bulk reclassification |
| CIBERSORT cell-ratio per state (Fig S4C) | CIBERSORT | depends on bulk + custom signature matrix |
| NicheNet upstream ligands → STAT4 (Fig 4A–B) | NicheNet (R) | depends on T-cell subset from above; many free choices |
| LASSO 7-gene & 8-gene models + survival (Fig 3D,G; 4D,G; 5C) | glmnet on GSE21374 | random 6:4 split (seed unspecified) → not exactly reproducible; downstream of WGCNA |
Rationale: per BRIEF 80/20 — reproduce the clearly-specified scRNA-seq outputs; the bulk-microarray arm hinges on an un-pinnable 2,611-sample list and a
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is an incomplete-but-honest reproduction: the code (PLOGS→DEAPLOG) and data (GSE145927+GSE109564) resolved and a Methods-faithful Scanpy pipeline was submitted as «our HPC» «job», but the room was finalized early on operator instruction while the job was still running, so C1/C2/C3 have no reproduced values — nothing was fabricated, but nothing was confirmed either. The only substantive gaps are on the input/method side: an unreported Leiden resolution/seed (making the '11 cell types' count non-deterministic) and an un-pinnable bulk-arm sample list. Severity and core-claim status (STAT4 enrichment in T cells, Fig 3F) are therefore undetermined, warranting an overall yellow pending the «infra» output.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.