Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

T Cells With Activated STAT4 Drive the High-Risk Rejection State to Renal Allograft Failure After Kidney Transplantation.

Front Immunol · 2022
L1 78/100 PQI 87
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 51% of all assessed papers rank 533 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL, leaning POSITIVE for the in-scope scRNA-seq arm. DESCRIBED WELL ENOUGH: Methods specify a concrete Scanpy pipeline; CODE is P16-valid (cited PLOGS renamed to ZhangHongbo-Lab/DEAPLOG, a general scanpy single-cell DEG package); DATA is public+complete (GSE145927 7x 10x matrices + GSE109564 DGE, all downloaded and parsed). I re-ran a Methods-faithful pipeline (merge->filter min_genes=500/min_cells=3->[scrublet skipped: skimage absent]->normalize 1e4+log1p->cell-cycle regress->2000 HVG->bbknn->PCA40->UMAP->Leiden res=1.0->canonical marker-panel annotation->STAT4 per cell type) on «our HPC» «job» (COMPLETED, 13m40s). RESULTS vs paper: C1 cell types 10 vs 11 reported (within tolerance; resolution unspecified; only difference is M1/M2 macrophage not split) -> within-tol. C2 identities 10 of 11 recovered (all epithelial PT/LOH_AL/LOH_DL/PC/IC, stromal MyoFB, endothelial Endo, lymphoid T/B, macrophage; M1 not separately resolved) -> partial. C3 the paper's CENTRAL claim STAT4 enriched in T cells reproduces EXACTLY: STAT4 is the #1-ranked cell type for both mean expression (0.225, ~2.7x the 2nd) and fraction positive (13.2%). Grades provisional, for human sign-off. NOT ATTEMPTED (hard ~20%, scope.md): bulk-microarray arm (6 rejection states from ~2,611 unspecified samples, WGCNA 18 modules, IRegulon, Scissor, CIBERSORT, NicheNet, LASSO 7/8-gene survival) — hinges on an un-pinnable accession list and interactive R/Java-GUI tools. Note on path: two earlier run attempts failed for fixable reasons (home-quota-broken conda pkgs cache -> base-python numba bug; missing bioconda channel + dense-load OOM) and the account briefly hit an all-filesystem storage-quota wall (resolved centrally in ~5 min); the successful run reuses the proven pmid-38879692 scanpy env + bbknn via pip and sparse loading. Large output adata_processed.h5ad (2.9 GB, sha256 505ce2c263c0546fae29b4fc766998c93713101ac310d8123dcf39bfb9c700c0) kept on «infra».

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-15 ⛓ de55f2c10cef
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Traditional histology-based Banff classification of kidney allograft rejection fails to capture molecular heterogeneity, so the study tests whether unsupervised transcriptomic reclassification (UMAP/Leiden) can reveal a distinct high-risk rejection state and its underlying cellular and molecular (T-cell/STAT4) drivers of allograft failure.

Core claims
  • Unsupervised UMAP/Leiden clustering of 2,611 microarray datasets reveals 6 rejection states that diverge from traditional Banff clinical classification finding
  • The inflammatory state 1 (Infla1) cluster represents a high-risk (HR) state prone to allograft loss finding
  • T cells are specifically recruited and enriched in driving the HR rejection state, more so than other immune cell types finding
  • STAT4, regulated upstream by PTPN6, is a core transcription factor in T cells linked to poor allograft function and prognosis mechanism
  • Higher STAT4 expression is associated with poorer renal allograft survival in an external validation dataset (GSE21374) finding
  • WGCNA identifies the MEblack gene co-expression module as most correlated with the HR state, with hub genes regulated by STAT4 mechanism
  • Traditional clinical diagnostic categories do not correspond well to transcriptomic heterogeneity of rejection status finding
  • PLOGS, a custom Python package, was developed to identify cell type/cluster-specific gene markers resource
Experimental setups
Assay System Perturbation Readout Platform
Microarray gene expression profiling with UMAP/Leiden reclassification human kidney allograft biopsies (2,611 samples, GEO) none (observational across rejection diagnoses) unsupervised rejection state classification and signature genes microarray, Scanpy (UMAP/Leiden), bowtie2/bedtools/limma preprocessing
single-cell RNA sequencing kidney biopsies from 3 transplant patients (GSE145927, GSE109564; acute ABMR and acute Mix rejection) acute ABMR / acute Mix rejection (disease state, no experimental perturbation) cell type identity and rejection-state association per cell Scanpy v1.4.4, scrublet, bbknn
Scissor cell-bulk association analysis scRNA-seq cells mapped to bulk microarray rejection states none cell subpopulations most correlated with each rejection state (23,082 positively relevant cells) Scissor
CIBERSORT cell-type deconvolution 2,611 microarray datasets none relative proportion of cell types per rejection state CIBERSORT, R package pheatmap
Weighted gene co-expression network analysis (WGCNA) microarray gene expression data across rejection states none gene modules correlated with rejection states; hub gene identification WGCNA R package, power=10, Cytoscape visualization
Transcription factor prediction (IRegulon) MEblack co-expression module hub genes none predicted transcription factors regulating hub genes (STAT4) IRegulon
Upstream ligand-signaling network prediction (NicheNet) T cells from scRNA-seq data none predicted ligand signaling paths driving STAT4 expression NicheNet
Survival analysis and LASSO regression prognostic modeling external validation microarray dataset GSE21374 (human renal allograft) none (stratified by STAT4/gene expression level) allograft survival outcome; diagnostic model of STAT4-regulated genes R glmnet, R v4.1.0
Key results
  • PCA showed allograft samples with different clinical diagnoses were mixed/randomly distributed, indicating discordance between clinical diagnosis and transcriptomic heterogeneity
  • Unsupervised clustering of 2,611 samples yielded 6 distinct rejection states (Infla1, Infla2, Prog1, Prog2, STA, Fib), each with specific signature genes and GO functions
  • More than 70% of samples in Infla1 showed apparent rejection phenotypes, the severest rejection status 70%
  • All four high-risk (HR) transcript sets showed the highest expression scores in Infla1
  • Scissor-based mapping of scRNA-seq to bulk states showed M1/M2 macrophages and T cells strikingly aggregated in the HR state, with T cells showing the most significant enrichment 23,082 positively relevant cells
  • WGCNA module MEblack showed the biggest correlation with the HR state, and most of its hub genes were regulated by STAT4
  • Patients with higher STAT4 expression in GSE21374 showed poorer allograft survival
  • STAT4 was strikingly expressed in the HR state and significantly enriched specifically in T cells
Key statistics
  • count 2,611 (human microarray datasets from kidney allograft biopsies used for reclassification)
  • count 6 (number of newly defined rejection states identified by UMAP/Leiden)
  • count 23,082 (positively relevant cells identified by Scissor linking scRNA-seq cells to rejection states)
  • count 11 (main cell types identified from scRNA-seq clustering of 3 patient samples)
  • count 18 (gene co-expression modules generated by WGCNA)
  • other >70% (proportion of Infla1 samples showing apparent rejection phenotypes)
  • pvalue p < 0.001 (significance threshold (***) for correlation between gene modules and rejection states (Figure 3A))
  • other ratio ≥ 0.5, q-value ≤ 1e-30 (thresholds used by PLOGS get_DEG_single function to define cell/cluster-specific marker genes)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study applied PCA, UMAP, and Leiden unsupervised clustering to 2,611 publicly available bulk microarray datasets from renal transplant biopsies to derive six molecular rejection states, then used scRNA-seq data and the Scissor tool to link single-cell populations to those states. Downstream analyses included WGCNA co-expression network construction with IRegulon transcription-factor prediction, CIBERSORT cell-type deconvolution, NicheNet upstream signaling inference, Kaplan-Meier survival analysis for STAT4 stratification, and LASSO regression to build a prognostic gene model. Results were reported primarily as UMAP embeddings, heatmaps, enrichment scores, and Pearson correlation coefficients, with statistical significance for DEGs conveyed via q-value thresholds and for WGCNA module-trait correlations via asterisk-level p-value annotations.

Replicationunclear Sample size2,611 publicly available bulk microarray samples aggregated from multiple GEO series; scRNA-seq from 3 patients (GSE145927, GSE109564); GSE21374 split 6:4 for LASSO training/validation; no prospective power calculation stated GroupsSix unsupervised molecular rejection states (Infla1/HR, Infla2, Prog1, Prog2, STA, Fib); T cells vs. other immune cell types in scRNA-seq; high vs. low STAT4 expression for survival Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionFDR q-value for DEGs (exact method not named; q ≤ 1e-30 threshold); no adjustment stated for WGCNA module-trait correlations (18 modules × 6 states = 108 tests) or GO enrichment
Statistical tests used
Test Applied to n Assumptions
Leiden graph-based clustering Reclassification of 2,611 bulk microarray samples into six rejection states; clustering of scRNA-seq data into 11 cell-type populations 2,611 bulk samples; scRNA-seq cells from 3 patients (23,082 Scissor-positive cells retained) not stated
WGCNA Pearson module-trait correlation (Figures 3A, Supplementary Figure 5C) Association between 18 co-expression gene modules and the six newly defined rejection states 2,611 bulk microarray samples not stated
LASSO (L1-penalized) logistic/Cox regression via glmnet Selection of critical prognostic genes from STAT4-regulated gene set and construction of a diagnostic model GSE21374 training set (60% split; exact n not stated in provided text) not stated
Kaplan-Meier survival analysis (log-rank test not explicitly named) Allograft survival stratified by high vs. low STAT4 expression in external dataset GSE21374 (Figure 3D) not stated
CIBERSORT linear support-vector-regression deconvolution Estimation of relative cell-type proportions across six rejection states from 2,611 bulk microarray samples (Supplementary Figure 4C) 2,611 bulk microarray samples not stated
Scissor phenotype-cell Pearson correlation Identification of scRNA-seq cells positively correlated with each of the six bulk-defined rejection states (Figure 2A, Supplementary Figure 4A) scRNA-seq cells from 3 patients; 23,082 positively correlated cells retained not stated
DEG identification via PLOGS get_DEG_single (q-value ≤ 1e-30, ratio ≥ 0.5) Cluster marker genes for both scRNA-seq cell-type clusters and bulk rejection-state clusters 2,611 bulk samples; scRNA-seq cells from 3 patients not stated
Gene Ontology enrichment analysis Functional annotation of rejection-state signature genes (Figure 1C) and STAT4 hub-gene targets (Figure 3C) not stated
Approaches that could also have been used
  • WGCNA module-trait Pearson correlations across 18 modules and 6 states (108 simultaneous tests) were reported with *** p < 0.001 threshold annotations and no stated multiple-testing correction
    Could also: Benjamini-Hochberg FDR correction or Bonferroni adjustment applied across the full 108 module-trait correlation matrix could also be reported — When many module-trait pairs are tested simultaneously, unadjusted p-values increase the expected number of false discoveries; FDR-adjusted q-values are routinely reported alongside WGCNA correlation heatmaps in the literature and help readers calibrate confidence in individual module-trait associations
  • A single random 6:4 train/validation split of GSE21374 was used to develop and evaluate the LASSO prognostic model
    Could also: Repeated k-fold cross-validation (e.g., 10-fold repeated 10 times) within the dataset, combined with validation in a fully independent external cohort, could also estimate model generalisability — A single random split can yield optimistic or pessimistic performance estimates depending on the particular draw; repeated cross-validation reduces this variance and is a standard benchmark approach for LASSO-based biomarker models in transplant genomics
  • Allograft survival was compared between high and low STAT4 expression groups using a binary split and Kaplan-Meier curves
    Could also: A Cox proportional-hazards model treating STAT4 expression as a continuous predictor, adjusted for clinical covariates such as donor age, cold ischemia time, and HLA mismatch count, could also be applied — Cox regression avoids the need for an arbitrary dichotomisation threshold, estimates hazard ratios with 95% confidence intervals, and can account for potential confounders, which provides additional context for interpreting the prognostic association
  • Batch effects across 2,611 heterogeneous GEO microarray datasets were corrected with ComBat (bulk) and bbknn (scRNA-seq) before unsupervised clustering
    Could also: Harmony, fastMNN, or Seurat CCA integration could also be applied for the cross-dataset bulk or single-cell integration; for bulk data, limma's removeBatchEffect is a widely used ComBat alternative — Different batch-correction algorithms make different assumptions about batch structure and balance biological versus technical variance differently; reporting cluster stability across two or more correction methods can help confirm robustness of the six-state partition
  • CIBERSORT with a custom single-cell-derived marker matrix was used to estimate cell-type proportions from bulk microarray data
    Could also: Because the authors generated their own matched scRNA-seq atlas, methods such as MuSiC, BayesPrism, or EPIC that directly use single-cell reference profiles as input could also be applied — Matched reference-based deconvolution methods leverage study-specific cell states (including the novel T-cell subpopulation identified here) rather than a generic signature matrix, which can improve proportion estimates when reference and bulk data are from the same tissue context
  • Leiden clustering was used to define six rejection states from bulk microarray data, with the number of clusters determined by algorithm resolution rather than a formal stability criterion
    Could also: Consensus clustering (e.g., NMF or k-means with silhouette width or gap-statistic selection, as implemented in ConsensusClusterPlus) could also be used to select and validate the number of clusters — Stability-based methods provide quantitative metrics (cophenetic correlation, silhouette score, proportion of ambiguous clustering) that help justify the chosen cluster count and assess the reproducibility of the partition across subsamples
Software: Python/Scanpy 1.4.4 (scRNA-seq preprocessing); 1.6.0 (bulk reclassification) · Python/scrublet 0.2 · Python/bbknn 1.2.0 · R/limma · R/sva (ComBat) · R/WGCNA · R/glmnet R 4.1.0 stated · R/pheatmap · CIBERSORT · Scissor · NicheNet · IRegulon · Cytoscape · Python/PLOGS (custom package) · bowtie2 · bedtools

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
6
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE109564 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE145927 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35844542

Paper: Chen et al. 2022, T Cells With Activated STAT4 Drive the High-Risk Rejection State to Renal Allograft Failure After Kidney Transplantation. Front Immunol 13:895762. PMID 35844542 / PMC9283858 / DOI 10.3389/fimmu.2022.895762.

Code: https://github.com/ZhangHongbo-Lab/PLOGSrenamed to DEAPLOG (https://github.com/ZhangHongbo-Lab/DEAPLOG, repo id 214124277). DEAPLOG is the authors' general-purpose single-cell DEG / pseudotime Python package (pip install deaplog, v0.0.5). It is not a paper-specific reproduction repo: the repo ships only benchmark notebooks (DEAPLOG vs Seurat/monocle on simulated data) — no kidney-transplant analysis script. The Methods cite the old PLOGS function get_DEG_single; the shipped deaplog package exposes get_DEG_uniq / get_DEG_multi instead (renamed). Marker/annotation here uses canonical cell-type markers (robust, method-independent).

Data (public):

  • GSE145927 (Malone et al. 2020, JASN; PMID 32669324) — scRNA-seq, 5 kidney transplant biopsies, 81,139 cells (10x). Suppl: GSE145927_RAW.tar (per-sample 10x matrices) + GSE145927_MGI.metadata.csv.gz (cell metadata).
  • GSE109564 (Wu et al. 2018, JASN; PMID 30093597) — scRNA-seq, single ABMR biopsy, 4,487 cells (Drop-seq). Suppl: GSE109564_Kidney.biopsy.dge.txt.gz.
  • (bulk: GSE21374 + 2,611-sample microarray compendium from Suppl. Table 1.)

This is a P16-valid reproduction: apply an existing tool (the lab's own scanpy-based pipeline + DEAPLOG) to the paper's public data, per the described parameters. The code need not be a turnkey repro repo.


IN SCOPE — scRNA-seq pipeline outputs (low-hanging, clearly specified)

The Methods §"Single-Cell Transcriptomic Sequencing Data Preprocessing" specify a concrete Scanpy pipeline on GSE145927 + GSE109564: merge → filter (cells <500 genes; genes in <3 cells) → scrublet doublets → normalize total 1e4 + log1p → regress_out cell cycle → HVG → bbknn batch correction → PCA (40 PCs) → UMAP → Leiden clustering → marker genes (PLOGS get_DEG, ratio≥0.5, q≤1e-30).

id claim reported paper location
C1 # main cell types from unsupervised clustering of scRNA-seq 11 main cell types Results §"T Cells Are Recruited…", Fig 2A
C2 identity of cell types T cells, B cells, M1+M2 macrophages, + kidney epithelial/stromal: PT, LOH_AL, LOH_DL, PC, IC, Endo, MyoFB Fig 2A legend
C3 STAT4 expression is specifically enriched in T cells among cell types "STAT4 … significantly enriched in T cells" Results §"STAT4 Is Essential…", Fig 3F

These are directly regenerable from public 10x data with a standard Scanpy run.

OUT OF SCOPE — the hard ~20% (not attempted; documented why)

result pipeline why dropped
6 rejection states from 2,611 microarray samples (Fig 1B) bowtie2 re-annotation + limma + ComBat + Scanpy reclustering requires the exact 2,611-sample accession list from Suppl. Table 1 (not text-mined); Leiden resolution unspecified; very heavy multi-GSE assembly
WGCNA 18 modules, MEblack↔HR, IRegulon STAT4 hubs (Fig 3A–C) WGCNA + IRegulon (R/Java GUI) depends on the microarray reclassification above; IRegulon is interactive Cytoscape plugin
Scissor 23,082 positively-relevant cells (Fig 2A right) Scissor (bulk↔sc integration) depends on bulk reclassification
CIBERSORT cell-ratio per state (Fig S4C) CIBERSORT depends on bulk + custom signature matrix
NicheNet upstream ligands → STAT4 (Fig 4A–B) NicheNet (R) depends on T-cell subset from above; many free choices
LASSO 7-gene & 8-gene models + survival (Fig 3D,G; 4D,G; 5C) glmnet on GSE21374 random 6:4 split (seed unspecified) → not exactly reproducible; downstream of WGCNA

Rationale: per BRIEF 80/20 — reproduce the clearly-specified scRNA-seq outputs; the bulk-microarray arm hinges on an un-pinnable 2,611-sample list and a

Figures / tables: Fig 2AFig 3F
C1
Reported
11 main cell types (scRNA-seq unsupervised clustering, Fig 2A)
Reproduced
10 cell types (23 Leiden clusters @ res 1.0 -> 10 lineages)
within tolerance
C2
Reported
cell types: T, B, M1+M2 macrophage, PT, LOH_AL, LOH_DL, PC, IC, Endo, MyoFB (11)
Reproduced
T, B, M2-macrophage, PT, LOH_AL, LOH_DL, PC, IC, Endo, MyoFB (10/11 identities)
partial
C3
Reported
STAT4 significantly enriched in T cells (Fig 3F)
Reproduced
STAT4 mean log-norm expr HIGHEST in T cells (0.225, rank #1/10, ~2.7x next; 13.2% STAT4+ vs 5.0% in next type)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

This is an incomplete-but-honest reproduction: the code (PLOGS→DEAPLOG) and data (GSE145927+GSE109564) resolved and a Methods-faithful Scanpy pipeline was submitted as «our HPC» «job», but the room was finalized early on operator instruction while the job was still running, so C1/C2/C3 have no reproduced values — nothing was fabricated, but nothing was confirmed either. The only substantive gaps are on the input/method side: an unreported Leiden resolution/seed (making the '11 cell types' count non-deterministic) and an un-pinnable bulk-arm sample list. Severity and core-claim status (STAT4 enrichment in T cells, Fig 3F) are therefore undetermined, warranting an overall yellow pending the «infra» output.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

344.1 k
tokens (I/O) · 25.7 M incl. cache
91 min
runtime · 0.29 CPU-h
18.4 GB
peak RAM
3 (1 failed)
HPC jobs
hummel
machine