Identification of a PRDM1-regulated T cell network to regulate atherosclerotic plaque inflammation.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (1:1, all 5 in-scope pipeline claims exact). Jin/Maas et al., Genome Medicine 2025, PRDM1-regulated T-cell network in atherosclerotic plaque. Described well enough to reproduce: the authors ship ordered R scripts (github.com/jha14/PRDM1_Tcell_atherosclerosis, MIT) + all processed objects on Zenodo (10.5281/zenodo.15852128, every used file md5-verified). NOTE the BRIEF code link github.com/jha14/PRDM1 404s — repo was renamed (same owner); resolved via GitHub user listing. On «our HPC» (R 4.5.3, WGCNA 1.74) I re-ran WGCNA from scratch on the shipped expression matrices and replayed the regulatory-network step: stable plaques -> 25 modules (paper 25), unstable -> 38 modules (paper 38); the 'greenyellow' module independently lands at color position 32 = UP-32 (the T-cell module), and top-100 GENIE3 INTERSECT top-100 ARACNe connectivity gives exactly 32 common TFs (paper: 32, Fig 5a) — the set contains precisely the key regulators the paper names in text (PRDM1, RUNX3, IRF7, EOMES, MYB, IRF4). Gene/sample shapes match (13,641 genes; 16+27 samples). WGCNA was recomputed (not copied from a shipped object), so the counts genuinely regenerate. NO fabrication signal. NOT ATTEMPTED (80/20): scRNA-seq Seurat clustering & AddModuleScore (06; >0.6 GB objects, stochastic, figure-level), drug repositioning (07; needs LINCS L1000 GCTX not on Zenodo), GO semantic-similarity heatmap & Bayesian-network edges (02/04; deprioritised), and all wet-lab (Prdm1-cKO mouse, IHC, flow; out of scope). Three job iterations: 2176121 (curl absent on node -> switched to python urllib), 2176124 (bioconductor-wgcna unsolvable -> switched to conda-forge r-wgcna), 2176128 succeeded in 91 s.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 100assessed: 2026-06-14 ⛓ 679a4760a2f8
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusUsing a network-based approach on human carotid plaque transcriptomes, this study asks which immune gene programs drive the transition from stable (low-risk) to unstable (rupture-prone, high-risk) atherosclerotic plaques, hypothesizing that a T cell regulatory network centered on PRDM1 controls plaque inflammation and destabilization.
- ★ A distinct gene co-expression module with a prominent T cell signature is enriched in unstable plaques and distinguishes high-risk from low-risk lesions. finding
- ★ PRDM1 (along with RUNX3 and IRF7) is a key T cell regulatory factor that is downregulated in plaque T cells of symptomatic versus asymptomatic patients, indicating a protective role. mechanism
- ★ T cell-specific Prdm1 deficiency in Western-type-diet-fed Ldlr knockout mice accelerates atherosclerotic plaque progression, demonstrating a causal protective regulatory role for T cell PRDM1. finding
- ★ In silico drug repurposing identifies EGFR inhibitors as promising therapeutic candidates targeting this PRDM1-regulated pathway. resource
- ★ An integrated WGCNA + Bayesian network + gene regulatory network pipeline applied to bulk plaque expression data can identify causal immune gene programs underlying plaque destabilization. method
- PRDM1 was prioritized as the key target because it lies downstream of IRF7. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Microarray (bulk transcriptomics) | Human carotid endarterectomy plaques (MaasHPS cohort, stable n=16, unstable n=27) | none (stable vs unstable disease state) | Genome-wide gene expression; differential expression, WGCNA, Bayesian network, regulatory network | Illumina HumanRef-8 v2.0 expression BeadChip; Illumina BeadStation 500 |
| Microarray (bulk transcriptomics) | Human carotid plaque tissue, BiKE cohort (atherosclerotic n=127, non-atherosclerotic artery n=10) | none (disease vs control tissue) | Differential gene expression | GSE21545 |
| Single-cell RNA sequencing | Human calcified atherosclerotic plaque (atherosclerotic core n=3, proximal adjacent portion n=3) | none | Single-cell differential gene expression | GSE159677 |
| Single-cell RNA sequencing | Human atherosclerotic plaque immune cells (symptomatic n=2, asymptomatic n=4) | none (symptomatic vs asymptomatic) | T cell subset signature validation; differential expression of PRDM1/RUNX3/IRF7 | GSE224273 |
| Loss-of-function atherosclerosis study (in vivo) | T cell-specific Prdm1-deficient, Western-type-diet-fed Ldlr knockout mice | T cell-specific Prdm1 knockout | Atherosclerotic plaque progression/size | — |
| Histology/morphometry/immunostaining | Human CEA plaque sections | none | Plaque area, necrotic core, CD3+ T cell density, CD68+ macrophages, CD31+ EC/microvessels, collagen, calcification | Leica Q500MC software; H&E, Sirius red, Alizarin red |
| In silico drug repurposing | Plaque transcriptomic network / PRDM1 pathway | none (computational) | Candidate therapeutic compounds (EGFR inhibitors) | — |
- ▲ A T cell-signature co-expression module is prominently enriched in unstable plaques and drives the stable-to-unstable phenotypic transition.
- ▼ RUNX3, IRF7 and especially PRDM1 are significantly downregulated in plaque T cells from symptomatic versus asymptomatic patients.
- ▲ T cell-specific Prdm1 deficiency accelerates plaque progression in WTD-fed Ldlr KO mice.
- – EGFR inhibitors identified as promising drug-repurposing candidates for the PRDM1 pathway.
- – Bayesian network of unstable plaque WGCNA clusters yielded a directed acyclic graph of 15 nodes (clusters) and 28 edges. 15 nodes, 28 edges
- other scale-free fitting index 0.84 (WGCNA scale-free topology fit for stable plaques at soft-thresholding power 6)
- other scale-free fitting index 0.81 (WGCNA scale-free topology fit for unstable plaques at soft-thresholding power 6)
- count 25 and 38 co-expression clusters (clusters after WGCNA on stable (25) and unstable (38) plaques)
- count 13,641 genes (genes retained after pre-processing for downstream analysis)
- count 1021 of 1639 transcription factors (candidate human TFs identified in dataset for regulatory network reconstruction)
- count 43 (16 stable, 27 unstable) (MaasHPS plaque segments used for transcriptional profiling)
- mean 72.84 ± 6.47 years (age (mean ± SD) of symptomatic CEA patients, all male)
- other bootstrap=1000, aggregate threshold 0.85, edge strength >0.7 (Bayesian network hill-climbing parameter settings)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This observational study applied Weighted Gene Co-expression Network Analysis (WGCNA) separately to stable (n=16) and unstable (n=27) human carotid artery plaque microarray samples, followed by Bayesian network inference on cluster eigengenes and dual gene regulatory network reconstruction (GENIE3 and ARACNe). Differential gene expression between plaque phenotypes was assessed with the limma linear modeling framework and Benjamini-Hochberg FDR correction; gene set overrepresentation and enrichment analyses characterized the co-expression clusters biologically. Computational findings were validated in external single-cell RNA-seq datasets and a mouse loss-of-function model (text truncated before mouse and scRNA-seq statistical details).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma linear model (moderated t-statistic) | Differential expression between stable (n=16) and unstable (n=27) plaques in MaasHPS microarray dataset; also applied to BiKE dataset | 43 (MaasHPS); 137 (BiKE) | not stated |
| WGCNA with Pearson's correlation and soft-thresholding power transformation | Co-expression network construction run separately for stable and unstable plaque subcohorts; scale-free fit indices 0.84 and 0.81 at power 6 | 16 (stable), 27 (unstable) | not stated |
| GSEA (gene set enrichment analysis) | Assessment of whether unstable plaque WGCNA clusters were enriched among genes ranked by log2 fold change across all 13,641 pre-processed genes | 13641 genes | not stated |
| Hypergeometric test | Evaluation of gene overlap between WGCNA co-expression clusters (alongside Jaccard Index) | null | not stated |
| Bayesian network inference via hill-climbing algorithm with bootstrap resampling (1000 replicates; aggregate threshold 0.85; edge strength cutoff 0.7) | Causal relationship inference between 15 unstable plaque WGCNA cluster eigengenes (n=27 samples) | 27 | not stated |
| GENIE3 (gradient-boosted tree variable importance) and ARACNe (mutual information) | Gene regulatory network reconstruction using all 43 samples; 1021 candidate transcription factors as regulators, 13,641 genes as targets | 43 | not stated |
-
Twenty-one of twenty-four patients contributed both a stable and an unstable plaque segment, creating a partially paired sample structure that is not explicitly modeled in the differential expression analysis↳ Could also: A paired or mixed-effects limma model with patient ID as a blocking factor (e.g., using duplicateCorrelation or a random-effects term) could also account for within-patient correlation — Modeling within-patient pairing partitions inter-individual variability from the phenotype contrast, which can increase sensitivity for detecting differential expression while better controlling type I error
-
WGCNA was run separately on the stable (n=16) and unstable (n=27) subcohorts, and the resulting networks were compared across phenotypes↳ Could also: WGCNA could also be run on the full merged dataset (n=43) with plaque phenotype as a module-trait correlation, followed by formal module preservation analysis (Zsummary statistic) to identify which modules are stable vs. phenotype-specific — Running on the full dataset increases statistical power for module detection; module preservation statistics provide a principled framework for distinguishing shared from phenotype-specific co-expression structure rather than relying on visual or ad-hoc comparison
-
Gene regulatory network reconstruction employed two data-driven methods (GENIE3 and ARACNe) whose outputs were combined↳ Could also: Knowledge-informed approaches such as DoRothEA or SCENIC could also be applied, incorporating curated TF-target priors alongside expression data — Database-informed methods constrain the regulatory search space using experimentally validated TF-target relationships, which can increase specificity for biologically interpretable edges and allow comparison with purely data-driven results
-
Bayesian network structure learning used a single algorithm (hill-climbing) with bootstrap confidence to select edges↳ Could also: Constraint-based (PC algorithm) or hybrid (MMHC) structure-learning approaches could also be applied and their edge sets compared with the hill-climbing result — Score-based and constraint-based methods make different assumptions about faithfulness and causal sufficiency; edges that are consistently recovered across algorithm families are generally considered more robust
-
Gene set enrichment (GSOA and GSEA) was conducted using Gene Ontology terms, with results organized by biological process, molecular function, and cellular component↳ Could also: Curated pathway databases such as KEGG or Reactome could also be queried alongside GO, and redundancy among enriched terms could be reduced using semantic clustering (e.g., via rrvgo or simplify from clusterProfiler) — GO and pathway databases partition the gene space differently, so triangulating across resources can highlight the most robust biological themes; semantic clustering reduces redundancy in large GO result sets to aid interpretation
-
Continuous morphometric and cell-density features (CD3+ T cell density, macrophage fractions, collagen content, etc.) are described in the methods but the statistical test used to compare these between stable and unstable plaques is not stated in the available text↳ Could also: Depending on distribution, either a two-sample t-test or Mann-Whitney U test (with BH correction across the panel of histological features) would be standard choices; a mixed-effects model could additionally account for the partially paired patient structure — For a panel of correlated histological endpoints measured on the same tissue sections, joint modeling or a correction for the number of comparisons would address family-wise error rate across the feature set
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41039608
Paper: Jin H, Maas SL, et al. "Identification of a PRDM1-regulated T cell network to regulate atherosclerotic plaque inflammation." Genome Medicine 2025. DOI 10.1186/s13073-025-01541-6 · PMCID PMC12490039.
Code: https://github.com/jha14/PRDM1_Tcell_atherosclerosis (MIT, default
branch main, pushed 2025-10-03). NOTE: the BRIEF link
github.com/jha14/PRDM1 404s — the repo was renamed to
PRDM1_Tcell_atherosclerosis (same owner jha14, confirmed via GitHub user
repo listing + code search). Authors' own code (P16 N/A — first-party repo).
Processed data (Zenodo): https://doi.org/10.5281/zenodo.15852128 — 17 files with MD5 checksums (exp_SP.rds, exp_UP.rds, genes.rds, GENIE3.rds, ARACNe.rds, GO_result_{SP,UP}.rds, gsea_wgcna.rds, bn_UP_bootstrap.rds, sc_*.rds, etc). Raw data (GEO): GSE163154 = "MaasHPS" microarray, 43 carotid plaque samples (16 stable / 27 unstable). scRNA cohorts: GSE159677, GSE224273.
The repo ships 8 ordered R scripts (codes/00..07) that load the Zenodo
processed objects and regenerate the figures. README pins versions (WGCNA 1.73,
bnlearn 4.5, GENIE3 1.6.0, Seurat 5.1.0, etc) — though the list is internally
inconsistent (WGCNA 1.73 ⇒ R≥4.x vs clusterProfiler 3.12 ⇒ R 3.6), so it is a
soft guide, not a pinned lockfile.
In scope (pipeline-derived, attempted)
| ID | Result | Pipeline | Inputs (Zenodo) | Tractability |
|---|---|---|---|---|
| C1 | WGCNA → 25 stable (SP-1..25) + 38 unstable (UP-1..38) co-expression clusters | WGCNA (01_wgcna.R): adjacency(power=6), TOMsimilarity, hclust avg, cutreeDynamic(minSize=30,deepSplit=2), mergeCloseModules(MEDissThres=0.25) | exp_SP.rds, exp_UP.rds | HIGH — deterministic, minutes |
| C2 | Top-100 TFs by connectivity to UP-32, GENIE3 ∩ ARACNe = 32 common TFs (Fig 5a) | 05_regulatory_network.R: rowSums of shipped GENIE3/ARACNe matrices over UP-32(greenyellow) columns, top-100 each, intersect | GENIE3.rds, ARACNe.rds, + moduleColors_UP from C1 | HIGH (depends on C1 module colors) |
| C3 | GSEA input = 13,641 pre-processed genes; cohort = 16 stable + 27 unstable | data-shape check | genes.rds, exp_SP/UP.rds | HIGH — trivial |
Out of scope / NOT attempted (with reason)
- scRNA-seq cell-type / T-subset clustering & AddModuleScore (06_scRNAseq.R, Seurat 5.1.0): heavy (sc_GSE159677.rds 500 MB + GSE224273), stochastic (UMAP/clustering), figure-level not a single pinnable number. 80/20 → skip.
- Drug repositioning (07): requires LINCS L1000 GCTX (GSE70138/GSE92742, tens of GB) NOT shipped on Zenodo → env/data unavailable for that step.
- GO semantic-similarity heatmap & Bayesian-network edge set (02/04): reproducible in principle from shipped GO_result_/bn_; deprioritised behind C1–C3 (the headline network-construction numbers). May attempt if cheap.
- Wet-lab (Prdm1-cKO mouse atherosclerosis, IHC, flow): out of scope.
Strategy
One «our HPC» SLURM job (partition std, conda env on «infra»): fetch the small Zenodo processed objects, re-run WGCNA (C1), recompute the TF intersection (C2), report gene/sample counts (C3). Compare to the paper's stated 25/38, 32, 13641.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean 1:1 reproduction on the authors' own public data + code: GSE163154 plus md5-verified Zenodo objects, with the authors' R pipeline replayed on «our HPC». All five in-scope headline construction numbers reproduce exactly — 25 stable / 38 unstable WGCNA modules, the 32 common GENIE3∩ARACNe TFs for the T-cell module UP-32 (recovering PRDM1, RUNX3, IRF7, EOMES, MYB, IRF4), 13,641 genes, and the 16+27 cohort. WGCNA was recomputed from scratch (not copied from a shipped object), strengthening the audit; no fabrication signal. The only caveat is scope (scRNA-seq, drug repositioning, and wet-lab were out of scope/not attempted), but every checkable value is on the authors' favorable side.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.