To Explore the Key Subgroup and Their Immune Microenvironment During the Formation of Coronary Plaque With scRNA-seq.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🔴The deviation was non-trivial in magnitude
- 🔴The central claim did not (fully) hold under reproduction
- 🔴Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
EXECUTED on «our HPC»; outcome MISMATCH. PMID 40454289 is a reanalysis of public GEO GSE196943 (FACS-sorted alphabeta T cells, 12 coronary plaques + 3 paired bloods; 23,858 genes x 26,352 cells). Paper ships NO own code; the registry GitHub link is the third-party tool broadinstitute/inferCNV (P16-valid) -- we reconstructed the standard Seurat v4 pipeline from the Methods text and ran inferCNV on the paper's data. The methods are described well enough to attempt all 5 pinned claims. Results: C1 (QC cell count 24,420) REPRODUCES within ~1% (24,670). But the paper's CENTRAL NK storyline does NOT reproduce: C3 -- only 6.3% NK-marker+ cells and ZERO NK-dominant clusters vs the reported 46% (11,185); recovering ~half NK cells from a T-cell-sorted input is biologically implausible and unobtainable from the shipped data. C4 -- inferCNV (T-cell reference) yields NO high-CNV subgroup (all per-group scores ~0.001-0.003, deviations of a few %); the highest is mast cells, not NK, so there is no basis to 'designate NK cells' by CNV. C5 -- no NK cluster exists to subcluster into 4 groups, and the four 'NK-subgroup markers' are housekeeping/ribosomal genes with GNB2L1 == RACK1 (the same gene labels two 'distinct' subgroups). C2 is only nominally true (5 lineages present but ~91% T cells). Provisional possible-fabrication flag: the 46% NK fraction and the NK subgroup/marker structure are not derivable from the data by a standard pipeline and read as clustering/annotation artifacts. NOT attempted (last-20% downstream, out of scope): CytoTRACE stemness, pseudotime/trajectory, CellChat cell-cell communication, GSEA. All grades are PROVISIONAL for human audit.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 33assessed: 2026-06-16 ⛓ 026c86674476
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18👤 1 human curator(s) · Level L2 2026-06-16
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe cellular heterogeneity underlying coronary plaque formation is not fully understood, and single-cell RNA sequencing can identify key immune cell subgroups (particularly NK cell subgroups) and their microenvironment interactions that contribute to coronary plaque development.
- ★ C1 RACK1+ NK cells are a crucial subgroup for understanding coronary plaque formation, exhibiting the highest cell stemness/differentiation potential and positioned at the start of the pseudotime trajectory finding
- ★ NK cells were identified as one of five major cell types (NK cells, T-cells, B-cells, MCs, myeloids) in coronary plaque tissue, with NK cell proportion notably greater in coronary artery samples than peripheral blood finding
- ★ NK cells cluster into four distinct subgroups (C0 GNB2L1+, C1 RACK1+, C2 HNRNPH1+, C3 ATP5E+) with heterogeneous distribution across donor sources, with C0 and C3 found only in coronary artery samples finding
- ★ Integrated single-cell analysis pipeline (scRNA-seq, CytoTRACE, Monocle, Slingshot, CellChat, SCENIC) can characterize NK cell subgroup differentiation trajectories, cell-cell interactions, and transcription factor regulatory networks in coronary plaque method
- ★ CellChat analysis identified three incoming and three outgoing intercellular signaling patterns among NK cell subgroups and other coronary plaque microenvironment cells mechanism
- C1 RACK1+ NK cells show distinct ligand-receptor interactions with other NK cell subgroups and cell types, and act as a key source/target in signaling patterns (e.g., outgoing pattern 2 with B-cells and myeloids) mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA sequencing (scRNA-seq) | human coronary arteries (12 donors) and matched peripheral blood (3 donors) | none | gene expression per cell, cell type/subgroup identity | Seurat (v4.3.0) |
| doublet removal / quality control | scRNA-seq dataset (GSE196943) | none | filtered high-quality single-cell transcriptomes | DoubletFinder |
| copy number variation inference (inferCNV) | NK cells vs T-cell reference | none | CNV signal per region to differentiate NK cell subgroup identity | inferCNV (Broad Institute) |
| differential gene expression / marker identification | five major coronary plaque cell types and four NK cell subgroups | none | cluster-specific marker genes (logFC > 0.25, expressed in >25% of cells) | Seurat FindAllMarker |
| GO, KEGG, and GSEA enrichment analysis | cell type/subgroup DEGs | none | enriched biological processes/pathways (adjusted p<0.05) | ClusterProfiler (v0.1.1) |
| trajectory/pseudotime analysis (CytoTRACE, Monocle, Slingshot) | NK cell subgroups (11,185 NK cells) | none | stemness score, differentiation trajectory, pseudotime gene expression | CytoTRACE; Monocle (DDRTree); Slingshot |
| cell-cell communication analysis | NK cell subgroups, B-cells, T-cells, MCs, myeloids | none | ligand-receptor interactions, incoming/outgoing signaling patterns | CellChat (v1.6.1) |
| gene regulatory network / transcription factor analysis (SCENIC) | NK cell subgroups | none | transcription factor enrichment and regulon activity | pySCENIC (v0.10.0), Python 3.7 |
- – 24,420 cells retained after QC, classified into 5 cell types: NK cells, T-cells, B-cells, MCs, myeloids
- ▲ NK cells and T-cells were predominant cell types; NK cell proportion notably higher in coronary artery than peripheral blood samples
- – 11,185 NK cells clustered into 4 subgroups: C0 GNB2L1+, C1 RACK1+, C2 HNRNPH1+, C3 ATP5E+
- – C0 and C3 NK cell subgroups found exclusively in coronary artery group, not peripheral blood
- – C2 HNRNPH1+ NK cells showed greatest stemness by violin plot analysis, but C1 RACK1+ NK cells showed the highest stemness by CytoTRACE and were located at the start of the pseudotime trajectory
- ▲ RACK1 gene showed high correlation with CytoTRACE score
- – Slingshot analysis identified one lineage (lineage 1) with differentiation endpoint at the C2 subgroup; C1 subgroup expression rose notably by the end of the pseudotime sequence while C0 and C3 decreased
- – CellChat identified 3 incoming signaling patterns and 3 outgoing signaling patterns involving different combinations of NK subgroups and other cell types
- count 24,420 cells (total cells retained after QC and batch effect removal, across 5 cell types)
- count 11,185 NK cells (NK cells identified via inferCNV, subsequently clustered into 4 subgroups)
- count 15 specimens (12 coronary artery donors + 3 blood donors) (sample source for scRNA-seq dataset GSE196943)
- count 4 NK cell subgroups (C0-C3) (subclustering result of NK cell population)
- other logFC > 0.25, expressed in >25% of cells in cluster (threshold for defining DEGs via FindAllMarker)
- pvalue adjusted p < 0.05 (significance threshold for GO/KEGG enrichment analysis)
- other 30 principal components (number of PCs selected for UMAP dimensionality reduction)
- other top 2000 highly variable genes (genes selected during data standardization/normalization)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This scRNA-seq study analyzed coronary plaque tissues and matched blood from 15 donors (12 coronary artery, 3 blood) using a bioinformatics pipeline built on Seurat for quality control, PCA/UMAP dimensionality reduction, and Wilcoxon-based differential gene expression to identify five cell types and four NK cell subgroups. Pseudotime trajectory analyses (CytoTRACE, Monocle/DDRTree, Slingshot), cell-cell interaction inference (CellChat), and transcription factor regulatory network reconstruction (pySCENIC) were applied to characterize the C1 RACK1+ NK cell subgroup. Results are conveyed primarily through visualizations (UMAP, violin, volcano, heatmap, dot, bar, and scatter plots), with adjusted p < 0.05 as the significance threshold for enrichment analyses and logFC > 0.25 with a minimum cell-fraction filter for DEGs; no classical group-level inferential statistics, effect sizes, or dispersion measures are reported.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Wilcoxon rank-sum test (Seurat FindAllMarkers default), cell-level | Differentially expressed genes across all five major cell types and across four NK cell subgroups | 24,420 total cells (all types); 11,185 NK cells (subgroup analysis) | not stated |
| Hypergeometric / Fisher's exact test (ClusterProfiler GO-BP and KEGG enrichment) | Pathway and biological process enrichment for each cell type and each NK cell subgroup | null | not stated |
| GSEA (permutation-based gene set enrichment) | Assessment of predefined gene sets between two biological conditions | null | not stated |
| CytoTRACE (regression-based transcriptomic complexity / stemness scoring) | Differentiation potential ranking of four NK cell subgroups | 11,185 NK cells | na |
| Monocle DDRTree (graph-based pseudotime trajectory inference) | Cell differentiation trajectory of NK cell subgroups | 11,185 NK cells | na |
| Slingshot (principal-curve lineage inference, pseudotime) | Lineage structure and expression dynamics of NK cell subgroups | 11,185 NK cells | na |
-
DEGs across NK cell subgroups were identified with Seurat's FindAllMarkers, which applies a Wilcoxon rank-sum test treating each cell as an independent observation↳ Could also: Pseudobulk differential expression using DESeq2 or edgeR, aggregating cells per donor before testing — Pseudobulk methods account for within-donor cell correlation; when many cells are drawn from few donors, cell-level tests can underestimate variance and inflate the number of significant findings, while pseudobulk approaches align statistical independence with the biological unit of replication
-
Batch effects across the 15 donors were corrected with Harmony before clustering and UMAP↳ Could also: Seurat RPCA or CCA integration, or scVI (variational autoencoder), or Scanorama — Different integration strategies make different assumptions about batch structure and biological variation; scVI explicitly models count-level distributions and can handle complex multi-donor designs, and benchmarking studies show that the best-performing method varies by dataset characteristics
-
Cell-cell communication was inferred with CellChat using its built-in curated ligand-receptor database↳ Could also: NicheNet (which also models downstream target gene expression) or LIANA (a consensus framework that aggregates CellChat, NicheNet, NATMI, and others) — Different tools use different ligand-receptor databases and scoring schemes; LIANA's consensus ranking across methods reduces dependence on any single database and provides more robust prioritization of interactions
-
Transcription factor regulon activity was reconstructed de novo using pySCENIC based on co-expression modules and DNA motif scoring↳ Could also: DoRothEA (confidence-graded curated regulons) or ChEA3 (ChIP-seq-informed TF enrichment) — Curated regulon databases leverage large compendia of experimental TF-binding evidence and can provide complementary evidence to de novo motif discovery, particularly for TFs with well-characterized binding motifs
-
NK cell identity was confirmed by applying inferCNV with T cells as a reference to distinguish NK cells by copy number variation signal↳ Could also: Canonical marker-gene-based classification (e.g., NCAM1/CD56, KLRD1, FCGR3A) combined with a reference-based automated classifier such as SingleR or scType — Reference-based classifiers provide probabilistic confidence scores for each cell's label and do not require a CNV signal, which may be noisy in non-malignant immune cells; using both approaches together can cross-validate assignments
-
Pseudotime directionality was inferred algorithmically by Monocle and Slingshot based on graph topology and a predefined root↳ Could also: RNA velocity (e.g., scVelo), which estimates differentiation direction from the ratio of unspliced to spliced transcripts within each cell — RNA velocity is data-driven and does not require selecting a root cell a priori; it can corroborate or refine the trajectory direction identified by pseudotime methods and provides an independent line of evidence for developmental ordering
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-40454289
Paper: "To Explore the Key Subgroup and Their Immune Microenvironment During the Formation of Coronary Plaque With scRNA-seq." Cardiol Res Pract 2025. PMID 40454289 / PMC12126265.
Code link in registry: https://github.com/broadinstitute/inferCNV — a third-party tool (broadinstitute inferCNV), NOT the authors' own analysis repo. Per brief P16 this is valid: we reproduce by running the standard scRNA-seq pipeline + inferCNV on the paper's data.
Data: GEO GSE196943 — public processed matrices
(GSE196943_counts_matrix.txt.gz raw counts, GSE196943_norm_matrix.txt.gz normalized).
IMPORTANT: GSE196943 is the ORIGINAL dataset from a different study
("Human coronary plaque T cells are clonally expanded and display cross-reacting
viral/self-antigen specificities"; 12 coronary donors + 3 paired blood, 15 GSM,
FACS-sorted αβ T cells). PMID 40454289 is a reanalysis of this public dataset.
In scope (pipeline-derived, attempted)
| # | Reported result | Pipeline | Reproducible? |
|---|---|---|---|
| C1 | 24,420 cells preserved after QC | Seurat QC (mito<20%, ery<5%, DoubletFinder) | yes — run pipeline, count |
| C2 | Five primary cell types (NK, T, B, MC, myeloid) | Seurat cluster + CellMarker annotation | partial — clustering + canonical markers |
| C3 | 11,185 NK cells extracted | annotation / subset | yes — count NK lineage fraction |
| C4 | inferCNV: T-cells as ref → high-CNV subgroup = NK | inferCNV (named tool) | attempt — central methodological claim |
| C5 | 4 NK subgroups C0–C3 (GNB2L1/RACK1/HNRNPH1/ATP5E) | subcluster NK | de-prioritized (last-20%) |
Out of scope (not attempted)
- Stemness (CytoTRACE), pseudotime/trajectory "state1-3", cell-cell communication (CellChat), GSEA/enrichment — downstream, less precisely specified (last 20%).
- Wet-lab / clinical / external-database (CellMarker DB curation) steps.
Red flags noted up front (for AUDIT, provisional)
- Data is FACS-sorted αβ T cells, yet paper reports 5 lineages incl. 11,185 NK (46% of 24,420) — biologically implausible from a T-cell-sorted input.
- inferCNV (a tumor-vs-normal CNV tool) used to "designate NK cells" as high-CNV relative to T cells — non-canonical, immune cells are not CNV-distinguishable this way.
- NK-subgroup "markers" GNB2L1/RACK1/HNRNPH1/ATP5E are housekeeping/ribosomal genes, not NK markers — and GNB2L1 IS RACK1 (alias of the same gene) labelled as two distinct subgroups (C0 vs C1). Strong artifact signal.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a reanalysis of public GEO GSE196943, so the input data and the exact reported values (24,420 QC cells; 11,185 NK = 46%; inferCNV NK designation; 4 NK subgroups) are directly comparable. Only the QC cell count (C1) reproduces (24,670, +1.0%); the entire NK-centric core fails on the same data — 6.3% NK-marker+ cells vs 46%, no high-CNV NK subgroup (highest is mast at 0.0029), and 'subgroup markers' that are housekeeping genes with GNB2L1 being the deprecated alias of RACK1. The deviation sits in the authors' annotation/computation logic, is severe (~7×, reversed CNV ranking), and is not derivable from the shared data — the NK fraction and subgroup structure read as clustering/annotation artifacts, warranting a fabrication-suspect flag on q5/q7/q8.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.