COVID-19 vaccination atlas using an integrative systems vaccinology approach.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce 1:1. Authors' own R code (CC0, commit 45f14be) ships the integer count matrix + sample metadata + DESeq2 outputs for 5 integrated GEO RNA-seq studies. I re-ran the central per-study DGE step for GSE206023 exactly as coded (DESeqDataSetFromMatrix design=~age+sex+condition; keep rowSums>=10; DESeq(); results() per coefficient) on «our HPC» (R 4.3.3 / DESeq2 1.42.0). Against the canonical shipped table gse206023_DEGs_wald.csv (the one the pipeline reads back), 98.3% of 106,866 gene x contrast log2FoldChange values match to <0.01 and padj correlate at r=0.994 -> the core DGE pipeline reproduces near-exactly. Two honest gaps: (a) the two smallest-sample D28 contrasts yield 2-3x more DEGs in my raw run, consistent with the paper's documented dynamic FDR filter for sub-median-n comparisons (a downstream step I did not apply); (b) FLAG: the paper's headline 562 samples / 245 participants is not exactly derivable from the shipped metadata, which has 554 sample rows / 225 unique participant_ids (per-study counts reconcile internally to 554) - a reviewer should adjudicate, not asserted as fabrication. NOT attempted (the hard 20%, downstream of DGE): full multi-study PCA (Fig3a 85.4% variance), Table 3 top-20 PC genes, GSEA bubble plots, xCell cell deconvolution (Fig4), the ML marker-gene selection (LHFPL2/FYN/TYMP/JUP/SERINC5/KIT/GRAMD1C), and re-deriving counts from raw GEO FASTQ. The reproduction used the authors' shipped count matrix so the DGE input matches theirs exactly.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 69assessed: 2026-06-14 ⛓ 1e2fd1534dc3
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether an integrative systems vaccinology approach applied to RNAseq data from multiple cohorts can reveal shared and unique immune signatures across different COVID-19 vaccine types, technologies, and regimens (including infection status), and how these compare to vaccines against other pathogens.
- ★ mRNA vaccines induce transient but strong immune responses after booster doses finding
- ★ Viral vector vaccines show sustained immune activation with infection, partially resembling the infection immune profile finding
- ★ Heterologous vaccination regimens demonstrate enhanced immune diversification compared to homologous regimens finding
- ★ Immune-related DEGs comprise less than 10% of total DEGs but are sufficient to establish condition-specific relationships finding
- ★ PCA of immune DEGs identifies three distinct clusters of immune responses corresponding to infection, vaccination in healthy individuals, and early infection/first-dose states finding
- ★ A COVID-19 vaccination atlas integrating 562 samples from 245 participants across five vaccines and regimens was constructed and compared with data from 13 vaccines against other diseases resource
- Shared immune markers (e.g., MARCKS, NFKBIA, IL1B, CCL3) are enriched across both infection and vaccination conditions, not exclusively in one finding
- Sex was manually imputed for GSE206023 using expression of X- and Y-linked genes (DDX3Y, KDM5D, RPS4Y1, UTY, XIST) method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq / differential gene expression analysis | PBMCs from human participants (245 participants, 5 cohorts) | COVID-19 vaccination (IN, SU, RNA, VV types; homologous/heterologous regimens; 1-3 doses) and/or SARS-CoV-2 infection | differentially expressed genes (DEGs), classified as up/down/neutral (|L2FC|>1) | — |
| Gene Set Enrichment Analysis (GSEA) | PBMCs, same cohorts | vaccination/infection conditions | enrichment of immune-related gene sets (ImmuneGO, BTM) and GO cellular/metabolic processes | — |
| Principal component analysis (PCA) with K-means clustering | immune DEGs from PBMC samples | vaccination/infection conditions | variance explained by PC1/PC2, cluster assignment | — |
| Spearman's correlation analysis | immune and non-immune DEGs across conditions | vaccination/infection conditions | condition-specific correlation clustering (FDR <0.10) | — |
| Sex imputation from gene expression (TPM) | PBMC samples (GSE206023) | none (metadata reconstruction) | expression of DDX3Y, KDM5D, RPS4Y1, UTY (Y-linked) and XIST (X-linked) | — |
| Immune cell population dynamics analysis | PBMC samples across cohorts | vaccination/infection conditions | cell population composition changes | — |
- ▲ Vaccinated individuals (e.g., two BNT doses) exhibited higher total DEGs post-infection compared to unvaccinated controls, driven by over 200 upregulated genes (BNT-I D0-5 vs. I D0-4) >200 genes
- – PC1 and PC2 explained 72.8% and 12.6% of variance respectively, together capturing 85.4% of total variance 72.8% + 12.6% = 85.4%
- – Immune-related genes reduced the analysis variable set from ~16,000 genes to ~1800 immune genes ~16,000 to ~1800
- – Reinfected patients displayed fewer upregulated DEGs despite similar downregulation trends compared to non-reinfected
- – DEG counts in hospitalized cohorts peaked at day 10 post-admission in partially vaccinated (BNT-I D10) patients but declined thereafter
- ▼ Immune-related genes constituted less than 10% of total DEGs and declined over time in vaccinated and infected individuals <10%
- ▲ Shared immune markers MARCKS, NFKBIA, IL1B, and CCL3 were enriched across both infection and vaccination conditions
- count 562 samples (Total samples analyzed across the atlas)
- count 245 participants (Total participants across five cohorts/studies)
- other PC1=72.8%, PC2=12.6% variance explained (PCA of immune DEGs across all conditions)
- mean median age 35 years (IQR: 17-83) (Overall study population age)
- count >200 upregulated genes (BNT-I D0-5 vs. I D0-4 comparison)
- pvalue FDR <0.10 (Threshold used for Spearman's correlation analysis of DEGs)
- fold_change vaccine efficacies: BNT 95.0%, MO 94.1%, ChAd 90% (Clinical trial efficacy of homologous vaccination regimens)
- fold_change BBIBP 78.1%, ZF2001 75.7% efficacy against symptomatic infection (Efficacy of classical inactivated/subunit vaccines)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study performed an integrative reanalysis of bulk RNAseq data from five publicly available GEO cohorts (562 samples, 245 participants) covering five COVID-19 vaccine types in homologous and heterologous regimens as well as infected groups. Differential gene expression analysis was applied stratified by study and condition, with immune-relevant DEGs filtered using ImmuneGO and BTM gene sets. Spearman's correlation (FDR <0.10), PCA with k-means clustering, and GSEA were used to identify condition-specific immune signatures; results are reported primarily as log2 fold changes, proportional DEG distributions, and percent variance explained by principal components.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Kruskal-Wallis test | Comparison of age distributions across the five GEO study cohorts (Fig. 1d) | 245 participants across 5 cohorts | not stated |
| Spearman's rank correlation | Correlation of shared immune and non-immune DEGs across all conditions to identify clustering by vaccine type, regimen, and infection status (Fig. 2b, Fig. S1a,b) | 562 samples aggregated across conditions | not stated |
| Differential gene expression analysis (specific algorithm not named in provided text) | Identification of DEGs per condition vs. baseline, stratified by study and infection status; |L2FC| ≤ 1 used as neutral threshold (Fig. 2a) | — | not stated |
| Principal components analysis (PCA) | Dimensionality reduction and immune response profiling across all conditions using ~1800 immune-related DEGs (Fig. 3) | 562 samples; ~1800 immune-related genes | na |
| K-means clustering | Identification of three distinct immune response clusters in PCA space (Fig. 3a) | — | not stated |
| Gene Set Enrichment Analysis (GSEA) | Pathway enrichment of DEGs using ImmuneGO and BTM gene sets across conditions (Fig. 1a, Fig. 2a) | — | not stated |
-
Differential gene expression analysis was applied across multiple conditions within integrated multi-cohort RNAseq data, but the specific DGE algorithm is not named in the provided text↳ Could also: Established bulk RNAseq DGE tools such as DESeq2 (Wald test), edgeR (likelihood ratio or quasi-likelihood F-test), or limma-voom could each be applied; in a multi-cohort integration, batch-aware frameworks (e.g., limma with study as a blocking covariate, or ComBat-seq for upstream batch correction) are also widely used — Naming and justifying the DGE algorithm aids reproducibility and allows readers to assess model assumptions; batch-aware approaches explicitly account for inter-study variability inherent in combining five GEO datasets from different populations and protocols
-
Spearman's correlation was used to characterize shared DEG patterns across conditions, with FDR <0.10 as the significance threshold↳ Could also: A more stringent FDR threshold (e.g., 0.05) is conventional in transcriptomic analyses; distance-based clustering metrics or mutual information could additionally capture non-monotonic co-expression relationships not detected by rank correlation — A 0.05 FDR threshold reduces the expected proportion of false discoveries; more general dependency measures may surface patterns where the relationship between conditions is not strictly monotonic
-
K-means clustering was applied to PCA scores to define three immune response clusters, with k fixed at 3↳ Could also: Hierarchical clustering (e.g., Ward linkage with Euclidean or correlation distance) or model-based clustering (e.g., Gaussian mixture models) could also partition PCA space; the number of clusters could be chosen data-adaptively using the gap statistic or average silhouette width — K-means assumes spherical, equally sized clusters and requires pre-specifying k; hierarchical approaches provide a dendrogram showing the full merging hierarchy, and data-adaptive k-selection methods give an objective basis for the chosen number of clusters
-
PCA was used as the primary dimensionality reduction and visualization method for immune response profiles across conditions↳ Could also: Non-linear methods such as UMAP or t-SNE could also visualize local structure among conditions; supervised approaches such as partial least squares discriminant analysis (PLS-DA) could emphasize variance related to predefined group labels — PCA captures global linear variance; non-linear embeddings can reveal cluster separation not visible in the first two PCs, which here together explained 85.4% of variance but may still compress finer between-condition structure
-
Sex was imputed for one dataset (GSE206023) using expression of four Y-linked and one X-linked marker genes normalized to TPM↳ Could also: Dedicated computational tools for sex inference from RNA-seq (e.g., checkSex in R, or XYalign) implement validated multi-gene classifiers with defined decision thresholds; alternatively, the dataset could be analyzed with sex treated as an unknown covariate or excluded from sex-stratified analyses — Purpose-built tools apply decision boundaries validated across large reference panels and provide documented, reproducible criteria, reducing the risk of misclassification from single-gene outliers or platform-specific TPM normalization effects
-
The Kruskal-Wallis test was used as an omnibus comparison of age distributions across the five cohorts↳ Could also: A post-hoc pairwise test (e.g., Dunn's test with Bonferroni or Benjamini-Hochberg correction) could follow the omnibus result to identify which specific cohort pairs differ in age, since age is a potential confound in the integrated analysis — The omnibus test indicates that at least one pair differs but does not specify which; pairwise follow-up tests identify the specific between-cohort age differences that may need to be addressed as covariates in downstream DGE or correlation models
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-40456760 (COVID-19 Vaccination Atlas, NPJ Vaccines 2025)
Paper: integrative systems-vaccinology atlas integrating 5 GEO RNA-seq studies (GSE189039, GSE199750, GSE201530, GSE201533, GSE206023) of PBMC transcriptomes after COVID-19 vaccination / infection. Code: github.com/wapsyed/covidvax_atlas (commit 45f14be, CC0). Data: GEO (GSE206023 is this RU's headline accession; the analysis integrates all five).
Pipeline-derived results (IN principle reproducible)
| Result | Pipeline / tool | In scope here? |
|---|---|---|
Per-study DEG tables (DESeq2 Wald, ~ age + sex + condition, FDR<0.05, |L2FC|>1) |
DESeq2 | YES — primary target (GSE206023) |
| Sample/participant counts (562 samples / 245 participants / 5 studies) | metadata count | YES — trivial data point |
| Immune-gene filtering (~16k→~1800 immune genes) | annotation join (ImmuneGO+BTM) | maybe (annotation files shipped) |
| PCA variance (Fig 3a 85.4%, Fig 5c 89%) + Table 3 top-20 PC genes | prcomp on DEG L2FC matrix | OUT (last-20%, downstream of DGE) |
| GSEA bubble plots (FDR<0.10) | fgsea/clusterProfiler | OUT (last-20%) |
| Cell populations (Fig 4) | xCell via immunedeconv | OUT (last-20%) |
| 7 marker genes LHFPL2/FYN/TYMP/JUP/SERINC5/KIT/GRAMD1C (abstract) | ML/PCA selection | OUT (ML notebook; note for fabrication check) |
Out of scope (not pipeline / not attempted)
- Wet-lab / clinical interpretation, vaccine-landscape curation (manual).
- The full multi-study integration, ML descriptor models, interactive Shiny apps.
Primary reproduction target (80/20)
Re-run the GSE206023 DESeq2 DGE step exactly as coded in
Article_Covid19VaxAtlas.Rmd (chunk ~L4000): subset all_matrices_counts.rds
to GSE206023 samples, build colData (condition merged D0→"H (D0)" baseline +
releveled, sex factor, age z-scaled), DESeqDataSetFromMatrix(design=~age+sex+condition),
keep rowSums>=10, DESeq(), extract results() per resultsName. Compare the
recomputed per-gene log2FoldChange / padj against the shipped
DGE analysis/gse206023_DEGs.csv. DESeq2 is deterministic → expect near-exact.
Rationale: this is the central, clearly-specified pipeline step, fully self-contained from shipped inputs. PCA/GSEA/xCell are downstream of it and are the optional hard 20%.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The central computational step — per-study DESeq2 DGE for GSE206023 on the authors' shipped CC0 count matrix — reproduces near-1:1 (98.3% of log2FC within 0.01; padj r=0.994; method params verbatim), so the core claim holds. The only divergences are explainable: the D28 small-n DEG counts differ 2-3x because we did not apply the paper's documented dynamic FDR filter (our methodology gap), and the extreme-gene Pearson dip is numerical instability. The one genuine authors'-side flag is that the headline 562 samples / 245 participants is not derivable from the shipped metadata (554 / 225), off by 8 samples / 20 participants — a reporting/availability inconsistency for human adjudication, not fabrication. Net: solid partial reproduction, deviations are moderate and mostly on the data-tally/reporting and our-method side.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.