Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

COVID-19 vaccination atlas using an integrative systems vaccinology approach.

NPJ Vaccines · 2025
L1 69/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
69/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 35% of all assessed papers rank 745 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce 1:1. Authors' own R code (CC0, commit 45f14be) ships the integer count matrix + sample metadata + DESeq2 outputs for 5 integrated GEO RNA-seq studies. I re-ran the central per-study DGE step for GSE206023 exactly as coded (DESeqDataSetFromMatrix design=~age+sex+condition; keep rowSums>=10; DESeq(); results() per coefficient) on «our HPC» (R 4.3.3 / DESeq2 1.42.0). Against the canonical shipped table gse206023_DEGs_wald.csv (the one the pipeline reads back), 98.3% of 106,866 gene x contrast log2FoldChange values match to <0.01 and padj correlate at r=0.994 -> the core DGE pipeline reproduces near-exactly. Two honest gaps: (a) the two smallest-sample D28 contrasts yield 2-3x more DEGs in my raw run, consistent with the paper's documented dynamic FDR filter for sub-median-n comparisons (a downstream step I did not apply); (b) FLAG: the paper's headline 562 samples / 245 participants is not exactly derivable from the shipped metadata, which has 554 sample rows / 225 unique participant_ids (per-study counts reconcile internally to 554) - a reviewer should adjudicate, not asserted as fabrication. NOT attempted (the hard 20%, downstream of DGE): full multi-study PCA (Fig3a 85.4% variance), Table 3 top-20 PC genes, GSEA bubble plots, xCell cell deconvolution (Fig4), the ML marker-gene selection (LHFPL2/FYN/TYMP/JUP/SERINC5/KIT/GRAMD1C), and re-deriving counts from raw GEO FASTQ. The reproduction used the authors' shipped count matrix so the DGE input matches theirs exactly.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 69
    assessed: 2026-06-14 ⛓ 1e2fd1534dc3
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether an integrative systems vaccinology approach applied to RNAseq data from multiple cohorts can reveal shared and unique immune signatures across different COVID-19 vaccine types, technologies, and regimens (including infection status), and how these compare to vaccines against other pathogens.

Core claims
  • mRNA vaccines induce transient but strong immune responses after booster doses finding
  • Viral vector vaccines show sustained immune activation with infection, partially resembling the infection immune profile finding
  • Heterologous vaccination regimens demonstrate enhanced immune diversification compared to homologous regimens finding
  • Immune-related DEGs comprise less than 10% of total DEGs but are sufficient to establish condition-specific relationships finding
  • PCA of immune DEGs identifies three distinct clusters of immune responses corresponding to infection, vaccination in healthy individuals, and early infection/first-dose states finding
  • A COVID-19 vaccination atlas integrating 562 samples from 245 participants across five vaccines and regimens was constructed and compared with data from 13 vaccines against other diseases resource
  • Shared immune markers (e.g., MARCKS, NFKBIA, IL1B, CCL3) are enriched across both infection and vaccination conditions, not exclusively in one finding
  • Sex was manually imputed for GSE206023 using expression of X- and Y-linked genes (DDX3Y, KDM5D, RPS4Y1, UTY, XIST) method
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq / differential gene expression analysis PBMCs from human participants (245 participants, 5 cohorts) COVID-19 vaccination (IN, SU, RNA, VV types; homologous/heterologous regimens; 1-3 doses) and/or SARS-CoV-2 infection differentially expressed genes (DEGs), classified as up/down/neutral (|L2FC|>1)
Gene Set Enrichment Analysis (GSEA) PBMCs, same cohorts vaccination/infection conditions enrichment of immune-related gene sets (ImmuneGO, BTM) and GO cellular/metabolic processes
Principal component analysis (PCA) with K-means clustering immune DEGs from PBMC samples vaccination/infection conditions variance explained by PC1/PC2, cluster assignment
Spearman's correlation analysis immune and non-immune DEGs across conditions vaccination/infection conditions condition-specific correlation clustering (FDR <0.10)
Sex imputation from gene expression (TPM) PBMC samples (GSE206023) none (metadata reconstruction) expression of DDX3Y, KDM5D, RPS4Y1, UTY (Y-linked) and XIST (X-linked)
Immune cell population dynamics analysis PBMC samples across cohorts vaccination/infection conditions cell population composition changes
Key results
  • Vaccinated individuals (e.g., two BNT doses) exhibited higher total DEGs post-infection compared to unvaccinated controls, driven by over 200 upregulated genes (BNT-I D0-5 vs. I D0-4) >200 genes
  • PC1 and PC2 explained 72.8% and 12.6% of variance respectively, together capturing 85.4% of total variance 72.8% + 12.6% = 85.4%
  • Immune-related genes reduced the analysis variable set from ~16,000 genes to ~1800 immune genes ~16,000 to ~1800
  • Reinfected patients displayed fewer upregulated DEGs despite similar downregulation trends compared to non-reinfected
  • DEG counts in hospitalized cohorts peaked at day 10 post-admission in partially vaccinated (BNT-I D10) patients but declined thereafter
  • Immune-related genes constituted less than 10% of total DEGs and declined over time in vaccinated and infected individuals <10%
  • Shared immune markers MARCKS, NFKBIA, IL1B, and CCL3 were enriched across both infection and vaccination conditions
Key statistics
  • count 562 samples (Total samples analyzed across the atlas)
  • count 245 participants (Total participants across five cohorts/studies)
  • other PC1=72.8%, PC2=12.6% variance explained (PCA of immune DEGs across all conditions)
  • mean median age 35 years (IQR: 17-83) (Overall study population age)
  • count >200 upregulated genes (BNT-I D0-5 vs. I D0-4 comparison)
  • pvalue FDR <0.10 (Threshold used for Spearman's correlation analysis of DEGs)
  • fold_change vaccine efficacies: BNT 95.0%, MO 94.1%, ChAd 90% (Clinical trial efficacy of homologous vaccination regimens)
  • fold_change BBIBP 78.1%, ZF2001 75.7% efficacy against symptomatic infection (Efficacy of classical inactivated/subunit vaccines)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study performed an integrative reanalysis of bulk RNAseq data from five publicly available GEO cohorts (562 samples, 245 participants) covering five COVID-19 vaccine types in homologous and heterologous regimens as well as infected groups. Differential gene expression analysis was applied stratified by study and condition, with immune-relevant DEGs filtered using ImmuneGO and BTM gene sets. Spearman's correlation (FDR <0.10), PCA with k-means clustering, and GSEA were used to identify condition-specific immune signatures; results are reported primarily as log2 fold changes, proportional DEG distributions, and percent variance explained by principal components.

Replicationbiological Sample size562 samples from 245 participants across five GEO cohorts; per-condition sample sizes visualized in Fig. 1c; no formal sample-size or power calculation described GroupsFive vaccine types (RNA, VV, IN, SU) × homologous/heterologous regimens × dose number (1–3) × infection status (unvaccinated, vaccinated, infected, vaccinated-then-infected, reinfected) at multiple longitudinal time points Pairingmixed Randomization/blindingnot stated DispersionIQR Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionFDR (specific algorithm, e.g. Benjamini-Hochberg, not named in provided text)
Statistical tests used
Test Applied to n Assumptions
Kruskal-Wallis test Comparison of age distributions across the five GEO study cohorts (Fig. 1d) 245 participants across 5 cohorts not stated
Spearman's rank correlation Correlation of shared immune and non-immune DEGs across all conditions to identify clustering by vaccine type, regimen, and infection status (Fig. 2b, Fig. S1a,b) 562 samples aggregated across conditions not stated
Differential gene expression analysis (specific algorithm not named in provided text) Identification of DEGs per condition vs. baseline, stratified by study and infection status; |L2FC| ≤ 1 used as neutral threshold (Fig. 2a) not stated
Principal components analysis (PCA) Dimensionality reduction and immune response profiling across all conditions using ~1800 immune-related DEGs (Fig. 3) 562 samples; ~1800 immune-related genes na
K-means clustering Identification of three distinct immune response clusters in PCA space (Fig. 3a) not stated
Gene Set Enrichment Analysis (GSEA) Pathway enrichment of DEGs using ImmuneGO and BTM gene sets across conditions (Fig. 1a, Fig. 2a) not stated
Approaches that could also have been used
  • Differential gene expression analysis was applied across multiple conditions within integrated multi-cohort RNAseq data, but the specific DGE algorithm is not named in the provided text
    Could also: Established bulk RNAseq DGE tools such as DESeq2 (Wald test), edgeR (likelihood ratio or quasi-likelihood F-test), or limma-voom could each be applied; in a multi-cohort integration, batch-aware frameworks (e.g., limma with study as a blocking covariate, or ComBat-seq for upstream batch correction) are also widely used — Naming and justifying the DGE algorithm aids reproducibility and allows readers to assess model assumptions; batch-aware approaches explicitly account for inter-study variability inherent in combining five GEO datasets from different populations and protocols
  • Spearman's correlation was used to characterize shared DEG patterns across conditions, with FDR <0.10 as the significance threshold
    Could also: A more stringent FDR threshold (e.g., 0.05) is conventional in transcriptomic analyses; distance-based clustering metrics or mutual information could additionally capture non-monotonic co-expression relationships not detected by rank correlation — A 0.05 FDR threshold reduces the expected proportion of false discoveries; more general dependency measures may surface patterns where the relationship between conditions is not strictly monotonic
  • K-means clustering was applied to PCA scores to define three immune response clusters, with k fixed at 3
    Could also: Hierarchical clustering (e.g., Ward linkage with Euclidean or correlation distance) or model-based clustering (e.g., Gaussian mixture models) could also partition PCA space; the number of clusters could be chosen data-adaptively using the gap statistic or average silhouette width — K-means assumes spherical, equally sized clusters and requires pre-specifying k; hierarchical approaches provide a dendrogram showing the full merging hierarchy, and data-adaptive k-selection methods give an objective basis for the chosen number of clusters
  • PCA was used as the primary dimensionality reduction and visualization method for immune response profiles across conditions
    Could also: Non-linear methods such as UMAP or t-SNE could also visualize local structure among conditions; supervised approaches such as partial least squares discriminant analysis (PLS-DA) could emphasize variance related to predefined group labels — PCA captures global linear variance; non-linear embeddings can reveal cluster separation not visible in the first two PCs, which here together explained 85.4% of variance but may still compress finer between-condition structure
  • Sex was imputed for one dataset (GSE206023) using expression of four Y-linked and one X-linked marker genes normalized to TPM
    Could also: Dedicated computational tools for sex inference from RNA-seq (e.g., checkSex in R, or XYalign) implement validated multi-gene classifiers with defined decision thresholds; alternatively, the dataset could be analyzed with sex treated as an unknown covariate or excluded from sex-stratified analyses — Purpose-built tools apply decision boundaries validated across large reference panels and provide documented, reproducible criteria, reducing the risk of misclassification from single-gene outliers or platform-specific TPM normalization effects
  • The Kruskal-Wallis test was used as an omnibus comparison of age distributions across the five cohorts
    Could also: A post-hoc pairwise test (e.g., Dunn's test with Bonferroni or Benjamini-Hochberg correction) could follow the omnibus result to identify which specific cohort pairs differ in age, since age is a potential confound in the integrated analysis — The omnibus test indicates that at least one pair differs but does not specify which; pairwise follow-up tests identify the specific between-cohort age differences that may need to be addressed as covariates in downstream DGE or correlation models
Software: BioArt; Canva (figure creation only)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40456760 (COVID-19 Vaccination Atlas, NPJ Vaccines 2025)

Paper: integrative systems-vaccinology atlas integrating 5 GEO RNA-seq studies (GSE189039, GSE199750, GSE201530, GSE201533, GSE206023) of PBMC transcriptomes after COVID-19 vaccination / infection. Code: github.com/wapsyed/covidvax_atlas (commit 45f14be, CC0). Data: GEO (GSE206023 is this RU's headline accession; the analysis integrates all five).

Pipeline-derived results (IN principle reproducible)

Result Pipeline / tool In scope here?
Per-study DEG tables (DESeq2 Wald, ~ age + sex + condition, FDR<0.05, |L2FC|>1) DESeq2 YES — primary target (GSE206023)
Sample/participant counts (562 samples / 245 participants / 5 studies) metadata count YES — trivial data point
Immune-gene filtering (~16k→~1800 immune genes) annotation join (ImmuneGO+BTM) maybe (annotation files shipped)
PCA variance (Fig 3a 85.4%, Fig 5c 89%) + Table 3 top-20 PC genes prcomp on DEG L2FC matrix OUT (last-20%, downstream of DGE)
GSEA bubble plots (FDR<0.10) fgsea/clusterProfiler OUT (last-20%)
Cell populations (Fig 4) xCell via immunedeconv OUT (last-20%)
7 marker genes LHFPL2/FYN/TYMP/JUP/SERINC5/KIT/GRAMD1C (abstract) ML/PCA selection OUT (ML notebook; note for fabrication check)

Out of scope (not pipeline / not attempted)

  • Wet-lab / clinical interpretation, vaccine-landscape curation (manual).
  • The full multi-study integration, ML descriptor models, interactive Shiny apps.

Primary reproduction target (80/20)

Re-run the GSE206023 DESeq2 DGE step exactly as coded in Article_Covid19VaxAtlas.Rmd (chunk ~L4000): subset all_matrices_counts.rds to GSE206023 samples, build colData (condition merged D0→"H (D0)" baseline + releveled, sex factor, age z-scaled), DESeqDataSetFromMatrix(design=~age+sex+condition), keep rowSums>=10, DESeq(), extract results() per resultsName. Compare the recomputed per-gene log2FoldChange / padj against the shipped DGE analysis/gse206023_DEGs.csv. DESeq2 is deterministic → expect near-exact.

Rationale: this is the central, clearly-specified pipeline step, fully self-contained from shipped inputs. PCA/GSEA/xCell are downstream of it and are the optional hard 20%.

Figures / tables: Table
C1
Reported
48 GSE206023 samples
Reproduced
48 samples entered DESeq2
exact
C2
Reported
shipped DESeq2 DEG table (gse206023_DEGs_wald.csv), deterministic
Reproduced
98.3% of 106866 gene x contrast log2FC within |d|<0.01; median d=0; padj Pearson r=0.994; L2FC Pearson 0.908
within tolerance
C3
Reported
per-contrast DEG counts from shipped _wald table
Reproduced
4/6 vaccine contrasts match within ~3%; 2/6 (D28, smallest n) differ 2-3x
partial
C4
Reported
562 samples / 245 participants / 5 studies (Results)
Reproduced
554 samples / 225 participants / 5 studies from shipped metadata
did not match
C5
Reported
DESeq2 Wald, ~age+sex+condition, FDR<0.05, |L2FC|>1, BH
Reproduced
confirmed verbatim in shipped code
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 69/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

The central computational step — per-study DESeq2 DGE for GSE206023 on the authors' shipped CC0 count matrix — reproduces near-1:1 (98.3% of log2FC within 0.01; padj r=0.994; method params verbatim), so the core claim holds. The only divergences are explainable: the D28 small-n DEG counts differ 2-3x because we did not apply the paper's documented dynamic FDR filter (our methodology gap), and the extreme-gene Pearson dip is numerical instability. The one genuine authors'-side flag is that the headline 562 samples / 245 participants is not derivable from the shipped metadata (554 / 225), off by 8 samples / 20 participants — a reporting/availability inconsistency for human adjudication, not fabrication. Net: solid partial reproduction, deviations are moderate and mostly on the data-tally/reporting and our-method side.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

195.5 k
tokens (I/O) · 10.2 M incl. cache
15 min
runtime · 0.03 CPU-h
1.1 GB
peak RAM
1
HPC jobs
hummel
machine