VIGET: A web portal for study of vaccine-induced host responses based on Reactome pathways and ImmPort data.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce WITHOUT author contact, and reproduced on the authors' OWN public data with their OWN pipeline code (P16 valid). VIGET = limma DE -> Reactome enrichment; backend data is public on Zenodo 7407195 (1.14 GB matrix + metadata), pipeline code on github.com/VIOLINet/immport-ws @ d4aa548. EXACT 1:1 on structural claims: integrated matrix 22,343 genes x 4,859 GSM (C0) and the yellow-fever cohort 482 samples / 4 ImmPort studies / 3 vaccines (C1). The case-study pathway pipeline reproduces the paper's full TEMPORAL pattern with the correct timepoint-specific top pathways (C13): day-7 Interferon signaling, day-14 Cell Cycle / Cell Cycle Mitotic (top-2), day-28 Neutrophil degranulation (#1) + Immune/Innate/IL-10. Pathway identities and ranks match; the exact Reactome FDR values differ (up to ~4 orders) because Reactome's database changed between the paper's ~2022 query and our 2026 query - a known version artifact (cf. Enrichr library shift). Floor-region FDRs (Neutrophil degranulation, Innate Immune System, Immune System @ d28) reproduce within ~2x. The paper's suspicious repeated FDR=1.80E-14 across 5 pathways is NOT fabrication: we reproduce the same shared-floor behaviour (our floor 3.03E-14), confirming it is the Reactome AnalysisService minimum-FDR floor. One pathway (UPR @ d14, C4) was not recovered as enriched -> mismatch. NOT attempted (the hard 20%): rebuilding the full 28-study / 4,859-sample ImmPort + 12-platform harmonisation (we consumed the released integrated matrix), the exact author covariate/correction toggles for each figure, and the web UI / FI-network rendering. Overall: PARTIAL - structurally exact, qualitatively faithful, exact enrichment FDRs version-shifted.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 64assessed: 2026-06-14 ⛓ 585d3337b251
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetCan a web-based portal that integrates ImmPort/GEO vaccine-response gene expression data with Reactome pathway and functional-interaction-network resources enable efficient, robust analysis of host immune responses to vaccination, and can such a tool reveal informative pathway-level response patterns (demonstrated via a yellow fever vaccine longitudinal study)?
- ★ VIGET is a web portal that lets users select vaccines/ImmPort studies, run differential gene expression analysis, and perform Reactome-based pathway enrichment and functional interaction network construction method
- ★ A longitudinal analysis of yellow fever vaccine responses using VIGET revealed a complex, time-varying activity pattern of immune system pathways annotated in Reactome finding
- ★ The authors curated a comprehensive vaccine-response gene expression resource covering 21 vaccines, 28 ImmPort studies, and 4,859 biosamples from 24 GSE datasets on 12 GPL platforms resource
- Batch effects exist across studies/platforms in the merged dataset and can be partially reduced using BBKNN correction finding
- ★ VIGET uses the Vaccine Ontology (VO) to build hierarchical classifications of vaccines by target pathogen and vaccine type method
- Reactome's Analysis Service (Binomial test with Benjamini-Hochberg FDR correction) is integrated into VIGET for pathway enrichment analysis method
- Influenza vaccine studies dominate the collected dataset relative to other vaccine types finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| differential gene expression analysis (limma) | human blood/PBMC samples from vaccinated subjects (ImmPort/GEO cohorts, e.g. yellow fever vaccine recipients) | vaccination at defined post-vaccination time points (e.g. 7, 14, 28 days vs day 0) | log2 fold change and adjusted p-values for gene expression | limma R package v3.50.1 on 12 GPL microarray platforms |
| pathway enrichment analysis | lists of differentially expressed genes | none | enriched Reactome pathways with corrected p-values | Reactome Analysis Service (Binomial test, Benjamini-Hochberg FDR) |
| functional interaction (FI) network construction | differentially expressed gene sets | none | network of functional gene-gene interactions | Reactome Functional Interaction service (reactome-fiviz) |
| PCA analysis | individual GSE gene expression datasets | none | detection of batch effects for manual annotation | R (processor.R script) |
| UMAP visualization with BBKNN batch correction | all 4,859 GSM samples across vaccines/platforms/GSEs | none | sample clustering pattern colored by vaccine, platform, and GSE | BBKNN (single-cell RNA-seq batch correction method) |
| microarray gene expression profiling (data collection) | human subjects across 28 ImmPort-annotated vaccine studies | various vaccines (influenza, yellow fever, meningococcal, HIV, M. tuberculosis, S. pneumoniae, Varicella-Zoster) | gene expression matrices mapped to official gene symbols | 12 GPL microarray platforms via GEO |
| boxplot distribution analysis | samples grouped by vaccine | none | interquartile range and outliers of expression values | — |
- – Final curated dataset comprises 22,343 genes across 4,859 GSM samples from 28 ImmPort studies, 24 GSE records, 12 GPL platforms, and 21 vaccines (20 VO ids)
- ▼ Filtering pipeline reduced ImmPort ExpSample records to a vaccination-focused gene expression cohort 5,817 of 11,692 samples retained
- ▲ Influenza virus vaccine studies dominate the dataset 3,317 GSM samples (68% of total) across 16 studies (57% of 28)
- – Yellow fever vaccine samples form the second-largest group 482 GSM samples across 4 studies
- – BBKNN correction reduced but did not fully eliminate batch effects (e.g., GPL10558 samples remained separated)
- – Two influenza vaccines showed much wider expression value distributions than other vaccines IQR 3-130 vs. typical 3-9
- – Longitudinal yellow fever vaccine analysis showed a complex activity response pattern across immune system pathways over the post-vaccination time course
- count 5,817 of 11,692 samples (samples retained after ImmPort filtering pipeline for vaccination focus)
- count 4,859 GSM samples; 22,343 genes (final merged gene expression matrix)
- count 28 ImmPort studies, 24 GSE records, 12 GPL platforms, 21 vaccines (20 VO ids) (coverage of the final curated dataset)
- other 68% of samples (3,317/4,859), 57% of studies (16/28) (proportion of dataset attributable to influenza vaccine studies)
- count 482 GSM samples, 4 studies (yellow fever vaccine sample coverage)
- other IQR 3-9 (typical) vs 3-130 (two influenza vaccines) (expression value distribution comparison across vaccines)
- other log2 fold change threshold > 0.2 or < -0.2 (gene selection criterion for pathway enrichment analysis in yellow fever case study)
- other Benjamini-Hochberg FDR-corrected p-values from Binomial test (statistical method used by Reactome Analysis Service for pathway enrichment)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
VIGET is a web portal for vaccine gene expression analysis that integrates limma-based differential gene expression (DEG) analysis with Reactome pathway enrichment. For the demonstration analysis of yellow fever vaccines, a paired limma linear model was fitted adjusting for vaccine, age, gender, race, platform, and batch, across five pairs of time points post-vaccination (days 0 vs 7, 14, 28; and consecutive intervals days 7 vs 14, 14 vs 28). Genes with |log2 fold change| > 0.2 were submitted to Reactome pathway enrichment using a binomial test with Benjamini-Hochberg FDR correction. The paper is primarily a tool-description paper; statistical results are conveyed through pathway-level visualizations rather than tabulated numeric summaries.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma empirical Bayes moderated linear model (moderated t-statistics) | Differential gene expression across five pairwise time-point comparisons in the yellow fever vaccine longitudinal analysis | Subset of 482 yellow fever GSM samples from 4 ImmPort studies; exact per-comparison n not stated | not stated |
| Binomial test (over-representation analysis) | Reactome pathway enrichment analysis applied to genes with |log2FC| > 0.2 from each DEG comparison | Number of submitted genes per comparison; not stated | not stated |
| PCA (principal component analysis) | Batch effect inspection across all 4,859 GSM samples; used to decide per-dataset batch annotation | 4,859 GSM samples | na |
| UMAP (Uniform Manifold Approximation and Projection) | Visualization of all samples colored by vaccine, platform, and GSE before and after BBKNN batch correction | 4,859 GSM samples | na |
-
Genes for pathway enrichment were selected using a log2 fold change threshold alone (|log2FC| > 0.2), independent of adjusted p-value↳ Could also: A combined threshold (e.g., adjusted p-value < 0.05 AND |log2FC| > threshold) or a fully ranked approach such as GSEA (Gene Set Enrichment Analysis) could also be used — Combined or ranked approaches incorporate statistical evidence alongside effect size; GSEA in particular avoids a binary cutoff and can detect pathway-level shifts that are distributed across many modestly changed genes — the authors explicitly chose the fold-change-only approach to increase sensitivity for immune response pathways
-
Pathway enrichment used an over-representation analysis (ORA) with a binomial test via Reactome↳ Could also: GSEA, Fisher's exact test-based ORA, or fgsea (fast pre-ranked GSEA in R) could also be applied to the same gene lists or rankings — Fisher's exact test is the most common ORA reference method, enabling direct cross-tool comparison; GSEA uses the full ranked list without a cutoff, which can improve detection of pathways with many small-effect genes
-
Batch effects were handled by including GSE-derived batch as a covariate in the limma linear model↳ Could also: Explicit batch correction methods such as ComBat (sva R package) or limma's removeBatchEffect could also be applied to the expression matrix prior to analysis — Covariate-based adjustment preserves within-batch variance for inference while explicit matrix correction is useful when the corrected values are needed for downstream steps such as clustering or visualization independent of limma
-
The five time-point comparisons were conducted as five separate DEG analyses without a stated cross-comparison multiplicity adjustment↳ Could also: A single repeated-measures model (e.g., limma with a time factor and contrast matrix, or a mixed-effects model) could also test all time points jointly with a single family-wise correction — A joint model explicitly accounts for correlation across repeated observations on the same subjects and controls the family-wise error rate across all comparisons simultaneously
-
Paired samples were incorporated via limma's blocking design (subjects paired across time points)↳ Could also: A linear mixed-effects model (e.g., lme4 or nlme in R) could also model repeated measures, accommodating subjects with incomplete time-point data — Mixed-effects models handle unbalanced designs where not all subjects have measurements at every time point, potentially retaining more data than a strict pair-wise complete-case approach
-
Batch effect assessment relied on visual inspection of PCA and UMAP plots↳ Could also: Quantitative metrics such as PVCA (Principal Variance Component Analysis) or kBET could also be used alongside visual methods — Quantitative batch metrics provide an objective, reportable measure of batch effect magnitude before and after correction, complementing the visual evidence shown in the figures
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
- A curated collection of transcriptome datasets... L1 62/100
- PulmonDB: a curated lung disease gene expressi...⚑ L1 53/100 ⚑
- Curation of over 10 000 transcriptomic studies... L1 80/100
- Exploration of the shared diagnostic genes and... L1 76/100
- Exploring the key genetic association between... L1 91/100
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37180100 (VIGET)
Paper: Brunson T, Sanati N, Huffman A, Masci AM, Zheng J, Cooke MF, Conley P,
He Y, Wu G. VIGET: A web portal for study of vaccine-induced host responses
based on Reactome pathways and ImmPort data. Front Immunol 2023.
PMID 37180100 · PMCID PMC10172660 · DOI 10.3389/fimmu.2023.1141030
Code: https://github.com/VIOLINet/immport-ws (Java/Spring REST API + R analysis
scripts; default branch master, commit d4aa548… 2024-01-31). Frontend (Vue):
github.com/VIOLINet/reactome-immport-web.
Backend data (public): Zenodo record 7407195 (DOI 10.5281/zenodo.7407195):
immport_vaccine_expression_matrix_mapped_merged_approved_genes_091421.csv(1.14 GB) — the integrated gene-expression matrix (≈22,343 genes × 4,859 GSM).ImmuneExposureGeneExpression_020922.csv(3.1 MB) — biosample / immune-exposure metadata. Brief's data accession: GEO GSE13486 — one of the yellow-fever GEO series folded into the integrated matrix (the case study spans 4 ImmPort studies / 24 GSE).
Nature of the paper
VIGET is a web portal (like a tool paper). Its pipeline, per Methods and the repo's R scripts, is:
- Differential gene expression with limma 3.50.1 on the integrated
expression matrix (
de_analysis/differential_exp_analysis.R): design~ immport_vaccination_time_groups + … + type_subtype(+ optional gpl / batch / paired corrections);lmFit → eBayes → topTable(coef=2, n=Inf). - Reactome pathway enrichment of the DE genes via Reactome's analysis web service: binomial test, p-values corrected by Benjamini-Hochberg FDR. Genes submitted are those with log2FC > 0.2 or < −0.2 (Methods).
- Functional-interaction network construction (visual; not a numeric result). There is no wet-lab content; per brief P16, running this pipeline (limma + Reactome) on the paper's own released data is a fully valid reproduction.
The reported result we target — yellow-fever case study (Figure 9, Tables S6/S9/S10)
The demonstration analyses 482 samples from 4 ImmPort YF studies (vaccines
YF-Vax, Stamaril, YF-17D vector), comparing post-vaccination timepoints to day 0.
Reactome enrichment tables report per-pathway FDR values. Pinned claims (see
original/claims.tsv):
- Day 14 vs 0 (Table S6 / Fig 9C): Cell Cycle FDR = 3.62E-14; Cell Cycle, Mitotic FDR = 3.62E-14; Unfolded Protein Response FDR = 3.50E-05.
- Day 28 vs 0 (Table S9): Interferon Signaling, Interferon alpha/beta signaling, Innate Immune System, Neutrophil degranulation, Interleukin-10 signaling all at FDR = 1.80E-14; Signaling by Interleukins 8.89E-12; Immune System & Cytokine Signaling 1.80E-12.
- Day 28 vs 0 (Table S10): Cell Cycle FDR = 2.38E-07; Cell Cycle, Mitotic 4.19E-09; UPR 3.11E-08.
Note on the repeated 1.80E-14: five distinct Day-28 pathways share the identical FDR. This is almost certainly Reactome's reported minimum FDR floor (smallest representable corrected p given the number of pathways tested), not five coincidentally-equal tests — recorded as a soft/floor value, flagged for the human reviewer (not a fabrication note per se, but a clamped value).
In scope (attempted)
- Reproduce the limma DE for Day-14-vs-0 and Day-28-vs-0 on the YF samples of the released matrix, then Reactome enrichment of the |log2FC|>0.2 gene set, and compare the FDR / rank of the named pathways (Cell Cycle, UPR, Neutrophil degranulation, IFN α/β, IL-10, Interferon Signaling, Innate Immune System).
- The qualitative temporal pattern (early IFN/innate → week-2 cell-cycle → recovery).
Out of scope / not attempted (with reason)
- Full database build (28 ImmPort studies → 4,859 biosamples, 24 GSE harmonised
to 22,343 genes): the integration/harmonisation across platforms is the upstream
ETL; we consume the released integrated matrix rather than rebuilding it
(
non_pipelinefor that step;
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Reproduced on the authors' own public data (Zenodo 7407195) with their own pipeline code, so input/method are 1:1. Structural claims are exact (C0 22,343×4,859; C1 482 samples/4 studies/3 vaccines) and the central temporal pathway program (C13) fully reproduces with correct timepoint-specific top pathways. The only deviations are (a) exact Reactome FDRs shifted by up to ~4 orders due to the 2022→2026 database version — a technical/expected metric-version artifact, including the reproducible FDR-floor that clears the repeated-1.80E-14 fabrication suspicion — and (b) one isolated miss (UPR @d14, C4). Deviations lie on the reference-database/version side, not the authors' integrity or the pipeline logic, so this is a solid reproduction with explainable deviations (yellow overall).
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.