Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comprehensive characterization of the antibody responses to SARS-CoV-2 Spike protein finds additional vaccine-induced epitopes beyond those for mild infection.

Elife · 2022
L1 57/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Any deviation was negligible
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
57/100
Reproducibility score
1.0 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 17% of all assessed papers rank 965 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction of Garrett/Galloway et al. eLife 73490 (Phage-DMS of SARS-CoV-2 Spike antibody epitopes). Repo matsengrp/phage-dms-vacc-analysis is exemplary (self-contained pipeline + vendored phippery + published figures). PIPELINE MECHANICS REPRODUCED FAITHFULLY: rebuilt bowtie2 reference from peptide_table, aligned (bowtie2 --trim3 32) and produced per-peptide counts that are BYTE-IDENTICAL to the authors' samtools idxstats/stats toolchain (validated), then phippery collect -> library-relative enrichment + scaled differential selection. STATIC CLAIMS GRADED: peptide library 24,840 vs paper 24,820 (MISMATCH +20); Moderna/Cohort1 = 49 (EXACT); HAARVI/Cohort2 = 65 vs 64 (off-by-1); 5 epitope regions + binding thresholds (EXACT). KEY DATA-AVAILABILITY FINDING: NCBI retained only normalized .sra (original per-sample tars purged). Per-sample identity recovered via SRA spot-group demux for only 44/633 samples (6 MiSeq runs by Illumina S-number + run 811 by dual-index barcode). The 4 HiSeq runs holding 573 samples - the ENTIRE Moderna vaccine cohort, main HAARVI cohort, and ALL beads controls - were deposited with only the single 8bp i7 index (12-14 distinct i7 for 173-200 dual-indexed samples), so per-sample demultiplexing is IMPOSSIBLE; the authors offer 'pre-processed enrichment data on request', confirming the raw deposit is not self-sufficient. Consequently the headline cohort-level results (vaccine-induced epitopes, per-cohort escape sites, all manuscript figures) cannot be independently reproduced from public data. Downstream demonstrated on the recoverable 30 HAARVI samples: scaled differential selection yields escape signal in the correct epitope regions (CTD 562 exact; NTD/FP/SH-H within +/-1-3) but not a precise per-residue match. NOT a pipeline/code failure - a public-data-deposit limitation. healthy=true (outcome correctly determined; compute ran).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 57
    assessed: 2026-06-21 ⛓ e160db6b83eb
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether SARS-CoV-2 infection and mRNA vaccination elicit antibodies that bind similar epitopes on the Spike protein and share similar escape pathways, given that most people will first encounter Spike via vaccination rather than infection.

Core claims
  • mRNA vaccination induces antibody binding to additional Spike epitopes (NTD and CTD in S1) beyond those seen after mild infection (FP and SH-H in S2) finding
  • Antibodies from hospitalized/severe infection show an epitope binding pattern similar to vaccinated individuals, including NTD and CTD binding finding
  • Vaccination induces a highly uniform escape profile at the SH-H epitope across individuals, whereas infection produces much more variable escape pathways at this epitope finding
  • The escape pathway for the FP epitope established after infection is not altered by subsequent vaccination finding
  • Covariates such as vaccine dose, vaccine type (mRNA-1273 vs BNT162b2), and age do not significantly affect antibody binding to the four identified epitopes finding
  • Phage-DMS (deep mutational scanning phage display) can be used to comprehensively profile linear antibody epitopes and escape mutations across the full Spike protein method
  • Binding magnitude to CTD and SH-H epitopes decreases between 36 and 119 days post first vaccine dose finding
  • Data and an interactive web tool from this study were made publicly available resource
Experimental setups
Assay System Perturbation Readout Platform
Phage-DMS (deep mutational scanning phage display library, immunoprecipitation followed by sequencing) Human serum, Moderna Trial Cohort (49 individuals; 100 µg or 250 µg mRNA-1273 dose) mRNA-1273 vaccination Enrichment of wild-type and mutant Spike peptides bound by serum antibodies (epitope binding and escape) T7 bacteriophage display library tiling Spike in 1-aa increments (31-aa peptides) with IP-sequencing
Phage-DMS (deep mutational scanning phage display library, immunoprecipitation followed by sequencing) Human serum, HAARVI Cohort (64 individuals: infected mild/severe, naive, some also vaccinated) Natural SARS-CoV-2 infection (mild or severe) and/or subsequent mRNA vaccination (mRNA-1273 or BNT162b2) Enrichment of wild-type and mutant Spike peptides bound by serum antibodies (epitope binding and escape) T7 bacteriophage display library tiling Spike in 1-aa increments (31-aa peptides) with IP-sequencing
Principal component analysis on wild-type peptide enrichment features Combined serum samples from Moderna Trial and HAARVI cohorts none (computational analysis of Phage-DMS data) Loading vectors/components identifying epitope regions (NTD, CTD, FP, SH-H) that differentiate sample groups
Key results
  • Hospitalized/severe infection and vaccinated individuals showed significantly higher antibody binding to NTD, CTD, and SH-H epitopes than nonhospitalized infected individuals
  • Nonhospitalized infected individuals showed significantly higher binding to the FP epitope than hospitalized or vaccinated individuals
  • No significant difference in epitope binding among the four regions between vaccinated individuals with vs. without prior infection p>0.05 (MWW)
  • Significantly decreased binding to CTD and SH-H epitopes at day 119 vs. day 36 post first vaccine dose p=0.008 (CTD), p=0.011 (SH-H)
  • No significant difference in epitope binding (NTD, CTD, FP, SH-H) between 100 µg and 250 µg mRNA-1273 dose groups
  • Apparent difference in SH-H binding by participant age did not survive multiple testing correction
  • SH-H escape pathway was highly uniform across vaccinated individuals but diverse across mildly infected individuals
  • FP escape pathway established by infection remained unchanged after subsequent vaccination
Key statistics
  • pvalue p=0.008 (CTD epitope binding, day 36 vs. day 119 post first vaccine dose, Wilcoxon rank-sum with Bonferroni correction)
  • pvalue p=0.011 (SH-H epitope binding, day 36 vs. day 119 post first vaccine dose, Wilcoxon rank-sum with Bonferroni correction)
  • count 49 individuals (34 at 100 µg, 15 at 250 µg dose) (Moderna Trial Cohort composition)
  • count 64 individuals (44 infected [39 nonhospitalized/mild, 5 hospitalized/severe], 20 naive) (HAARVI Cohort composition)
  • count 24 of 44 infected individuals also sampled post-vaccination (HAARVI Cohort vaccination follow-up sampling)
  • count n=64 samples at day 36 and n=64 at day 119 post vaccination (Moderna Trial Cohort timepoint sampling)
  • pvalue p>0.05 (No significant epitope binding difference between vaccinated groups with vs. without prior infection (MWW))
  • other 31-amino-acid peptides tiled in 1-amino-acid increments (Design of Spike Phage-DMS library)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used Phage-DMS (phage display deep mutational scanning) to profile serum antibody epitopes and escape sites across SARS-CoV-2 Spike in two cohorts (total n=113 individuals with diverse infection/vaccination histories). Principal component analysis was applied to identify epitope regions driving inter-sample differences, after which enrichment values were summed within each identified epitope region and compared across groups using nonparametric rank-sum tests. Multiple testing was addressed via Bonferroni correction, and results were reported as boxplots with p-values, without explicit effect size or confidence interval estimates.

Replicationbiological Sample sizeModerna Trial Cohort: n=49 (34 at 100 µg, 15 at 250 µg); HAARVI Cohort: n=64 (44 infected, 20 naive; 44 vaccinated); subgroup sizes noted per comparison; no formal power calculation stated Groupsnonhospitalized/mild infected, hospitalized/severe infected, vaccinated-naive, vaccinated-convalescent, unvaccinated-naive Pairingmixed Randomization/blindingnot stated DispersionIQR Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionBonferroni
Statistical tests used
Test Applied to n Assumptions
Principal component analysis (PCA) Identification of Spike epitope regions differentiating infected, vaccinated, and naive sample groups (Figure 2B, Figure 2—figure supplement 1) not stated
Mann-Whitney-Wilcoxon (MWW) test Pairwise comparisons of summed wild-type enrichment within four epitope regions between nonhospitalized infected and all other groups (Figure 2C); also used for vaccinated-with vs vaccinated-without prior infection comparison not stated
Wilcoxon rank-sum test with Bonferroni correction Covariate analyses: timepoint (day 36 vs day 119), dose (100 µg vs 250 µg), and age group comparisons within the Moderna Trial Cohort (Figure 3A–C); 36 group comparisons total n=64 per timepoint (Moderna Trial Cohort) not stated
Mann-Whitney test with Bonferroni correction HAARVI subgroup comparisons by time post symptom onset and by vaccine type (Figure 3—figure supplement 1) not stated
Approaches that could also have been used
  • Group comparisons used pairwise Mann-Whitney-Wilcoxon tests anchored to the nonhospitalized infected group, with Bonferroni correction
    Could also: A Kruskal-Wallis test (nonparametric one-way ANOVA analogue) followed by Dunn's post-hoc test with FDR (Benjamini-Hochberg) adjustment could also compare all groups simultaneously — A global omnibus test before pairwise contrasts is a standard approach that first establishes whether any group difference exists; BH-FDR is less conservative than Bonferroni and can increase power to detect true differences when the number of comparisons is moderate
  • Several participants contributed multiple serum samples at different timepoints, and some individuals appear in more than one comparison group
    Could also: Linear mixed-effects models (e.g., lme4 in R) or generalized estimating equations (GEE) could account for within-individual correlation across repeated measurements — When the same individual contributes samples at multiple timepoints, observations are not independent; mixed models explicitly model this structure and can improve precision of estimates for time and covariate effects
  • PCA was used to identify epitope regions of interest, and enrichment values within those regions were then summed to create per-sample summary scores for downstream testing
    Could also: Sparse PCA, UMAP, or penalized regression (e.g., LASSO) applied to the full enrichment matrix could also prioritize or select epitope features in a single step — These methods can identify non-linear structure or perform variable selection in a more integrated manner, though PCA with manual region definition provides high interpretability and biological legibility
  • Dispersion was conveyed via boxplots (interquartile range and median), and no confidence intervals were reported for group differences
    Could also: Bootstrap confidence intervals on group medians or mean differences could also quantify uncertainty around the central estimates — CIs convey both the direction and precision of an estimated difference and allow readers to assess practical significance independently of p-values, complementing the hypothesis-test results
  • Multiple covariates (age, dose, vaccine type, timepoint) were examined in separate pairwise tests rather than jointly
    Could also: A multivariable regression model (e.g., linear regression or quantile regression on enrichment scores) could assess independent effects of age, dose, type, and timepoint simultaneously — Joint modeling controls for confounding between covariates—for example, age may be correlated with dose group—and provides covariate-adjusted estimates of each factor's contribution
  • Bonferroni correction was applied to 36 comparisons in the covariate analyses
    Could also: Benjamini-Hochberg false discovery rate (FDR) control at 5% could also be applied to the same family of tests — Bonferroni controls the family-wise error rate (probability of any false positive), which is appropriate when individual false positives are costly; BH-FDR controls the expected proportion of false positives and is less conservative in exploratory analyses with many tests, potentially identifying additional associations worth follow-up
Software: GitHub repository (matsengrp/phage-dms-vacc-analysis) swh:1:rev:d4c770ad49ed2f8ab31e499265dd02273cff6f86

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35072628 (Garrett et al. 2022, eLife 73490)

Title: Comprehensive characterization of the antibody responses to SARS-CoV-2 Spike protein finds additional vaccine-induced epitopes beyond those for mild infection.

Code: https://github.com/matsengrp/phage-dms-vacc-analysis (default branch master) Data: SRA PRJNA765705 (11 runs, ~34 GB; each run = a .tar of demultiplexed per-sample fastq.gz) Method: Phage-DMS (deep mutational scanning by phage display) of SARS-CoV-2 Spike.

What the repo ships (fully self-contained, fastq → figures)

  • nextflow-pipeline-config/peptide_table.csv — the phage-display peptide library (24,840 rows: WT + single-aa-mutant 31-mers tiling Spike). This is the bowtie2 reference.
  • nextflow-pipeline-config/sample_table.csv — 633 samples (metadata: cohort, control_status {empirical|library|beads_only}, library_batch {SPIKE1|SPIKE2}, visit, dose, fastq_filename …).
  • nextflow-pipeline-config/PhIP-analysis.nf — Nextflow (PhIP-flow) pipeline.
  • analysis-image-template/phippery/ — VENDORED exact phippery version + requirements.txt.
  • analysis-scripts/*.py — enrichment + plotting (produce every manuscript figure).
  • Manuscript-Figures/SPIKE1, SPIKE2/{pdf,png,tiff} — the PUBLISHED figure outputs (ground truth to match).
  • SRA/sra-metadata.tsv — maps each SRA run/tar → contained per-sample fastq filenames.

Pipeline (PhIP-analysis.nf), each step reproducible 1:1

  1. phippery peptide-md-to-fasta : peptide_table.csv → peptides.fasta (Oligo column).
  2. bowtie2-build peptides.fasta → index.
  3. zcat fastq | bowtie2 --trim3 32 --threads 12 -x peptide - → per-sample SAM (reads 125 bp, trim 32 ⇒ 93 bp coding region; default end-to-end, exact seed = "zero mismatches").
  4. samtools idxstats → per-peptide counts; samtools stats → alignment stats.
  5. phippery collect-phip-data.phip (pickled xarray Dataset: counts × peptides × samples).
  6. layer-enrichment-stats.py → cpm, enrichment (vs library), standardized_enrichment (vs library+beads), differential_selection_wt_mut (scaled_by_wt=True, smoothing_flank_size=1) → layered-analysis.phip.
  7. Per batch (SPIKE1/SPIKE2): pca, heatmap-boxplot, logopairs (scaled diff-sel logos), threshold-epi-binding-barplot (summed WT enrichment per epitope vs binding threshold), haarvi/nih subgroups.

IN SCOPE (pipeline-derived, computational — attempt to reproduce)

  • R1 Peptide library size (24,840 in table vs paper "24,820") — profiling.
  • R2 Sample/cohort counts vs paper (Moderna 49; HAARVI 64; library/beads controls).
  • R3 Alignment statistics (reads mapped / coverage) — alignment-stats.py.
  • R4 Enrichment + standardized enrichment + scaled differential selection (the .phip data layers).
  • R5 Five epitope regions (NTD, CTD/CTD-N, FP, stem helix-HR2) + their binding thresholds (epitopes.py: NTD 40, CTD 200, CTD-N 120, FP 100, SH-H 150).
  • R6 Escape sites (peaks of scaled differential selection), per paper Results: NTD 291,294–297,300–302,304 · CTD 561,562 · FP 819,820,822,823 · SH-H 1148,1152,1155,1156.
  • R7 Regenerate the manuscript figures (epitope_wt_thresholds, logopairs, heatmap-boxplot, pca, haarvi, nih) for SPIKE1 & SPIKE2 and compare to Manuscript-Figures/.

OUT OF SCOPE (not pipeline-derived)

  • Wet-lab: phage library construction, serum/plasma collection, IP, sequencing (experimental).
  • Clinical cohort recruitment / metadata (HAARVI, Moderna trial) — external.
  • "Pre-processed enrichment data available upon request" — NOT deposited; we REGENERATE it from raw fastq (the raw fastq IS deposited on SRA), so this is not a blocker.

Reproduction approach

Manual faithful port of the Nextflow steps (same exact commands/args) via a conda env on «our HPC» «infra» (avoids Nextflow+Singularity+quay-container fragility on «infra»). Heavy alignment of 633 samples runs as a SLURM job; all data stays on «infra». Third-party-tool validity (P16) is moot — this is the authors' o

Figures / tables: Figures
R1
Reported
24,820 designed peptides
Reproduced
24,840 rows (1242 WT + 23598 mut)
did not match
R2a
Reported
Moderna cohort 49 individuals
Reproduced
Cohort 1 = 49 unique participant_ID
exact
R2b
Reported
HAARVI cohort 64 individuals
Reproduced
Cohort 2 = 65 unique participant_ID
partial
R3
Reported
bowtie2 --trim3 32 end-to-end, 125bp
Reproduced
reproduced on 44 recoverable samples; 72-92% align; counts byte-identical to samtools idxstats
within tolerance
R4
Reported
enrichment + std-enrichment + scaled diff-selection
Reproduced
library-relative enrichment + scaled diff-sel on 30 HAARVI samples; standardized(beads) not reproducible
partial
R5
Reported
5 epitope regions; thresholds NTD40 CTD200 CTD-N120 FP100 SH-H150
Reproduced
epitopes.py defines exactly these
exact
R6
Reported
escape sites NTD/CTD/FP/SH-H
Reproduced
HAARVI subset escape in correct regions (CTD 562 exact; others +/-1-3); vaccine escape unreproducible
partial
R7
Reported
manuscript figures SPIKE1/SPIKE2
Reproduced
NOT reproduced (cohorts on non-demultiplexable HiSeq runs)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 57/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

578.7 k
tokens (I/O) · 54.1 M incl. cache
169 min
runtime · 22.68 CPU-h
2.2 GB
peak RAM
1 (1 failed)
HPC jobs
hummel
machine