High-resolution profiling of pathways of escape for SARS-CoV-2 spike-binding antibodies.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce. Phage-DMS pipeline (phip-flow bowtie align -> phippery enrichment -> differential selection) with a public repo that conveniently SHIPS its small intermediate/result tables. Reproduced 5 of 6 targeted pipeline-derived claims by recomputing from those shipped artifacts WITHOUT the heavy run: C1 (24,820 peptides, EXACT), C5 (bowtie -n 0, 93 bp trim, EXACT), C6 (18 pts, r 0.96-0.27, excl. 7/16/17, EXACT), and C3+C4 (the two immunodominant escape epitopes = fusion peptide 809-834 and linker/HR2 1140-1168, recomputed as the top-2 site peaks of summed positive scaled_diff_sel across all 18 plasmas; coordinate offset confirmed by WT-residue identity to UniProt P0DTC2). No value appears fabricated -- all six match the deposited data. NOT done: (a) C2 library representation 96.0/95.9% (needs control-fastq alignment); (b) the full INDEPENDENT from-raw-fastq run of phip-flow on PRJNA715823 (~3.4 GB; lightweight and feasible) -- both blocked because the «our HPC» account «user» had its per-user storage quota EXHAUSTED on both «infra» and home at run time (touch fails everywhere; FS only 18% full -> quota wall, not capacity; shared by ~6 live rooms + ~858 historical dirs + ~46 GB conda envs). Did not delete anything (would break other live rooms); alerted the operator. C3/C4 grades are 'within-tol' (consistency check against the authors' shipped scaled_diff_sel, not an independent re-derivation from reads). ALL grades PROVISIONAL -- human review required.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 87assessed: 2026-06-18 ⛓ 321d90d74120
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusHow can SARS-CoV-2 evolve to escape antibody-mediated immune protection? The paper uses Phage-DMS to define the complete profile of single mutations across the spike (S) protein that reduce binding by COVID-19 convalescent plasma antibodies.
- ★ Phage-DMS comprehensively maps the effect of all possible single mutations across the SARS-CoV-2 spike protein on polyclonal plasma antibody binding, defining antibody escape pathways. method
- ★ Antibody binding to linear epitopes was common in two immunodominant regions: the fusion peptide (FP, aa 809-834) and the linker region upstream of HR2 (aa 1,140-1,168). finding
- ★ Escape mutations were variable within the immunodominant regions and there was individual person-to-person variation in both epitopes targeted and specific escape profiles. finding
- ★ Most patients lacked antibodies binding linear RBD peptides despite having RBD-binding antibodies, indicating RBD antibodies target mainly conformational/glycosylated epitopes missed by Phage-DMS. finding
- ★ A substantial range (0%-59%) of plasma neutralization activity is directed at non-RBD epitopes, with S2-subunit binding correlating with residual neutralization after RBD depletion. finding
- Antibody escape pathways predicted by Phage-DMS allow forecasting of antibody-mediated virus evolution at the individual level. finding
- A Spike Phage-DMS library of 24,820 unique 31-aa peptides tiling the S ectodomain (including D614G context) at single-aa resolution. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Phage-DMS (phage display deep mutational scanning, immunoprecipitation + deep sequencing) | T7 phage-displayed peptide library tiling SARS-CoV-2 Wuhan Hu-1 S ectodomain; COVID-19 convalescent patient plasma (HAARVI cohort, n=18) | all possible single amino acid mutations in peptides | peptide enrichment and scaled differential selection (mutant vs wild-type binding) | T7 phage display vector; Illumina deep sequencing |
| ELISA | COVID-19 patient plasma; RBD, Spike, and S2 subunit proteins | mock depletion vs RBD-binding antibody depletion | plasma antibody binding (area under the curve, AUC) | — |
| Pseudovirus neutralization assay | pseudotyped lentivirus expressing SARS-CoV-2 S protein; mock-depleted vs RBD-depleted plasma | depletion of RBD-binding antibodies | neutralization titer 50% (NT50) / residual neutralization | — |
- ▲ FP region (aa 809-834) and linker/HR2 region (aa 1,140-1,168) were the most enriched immunodominant linear epitopes across patients
- ▲ Patient 5 uniquely showed strong enrichment of an HR2 epitope (aa 1,167-1,191) not seen in the other 14 patients
- – Less common epitopes targeted individually: NTD (aa 255-280) patient 12; RBD (aa 485-500) and downstream of RBD (aa 540-573) patient 15; upstream of S1/S2 cleavage site (aa 620-644) patient 3
- ▲ Paired day 30 and day 60 p.s.o. samples were well correlated in peptide enrichment, significantly better than randomly paired samples median r=0.87 vs 0.29
- – RBD antibody depletion did not affect plasma binding to whole S protein, indicating most S-binding antibodies target non-RBD epitopes
- – A wide range of plasma neutralization activity is directed at non-RBD epitopes 0% to 59%
- – Binding to S2 subunit correlated with NT50 after RBD depletion, while FP/HR2 peptide enrichment did not correlate with residual neutralization
- – Final duplicate libraries contained a high percentage of all unique designed sequences 96.0% and 95.9%
- count 24,820 unique peptides (peptides designed in the Spike Phage-DMS library, 31 aa long overlapping by 30 aa)
- correlation median 0.87 (Pearson correlation of peptide enrichment between paired day 30 and day 60 p.s.o. samples)
- correlation median 0.29 (Pearson correlation of randomly paired samples)
- pvalue p = 1.2e-09 (Wilcoxon rank sum test, paired vs random sample correlation)
- correlation 0.96 to 0.27 (range of biological replicate correlation across patient samples)
- count 96.0% and 95.9% (percent of unique designed sequences present in library 1 and library 2)
- count 18 patients (COVID-19 patients from HAARVI cohort; 3 excluded (7,16,17) leaving 15)
- other 0% to 59% (residual neutralization activity directed at non-RBD epitopes after RBD depletion)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study used Phage-DMS (phage display deep mutational scanning) to profile antibody binding and escape across the SARS-CoV-2 spike protein in convalescent plasma from 18 COVID-19 patients sampled at ~day 30 and ~day 60 post-symptom onset. Quality control relied on Pearson's correlation between biological replicate phage-IP experiments, with samples below a threshold excluded. Group comparisons were performed with the Wilcoxon rank-sum test. The main results were reported as scaled differential selection values—a continuous, internally derived effect metric—visualized in heatmaps and line plots, supplemented by ELISA area-under-the-curve (AUC) measurements.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson's correlation coefficient | Quality control: correlation of peptide enrichment values between biological replicate Phage-DMS experiments; also used to assess correlation between paired day-30 and day-60 samples and between replicate correlation and within-patient time-point correlation | 18 patients initially; 15 included after QC | not stated |
| Wilcoxon rank-sum test (two-sample) | Comparison of correlation values between true paired patient samples (day 30 vs. day 60) versus randomly paired samples | 15 patients with paired samples (after QC exclusions) | not stated |
| ELISA area under the curve (AUC) — descriptive/comparative | Plasma binding to RBD, Spike, and S2 subunit before and after RBD-antibody depletion; comparison of mock-depleted vs. RBD-depleted neutralization titer (NT50) | 15 patients (day-30 samples) | not stated |
| Pearson's correlation coefficient | Correlation between S2 ELISA AUC and residual NT50 after RBD depletion; correlation between FP/HR2 enrichment and residual NT50 | 15 patients | not stated |
-
Pearson's correlation was used for quality control of replicate experiments and for all within-patient and between-sample correlation analyses↳ Could also: Spearman's rank correlation could also be used for the same QC comparisons — Spearman's does not assume linearity or normality of the enrichment distribution; with highly skewed count-based enrichment data from sequencing, rank-based correlation is often more robust to outlying high-enrichment peptides and does not require the homoscedasticity assumed by Pearson
-
A fixed Pearson's R threshold of 0.5 was used to exclude samples with poor replicate reproducibility↳ Could also: A data-adaptive or sensitivity-analysis approach could also be applied, such as repeating analyses at multiple thresholds (e.g., 0.4, 0.5, 0.6) or using a formal outlier test on replicate correlations — Reporting results under multiple thresholds or using a formal criterion would allow readers to assess how much the choice of 0.5 influences downstream findings and which patients are sensitive to that cutoff
-
The Wilcoxon rank-sum test was used for a single comparison (paired vs. randomly paired correlations); no multiplicity correction was applied across the several correlation-based tests reported↳ Could also: A permutation test could also be used for the same comparison, and a Bonferroni or Benjamini-Hochberg correction could be applied across the family of correlation tests — Permutation tests make no distributional assumptions and are naturally suited to small-n resampling contexts; an FDR or FWER correction across the reported p-values would help readers gauge the expected number of false positives when several tests are reported
-
ELISA results were reported as AUC summaries and compared descriptively or via Pearson correlation, without a formal inferential test for group differences↳ Could also: A paired t-test or Wilcoxon signed-rank test (given the paired before/after-depletion design) could also be applied to formalize the AUC comparisons — Formal tests with the matched structure of the depletion experiment (same patient before and after) would provide a p-value and an effect estimate (mean or median difference) that complement the visual display, particularly useful given n = 15
-
Results across the 24,820-peptide library were summarized with scaled differential selection values without a stated significance threshold or multiplicity correction↳ Could also: An empirical null distribution (e.g., derived from synonymous-codon controls or pre-pandemic serum) with an FDR-based significance threshold could also be applied — Defining a data-derived significance cutoff for differential selection would allow explicit calling of 'escape' versus 'no-effect' mutations, reducing reliance on visual inspection of heatmaps and facilitating cross-study comparisons
-
Individual patient profiles were presented qualitatively (heatmaps, representative cases) without a formal summary statistic for between-patient variability↳ Could also: An entropy or information-theoretic metric per site, or mixed-effects modeling treating patient as a random effect, could also quantify inter-individual variability in escape profiles — A per-site summary of between-patient heterogeneity (e.g., variance in scaled differential selection across patients) would quantify the degree of individual variation that the paper describes qualitatively, enabling statistical comparison of variability across epitope regions
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34010620
Paper: Garrett ME et al. "High-resolution profiling of pathways of escape for SARS-CoV-2 spike-binding antibodies." Cell 2021. DOI 10.1016/j.cell.2021.04.045. Code: https://github.com/meghangarrett/Spike-Phage-DMS Data: SRA PRJNA715823 (raw Illumina fastq, Phage-DMS phage-immunoprecipitation seq)
Method = Phage-DMS pipeline (all pipeline-derived results are in scope)
- Library design (
library-design/): Python scripts generate the 24,820 oligonucleotide peptide tiles for the Spike Phage-DMS library (WT + all single mutants across the S ORF, 31-aa tiles). NO sequencing data needed → cheapest reproducible claim. - Alignment pipeline (
alignment-pipeline/): Nextflow phip-flow (Matsen lab) → demultiplex, trim reads to 93 bp, Bowtie end-to-end align with 0 mismatches, SAMtools count per-peptide reads → enrichment matrices → phippery organizes into an xarray dataset. Needs raw fastq (PRJNA715823). - Analysis & plotting (
analysis-and-plotting/): Python + R compute enrichment, differential selection (log-fold change mutant vs WT peptide at a locus) and scaled differential selection (× WT enrichment), and the escape profiles / epitope figures.
In scope (pipeline-derived → attempt)
| id | reported result | where | pipeline | tractability |
|---|---|---|---|---|
| C1 | 24,820 unique peptides designed in the library | Results/STAR Methods | library-design scripts | HIGH — no seq data, pure compute |
| C2 | Library representation 96.0% / 95.9% across duplicate libraries | Results | phip-flow align of library/input controls | MED — needs control fastq |
| C3 | Immunodominant escape epitope = fusion peptide aa 809–834 | Results/Fig | full phip-flow + diff-selection | MED-HARD — needs Ab sample fastq |
| C4 | Immunodominant escape epitope = linker/HR2 aa 1,140–1,168 | Results/Fig | full phip-flow + diff-selection | MED-HARD |
| C5 | Bowtie end-to-end, 0 mismatch, reads trimmed to 93 bp (method param) | STAR Methods | phip-flow config | verify config, not a number |
| C6 | 18 COVID-19 patients; 3 (pt 7,16,17) excluded on QC corr | Results | sample-level QC in analysis | MED — derivable from counts |
Out of scope (not pipeline-derived → not attempted)
- Wet-lab Phage-DMS immunoprecipitation, ELISA, neutralization assays.
- dms-view interactive visualization (display only).
- Any clinical/manual interpretation.
Datasets to profile
- PRJNA715823 (SRA): raw fastq, Phage-DMS PhIP-seq, Illumina. N samples TBD from SRA Run table (Ab/serum IPs + library/input/beads controls across 2 reps).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.