Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

High-resolution profiling of pathways of escape for SARS-CoV-2 spike-binding antibodies.

Cell · 2021
L1 87/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
87/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 72% of all assessed papers rank 301 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce. Phage-DMS pipeline (phip-flow bowtie align -> phippery enrichment -> differential selection) with a public repo that conveniently SHIPS its small intermediate/result tables. Reproduced 5 of 6 targeted pipeline-derived claims by recomputing from those shipped artifacts WITHOUT the heavy run: C1 (24,820 peptides, EXACT), C5 (bowtie -n 0, 93 bp trim, EXACT), C6 (18 pts, r 0.96-0.27, excl. 7/16/17, EXACT), and C3+C4 (the two immunodominant escape epitopes = fusion peptide 809-834 and linker/HR2 1140-1168, recomputed as the top-2 site peaks of summed positive scaled_diff_sel across all 18 plasmas; coordinate offset confirmed by WT-residue identity to UniProt P0DTC2). No value appears fabricated -- all six match the deposited data. NOT done: (a) C2 library representation 96.0/95.9% (needs control-fastq alignment); (b) the full INDEPENDENT from-raw-fastq run of phip-flow on PRJNA715823 (~3.4 GB; lightweight and feasible) -- both blocked because the «our HPC» account «user» had its per-user storage quota EXHAUSTED on both «infra» and home at run time (touch fails everywhere; FS only 18% full -> quota wall, not capacity; shared by ~6 live rooms + ~858 historical dirs + ~46 GB conda envs). Did not delete anything (would break other live rooms); alerted the operator. C3/C4 grades are 'within-tol' (consistency check against the authors' shipped scaled_diff_sel, not an independent re-derivation from reads). ALL grades PROVISIONAL -- human review required.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 87
    assessed: 2026-06-18 ⛓ 321d90d74120
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

How can SARS-CoV-2 evolve to escape antibody-mediated immune protection? The paper uses Phage-DMS to define the complete profile of single mutations across the spike (S) protein that reduce binding by COVID-19 convalescent plasma antibodies.

Core claims
  • Phage-DMS comprehensively maps the effect of all possible single mutations across the SARS-CoV-2 spike protein on polyclonal plasma antibody binding, defining antibody escape pathways. method
  • Antibody binding to linear epitopes was common in two immunodominant regions: the fusion peptide (FP, aa 809-834) and the linker region upstream of HR2 (aa 1,140-1,168). finding
  • Escape mutations were variable within the immunodominant regions and there was individual person-to-person variation in both epitopes targeted and specific escape profiles. finding
  • Most patients lacked antibodies binding linear RBD peptides despite having RBD-binding antibodies, indicating RBD antibodies target mainly conformational/glycosylated epitopes missed by Phage-DMS. finding
  • A substantial range (0%-59%) of plasma neutralization activity is directed at non-RBD epitopes, with S2-subunit binding correlating with residual neutralization after RBD depletion. finding
  • Antibody escape pathways predicted by Phage-DMS allow forecasting of antibody-mediated virus evolution at the individual level. finding
  • A Spike Phage-DMS library of 24,820 unique 31-aa peptides tiling the S ectodomain (including D614G context) at single-aa resolution. resource
Experimental setups
Assay System Perturbation Readout Platform
Phage-DMS (phage display deep mutational scanning, immunoprecipitation + deep sequencing) T7 phage-displayed peptide library tiling SARS-CoV-2 Wuhan Hu-1 S ectodomain; COVID-19 convalescent patient plasma (HAARVI cohort, n=18) all possible single amino acid mutations in peptides peptide enrichment and scaled differential selection (mutant vs wild-type binding) T7 phage display vector; Illumina deep sequencing
ELISA COVID-19 patient plasma; RBD, Spike, and S2 subunit proteins mock depletion vs RBD-binding antibody depletion plasma antibody binding (area under the curve, AUC)
Pseudovirus neutralization assay pseudotyped lentivirus expressing SARS-CoV-2 S protein; mock-depleted vs RBD-depleted plasma depletion of RBD-binding antibodies neutralization titer 50% (NT50) / residual neutralization
Key results
  • FP region (aa 809-834) and linker/HR2 region (aa 1,140-1,168) were the most enriched immunodominant linear epitopes across patients
  • Patient 5 uniquely showed strong enrichment of an HR2 epitope (aa 1,167-1,191) not seen in the other 14 patients
  • Less common epitopes targeted individually: NTD (aa 255-280) patient 12; RBD (aa 485-500) and downstream of RBD (aa 540-573) patient 15; upstream of S1/S2 cleavage site (aa 620-644) patient 3
  • Paired day 30 and day 60 p.s.o. samples were well correlated in peptide enrichment, significantly better than randomly paired samples median r=0.87 vs 0.29
  • RBD antibody depletion did not affect plasma binding to whole S protein, indicating most S-binding antibodies target non-RBD epitopes
  • A wide range of plasma neutralization activity is directed at non-RBD epitopes 0% to 59%
  • Binding to S2 subunit correlated with NT50 after RBD depletion, while FP/HR2 peptide enrichment did not correlate with residual neutralization
  • Final duplicate libraries contained a high percentage of all unique designed sequences 96.0% and 95.9%
Key statistics
  • count 24,820 unique peptides (peptides designed in the Spike Phage-DMS library, 31 aa long overlapping by 30 aa)
  • correlation median 0.87 (Pearson correlation of peptide enrichment between paired day 30 and day 60 p.s.o. samples)
  • correlation median 0.29 (Pearson correlation of randomly paired samples)
  • pvalue p = 1.2e-09 (Wilcoxon rank sum test, paired vs random sample correlation)
  • correlation 0.96 to 0.27 (range of biological replicate correlation across patient samples)
  • count 96.0% and 95.9% (percent of unique designed sequences present in library 1 and library 2)
  • count 18 patients (COVID-19 patients from HAARVI cohort; 3 excluded (7,16,17) leaving 15)
  • other 0% to 59% (residual neutralization activity directed at non-RBD epitopes after RBD depletion)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study used Phage-DMS (phage display deep mutational scanning) to profile antibody binding and escape across the SARS-CoV-2 spike protein in convalescent plasma from 18 COVID-19 patients sampled at ~day 30 and ~day 60 post-symptom onset. Quality control relied on Pearson's correlation between biological replicate phage-IP experiments, with samples below a threshold excluded. Group comparisons were performed with the Wilcoxon rank-sum test. The main results were reported as scaled differential selection values—a continuous, internally derived effect metric—visualized in heatmaps and line plots, supplemented by ELISA area-under-the-curve (AUC) measurements.

Replicationbiological Sample size18 COVID-19 convalescent patients enrolled; 3 excluded by QC (replicate correlation < 0.5); final n = 15; two phage library biological replicates used throughout; no formal power calculation stated GroupsCOVID-19 convalescent patients (paired day ~30 and ~60 p.s.o.) vs. each other and vs. pre-pandemic controls (for ELISA lower limit of detection); true paired samples vs. randomly paired samples for reproducibility assessment Pairingpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Pearson's correlation coefficient Quality control: correlation of peptide enrichment values between biological replicate Phage-DMS experiments; also used to assess correlation between paired day-30 and day-60 samples and between replicate correlation and within-patient time-point correlation 18 patients initially; 15 included after QC not stated
Wilcoxon rank-sum test (two-sample) Comparison of correlation values between true paired patient samples (day 30 vs. day 60) versus randomly paired samples 15 patients with paired samples (after QC exclusions) not stated
ELISA area under the curve (AUC) — descriptive/comparative Plasma binding to RBD, Spike, and S2 subunit before and after RBD-antibody depletion; comparison of mock-depleted vs. RBD-depleted neutralization titer (NT50) 15 patients (day-30 samples) not stated
Pearson's correlation coefficient Correlation between S2 ELISA AUC and residual NT50 after RBD depletion; correlation between FP/HR2 enrichment and residual NT50 15 patients not stated
Approaches that could also have been used
  • Pearson's correlation was used for quality control of replicate experiments and for all within-patient and between-sample correlation analyses
    Could also: Spearman's rank correlation could also be used for the same QC comparisons — Spearman's does not assume linearity or normality of the enrichment distribution; with highly skewed count-based enrichment data from sequencing, rank-based correlation is often more robust to outlying high-enrichment peptides and does not require the homoscedasticity assumed by Pearson
  • A fixed Pearson's R threshold of 0.5 was used to exclude samples with poor replicate reproducibility
    Could also: A data-adaptive or sensitivity-analysis approach could also be applied, such as repeating analyses at multiple thresholds (e.g., 0.4, 0.5, 0.6) or using a formal outlier test on replicate correlations — Reporting results under multiple thresholds or using a formal criterion would allow readers to assess how much the choice of 0.5 influences downstream findings and which patients are sensitive to that cutoff
  • The Wilcoxon rank-sum test was used for a single comparison (paired vs. randomly paired correlations); no multiplicity correction was applied across the several correlation-based tests reported
    Could also: A permutation test could also be used for the same comparison, and a Bonferroni or Benjamini-Hochberg correction could be applied across the family of correlation tests — Permutation tests make no distributional assumptions and are naturally suited to small-n resampling contexts; an FDR or FWER correction across the reported p-values would help readers gauge the expected number of false positives when several tests are reported
  • ELISA results were reported as AUC summaries and compared descriptively or via Pearson correlation, without a formal inferential test for group differences
    Could also: A paired t-test or Wilcoxon signed-rank test (given the paired before/after-depletion design) could also be applied to formalize the AUC comparisons — Formal tests with the matched structure of the depletion experiment (same patient before and after) would provide a p-value and an effect estimate (mean or median difference) that complement the visual display, particularly useful given n = 15
  • Results across the 24,820-peptide library were summarized with scaled differential selection values without a stated significance threshold or multiplicity correction
    Could also: An empirical null distribution (e.g., derived from synonymous-codon controls or pre-pandemic serum) with an FDR-based significance threshold could also be applied — Defining a data-derived significance cutoff for differential selection would allow explicit calling of 'escape' versus 'no-effect' mutations, reducing reliance on visual inspection of heatmaps and facilitating cross-study comparisons
  • Individual patient profiles were presented qualitatively (heatmaps, representative cases) without a formal summary statistic for between-patient variability
    Could also: An entropy or information-theoretic metric per site, or mixed-effects modeling treating patient as a random effect, could also quantify inter-individual variability in escape profiles — A per-site summary of between-patient heterogeneity (e.g., variance in scaled differential selection across patients) would quantify the degree of individual variation that the paper describes qualitatively, enabling statistical comparison of variability across epitope regions
Software: dms-view (online visualization tool, GitHub) · BioRender.com · Custom computational pipeline (cloning/sequencing alignment, 0-mismatch stringent alignment)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34010620

Paper: Garrett ME et al. "High-resolution profiling of pathways of escape for SARS-CoV-2 spike-binding antibodies." Cell 2021. DOI 10.1016/j.cell.2021.04.045. Code: https://github.com/meghangarrett/Spike-Phage-DMS Data: SRA PRJNA715823 (raw Illumina fastq, Phage-DMS phage-immunoprecipitation seq)

Method = Phage-DMS pipeline (all pipeline-derived results are in scope)

  1. Library design (library-design/): Python scripts generate the 24,820 oligonucleotide peptide tiles for the Spike Phage-DMS library (WT + all single mutants across the S ORF, 31-aa tiles). NO sequencing data needed → cheapest reproducible claim.
  2. Alignment pipeline (alignment-pipeline/): Nextflow phip-flow (Matsen lab) → demultiplex, trim reads to 93 bp, Bowtie end-to-end align with 0 mismatches, SAMtools count per-peptide reads → enrichment matrices → phippery organizes into an xarray dataset. Needs raw fastq (PRJNA715823).
  3. Analysis & plotting (analysis-and-plotting/): Python + R compute enrichment, differential selection (log-fold change mutant vs WT peptide at a locus) and scaled differential selection (× WT enrichment), and the escape profiles / epitope figures.

In scope (pipeline-derived → attempt)

id reported result where pipeline tractability
C1 24,820 unique peptides designed in the library Results/STAR Methods library-design scripts HIGH — no seq data, pure compute
C2 Library representation 96.0% / 95.9% across duplicate libraries Results phip-flow align of library/input controls MED — needs control fastq
C3 Immunodominant escape epitope = fusion peptide aa 809–834 Results/Fig full phip-flow + diff-selection MED-HARD — needs Ab sample fastq
C4 Immunodominant escape epitope = linker/HR2 aa 1,140–1,168 Results/Fig full phip-flow + diff-selection MED-HARD
C5 Bowtie end-to-end, 0 mismatch, reads trimmed to 93 bp (method param) STAR Methods phip-flow config verify config, not a number
C6 18 COVID-19 patients; 3 (pt 7,16,17) excluded on QC corr Results sample-level QC in analysis MED — derivable from counts

Out of scope (not pipeline-derived → not attempted)

  • Wet-lab Phage-DMS immunoprecipitation, ELISA, neutralization assays.
  • dms-view interactive visualization (display only).
  • Any clinical/manual interpretation.

Datasets to profile

  • PRJNA715823 (SRA): raw fastq, Phage-DMS PhIP-seq, Illumina. N samples TBD from SRA Run table (Ab/serum IPs + library/input/beads controls across 2 reps).
Figures / tables: figures
C1
Reported
24,820 unique peptides in the Spike Phage-DMS library
Reproduced
24,820 (master_oligos.txt = 24,821 lines - 1 header; peptide_metadata.csv = 24,840 = 24,820 + 20 controls)
exact
C2
Reported
library representation 96.0% and 95.9% across duplicate libraries
Reproduced
not attempted (requires phip-flow alignment of library-control fastq; storage-quota blocked)
partial
C3
Reported
immunodominant escape epitope = fusion peptide aa 809-834
Reproduced
dominant escape peak at spike 809-834 (summed positive scaled_diff_sel window = 15034; top sites 811/828/824/812/821/809)
within tolerance
C4
Reported
immunodominant escape epitope = linker/HR2 aa 1140-1168
Reproduced
second escape peak at spike 1140-1168 (window = 8811; top sites 1145/1158/1161/1154)
within tolerance
C5
Reported
Bowtie end-to-end, 0 mismatches, reads trimmed to 93 bp; SAMtools counts
Reproduced
bowtie -n 0 -l 93 --trim3 32 (125->93bp) --norc --best --sam; num_mm=0, tile_length=93; counts via samtools idxstats
exact
C6
Reported
18 COVID-19 patients; 3 excluded (pt 7,16,17); replicate correlation 0.96 to 0.27
Reproduced
18 distinct participant_IDs (1-18); range max 0.96 (pt1) min 0.27 (pt17); 3 lowest mean-corr = pt 17(0.28)/16(0.36)/7(0.37)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 87/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

210.5 k
tokens (I/O) · 14.5 M incl. cache
39 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.