Target identification for repurposed drugs active against SARS-CoV-2 via high-throughput inverse docking.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for the DEPOSITED pipeline. The deterministic HARDs-construction data pipeline (sqlite3 build.sql -> hards.sql -> extract.sql on ChEMBL v27) reproduces 1:1 on «our HPC»: 158 HARDs (C1) and 11 triple-screen compounds (C3) regenerate with identical compound sets/values, and the three per-screen ranked CSVs are byte-for-byte md5-identical (C3b). The ONLY textual difference in the .dat files is sqlite float-repr precision on a derived column (max numeric delta 3.6e-15 = bit-identical IEEE doubles). This is a clean reproduction of the data-mining stage. The paper's HEADLINE inverse-docking results (docking Z-scores, preferential targets, 66/75/81/88% validation recovery, AUC 0.95, Apilimod Z=-2.21) are NOT reproducible from the deposit and were not attempted 1:1: the Vinardo-beta scoring binary is unobtainable (on request) and neither the raw docking scores nor the validation benchmark are deposited — an auditability/availability gap, with no fabrication evidence in the reproducible part. Finding: repo checksums.dat pins a WRONG sha256 for ChEMBL v27 (same hash as v26); actual EBI tarball sha256=5abce60d...; tarball integrity OK against canonical EBI source.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 40assessed: 2026-06-18 ⛓ b56d1c52d8a7
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetDrugs already shown by consensus of multiple high-throughput screens to have anti-SARS-CoV-2 activity act through specific, currently unknown viral or human protein targets that can be identified using an improved multi-scoring-function inverse-docking (INDO) protocol.
- ★ Combining Vinardo, Ledock, and Korp-PL scoring functions (via averaged Z-scores) improves correct target identification over any single scoring function. method
- ★ TMPRSS2 and PIKfyve are the most common preferential human targets among the repurposed anti-SARS-CoV-2 drugs tested. finding
- ★ Helicase and PLpro are the most common preferential viral targets among the repurposed drugs tested. finding
- ★ All compounds that preferentially select TMPRSS2 are known serine protease inhibitors. finding
- ★ All compounds that preferentially select PIKfyve are known tyrosine kinase inhibitors. finding
- ★ The INDO protocol was validated using a curated test set of known drug-target crystallographic complexes, including antiviral and non-antiviral compounds. method
- ★ A curated 'HARD' list of 152 repurposed drugs (from consensus of three HTS studies) was subjected to inverse docking against 18 SARS-CoV-2 and 6 human protein targets. resource
- Detailed structural analysis of docking poses reveals molecular interaction patterns explaining why TMPRSS2 and PIKfyve arise as preferential targets. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Inverse docking (multi-scoring-function validation) | 209 curated protein-ligand crystallographic complexes (PDBBIND refined 2018 set) | none | top-1/top-N correct target recovery rate, ROC AUC | Vinardo (beta), Ledock, Korp-PL |
| Inverse docking (HARD list target identification) | 152 repurposed drugs vs 18 SARS-CoV-2 proteins and 6 human proteins (65 structures, 35 sites) | drug (repurposed compound) docking, no biological perturbation | consensus Z-score-ranked preferred protein target per drug | Vinardo (beta), Ledock, Korp-PL |
| Molecular dynamics trajectory analysis | Glycosylated SARS-CoV-2 spike (S) protein trimer, closed state | none | atomic density of glycans at interprotomer interfaces | Shaw Research 10 μs MD simulation |
| Exploratory/site-identification docking | Interior cavities of SARS-CoV-2 spike protein | 300 randomly selected FDA-approved probe drugs (MW 400-600 Da, logP 0-6) | clustering of docked ligand centers of mass to reveal druggable internal sites | SuperDrug2 database structures; docking software as above |
| Homology/structure modeling | TMPRSS2, Nsp6, ExoN, M protein (targets lacking experimental structures) | none | predicted three-dimensional protein structure for use as docking target | Swiss-Model, AlphaFold, I-TASSER (Zhang lab) |
| High-throughput screening (source data, not performed in this study) | Cultured SARS-CoV-2-infected cells | large compound libraries (FDA/EMA-approved and clinical-trial drugs) | anti-SARS-CoV-2 activity used to build the HARD drug list | — |
- ▲ Combined use of Ledock, Korp-PL, and Vinardo recovered the correct crystallographic target as top-1/top-3/top-5/top-10 in the validation test set 66%, 75%, 81%, 88%
- – ROC analysis showed strong discrimination of true from false target predictions AUC = 0.95 (all predictions), AUC = 0.85 (top-1 only)
- – Among preferential targets selected across the HARD drug list, TMPRSS2 and PIKfyve (human enzymes) were most frequently chosen
- – Helicase and PLpro (viral enzymes) were the next most frequently selected preferential targets after TMPRSS2 and PIKfyve
- – HARD list construction from three independent HTS datasets yielded compounds meeting activity and molecular weight criteria 158 drugs reduced to 152 after MW filter
- – Optimal Z-score cutoffs separating true from false positive target assignments were identified from maximal TPR-FPR difference Z = -0.90 (all predictions), Z = -1.50 (top-1 only)
- – Target set for the INDO study comprised multiple structures/binding sites per protein to capture conformational diversity 65 protein structures representing 35 distinct sites
- count 209 crystallographic protein-ligand complexes (PDBBIND-derived INDO validation test set)
- count 43,681 docking calculations (209×209) per docking program (total pairwise dockings performed for test set validation)
- other top-1 66%, top-3 75%, top-5 81%, top-10 88% (correct target recovery rate using combined Ledock/Korp-PL/Vinardo scoring)
- other AUC = 0.95 (all predictions); AUC = 0.85 (top-1 predictions) (ROC curve discrimination performance of INDO procedure)
- other Z-score thresholds of -0.90 and -1.50 (average Z-scores giving maximal TPR-FPR difference)
- count 152 compounds (from an initial 158-drug HARD list) (final repurposed drug set subjected to INDO after molecular weight filtering)
- count 18 SARS-CoV-2 proteins and 6 human proteins as targets (protein target set used in the INDO study)
- count 65 protein structures representing 35 distinct binding sites (structural diversity of targets used for INDO)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a computational cheminformatics study employing inverse docking (INDO) with three docking programs (Vinardo, Ledock, Korp-PL) whose scores are Z-score-normalized and combined via arithmetic mean to rank 35 binding sites for each of 152 repurposed anti-SARS-CoV-2 compounds. Validation was performed against 209 crystallographic protein–ligand complexes from PDBBIND, with performance reported as top-N recovery fractions and ROC/AUC values. No traditional inferential hypothesis tests were applied; the primary analytical outputs are point-estimate recovery fractions and AUC metrics.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Z-score normalization with arithmetic-mean consensus of docking scores across three scoring functions (Vinardo, Ledock, Korp-PL) | Core INDO ranking pipeline applied to both the 209-complex validation set and the 152-compound HARD list against 65 protein structures | 209 (test set); 152 compounds × 65 structures (HARD list) | not stated |
| Top-N recovery rate (fraction of ligands correctly selecting their crystallographic target within top 1, 3, 5, or 10) | Validation of INDO protocol against the PDBBIND-derived test set (Fig. S1) | 209 | na |
| ROC curve analysis with area under the curve (AUC) | Discrimination performance evaluation of the INDO procedure on the PDBBIND test set (Fig. S2); two scenarios: all predictions (AUC = 0.95) and top-1 predictions only (AUC = 0.85) | 209 | not stated |
| Tanimoto-coefficient binning clustering (threshold = 0.6) via Chemmine tools | Removal of molecular redundancy when constructing the 209-compound PDBBIND test set | — | not stated |
| K-means / center-of-mass clustering of docked probe poses | Identification of six internal binding sites in the SARS-CoV-2 spike protein using 300 randomly selected FDA-approved probe molecules (Figs. S8–S9) | 300 | not stated |
-
Top-N recovery fractions and AUC values are reported as single point estimates computed over the full 209-complex test set↳ Could also: Bootstrap resampling of the 209-complex test set (e.g., 1000 iterations) could provide 95% confidence intervals for both recovery fractions and AUC values — With n = 209, the binomial 95% CI on a 66% top-1 recovery rate spans roughly ±6 percentage points; reporting these intervals would allow readers to judge whether observed differences between scoring-function combinations exceed sampling variability
-
Arithmetic mean of Z-scores was selected over exponential consensus average ranking by comparing both methods on the same full test set↳ Could also: k-fold cross-validation within the 209-complex test set would allow both consensus strategies to be compared with uncertainty estimates on their performance difference — Selecting the better of two methods by evaluating both on the same dataset that informs the choice introduces an optimistic bias in the reported performance of the selected method; cross-validation provides a less biased estimate of expected performance on new data
-
Z-score normalization (mean ± SD) was applied to correct scoring-function bias across proteins↳ Could also: Robust Z-scores using the median and median absolute deviation (MAD), or percentile-rank normalization, could also be used — Docking score distributions can be skewed or contain extreme values for flexible or unusually large binding sites; rank-based or MAD-based normalizations are less sensitive to such outliers and may yield more stable rankings
-
A single Tanimoto similarity threshold of 0.6 was used to de-replicate the test set↳ Could also: Reporting top-N recovery at multiple thresholds (e.g., 0.4, 0.6, 0.8) or using scaffold-based clustering (Bemis–Murcko) would allow assessment of how the diversity of the test set affects apparent difficulty — The chosen threshold determines both test-set size and chemical diversity, which directly affects the top-N recovery metric; a sensitivity analysis across thresholds would show whether the reported 66–88% recovery range is stable across reasonable de-replication choices
-
The HARD list was defined with a fixed top-25% activity cutoff applied uniformly to each of the three HTS datasets↳ Could also: A sensitivity analysis at alternative cutoffs (e.g., top-10%, top-33%) or a continuous activity-weighted scheme could also define the compound set — Reporting how the number of included compounds and the resulting target distribution shift under alternative thresholds would help readers assess whether the prioritization of targets such as TMPRSS2 and PIKfyve is robust to the specific cutoff chosen
-
Three hundred FDA-approved probe molecules used to discover internal spike-protein binding sites were selected randomly from a MW/logP-filtered set↳ Could also: A diversity-maximizing algorithm (e.g., MaxMin selection using Morgan fingerprint distances) could also be applied instead of random sampling — Random sampling from a filtered pool may by chance over-represent particular chemotypes; a diversity-maximizing selection would provide more uniform coverage of chemical space for the exploratory binding-site identification step, potentially revealing sites with different physicochemical preferences
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34825285
Paper: Ribone SR, Paz SA, Abrams CF, Villarreal MA. "Target identification for repurposed drugs active against SARS-CoV-2 via high-throughput inverse docking." J Comput Aided Mol Des 36, 25–37 (2022). DOI 10.1007/s10822-021-00432-3.
Code: https://github.com/alexispaz/HTIDocking-Cov2 (mirror: zenodo 10.5281/zenodo.5557825) Type: data + scripts repository (Shell/SQL). Authors' own repo (not a 3rd-party tool).
What the paper does (pipeline overview)
- HARDs list construction (data pipeline, SQL on ChEMBL v27). Three published
high-throughput SARS-CoV-2 drug-repurposing cell screens — Ellinger, Heiser,
Touret — are deposited in ChEMBL v27 (doc_id 110246/110267/110230). The repo's
hards/SQL scripts rank compounds within each screen, then keep compounds that rank in the top fraction of ≥2 screens ("HARDs"). Output = the candidate drug set that is subsequently docked. THIS IS FULLY DEPOSITED AND DETERMINISTIC. - Inverse docking (compute pipeline). Each HARD drug is docked against 65 protein structures (24 targets / 35 binding sites: 18 viral + 6 human). Docking with Vinardo (beta) and Ledock; top-10 Vinardo poses re-scored with KORP-PL. Ligands prepared with Gypsum-DL. Per-program scores → Z-score normalised per target → averaged → targets with mean Z ≤ −1.0 reported as preferential ("top-1") targets.
- Validation (compute). Self-docking benchmark on 209 PDBbind crystallographic complexes (top-N native-target recovery), plus 14 known SARS-CoV-2 inhibitor controls.
In scope (attempted)
| # | Result | Pipeline | Reproducibility |
|---|---|---|---|
| S1 | HARDs list = 158 compounds (active in ≥2 screens, rank cut<25 & ave<25) | hards.sh → build.sql+hards.sql+extract.sql (sqlite3 on ChEMBL v27) |
HIGH — deterministic SQL, ChEMBL v27 checksum-pinned, reference outputs shipped in output_27/. Compare row-for-row. |
| S2 | 11 compounds active in all 3 screens (hards3.dat) |
same | HIGH — same pipeline |
| S3 | Per-screen ranked compound counts (Ellinger/Heiser/Touret CSVs) | same | HIGH |
| S4 | Docking sanity: re-dock ≥1 control complex (e.g. Apilimod→PIKfyve, GRL-0617→PLpro) | smina --scoring vinardo on deposited structures |
MEDIUM — smina reimplements Vinardo; box specs from Suppl. Table S6 |
Out of scope / not attempted (with reason)
- Full inverse-docking Z-score table & preferential-target counts (TMPRSS2 7, PIKfyve 6, Helicase 5, …; Apilimod Z=−2.21). BLOCKERS: (a) Vinardo "beta version" is "available upon request" — not publicly obtainable; (b) the raw docking scores are NOT deposited (repo ships only inputs: structures + drug list, no docking output matrix), so the Z-scores cannot be regenerated 1:1 without re-running the full ~10k×3 docking campaign with the exact (unobtainable) scoring binary and exact box definitions. Partial sanity-docking with smina-vinardo is attempted (S4) but is a different tool, so it cannot reproduce the exact numbers — only test plausibility.
- Validation top-N recovery (66/75/81/88 %) and AUC 0.95. Requires the 209-complex PDBbind self-docking benchmark with the same (unobtainable) Vinardo-beta + KORP-PL stack. Not deposited as a result file. Not attempted 1:1.
- Wet-lab / external screen data generation (Ellinger/Heiser/Touret experiments) — external, out of scope.
Datasets profiled (same pass)
- D1 GitHub/Zenodo deposit
HTIDocking-Cov2:structures/(PDB) +hards/output_27/(HARDs outputs). - D2 ChEMBL v27 SQLite (upstream dependency; checksum 6f23acbf… pinned in
hards.sh). - D3 The 3 source screens (Ellinger/Heiser/Touret) as embedded ChEMBL documents.
Hard-compute plan («our HPC»/«infra» only)
- Front1: download
chembl_27_sqlite.tar.gz(~3.3 GB) into «infra» work dir, verify checksum, extract. - Run
build.sql/hards.sql/extract.sqlvia sqlite3 («infra»); compare regener
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is the authors' own data+SQL repository, and the HARDs-selection pipeline (C1=158, C3=11) is fully deposited, deterministic on ChEMBL v27, and being reproduced 1:1 (pending). The inverse-docking core (validation recovery 66/75/81/88%, AUC 0.95/0.85, 52 preferential HARDs, Apilimod -2.21) is not reproducible: the 'Vinardo beta' engine is request-only and the raw docking scores were never deposited, and the structure deposit is short (49 of 65). The problem sits on the authors'/data-availability side (non-deposited scores + unobtainable tool + under-deposit), not on a computed discrepancy — so the central claims are unverified rather than disproven. Severity is moderate and bounded by availability, hence a yellow/partial overall rather than a fabrication-suspect red.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.