Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Target identification for repurposed drugs active against SARS-CoV-2 via high-throughput inverse docking.

J Comput Aided Mol Des · 2021
L1 63/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
63/100
Reproducibility score
0.6 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 24% of all assessed papers rank 875 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for the DEPOSITED pipeline. The deterministic HARDs-construction data pipeline (sqlite3 build.sql -> hards.sql -> extract.sql on ChEMBL v27) reproduces 1:1 on «our HPC»: 158 HARDs (C1) and 11 triple-screen compounds (C3) regenerate with identical compound sets/values, and the three per-screen ranked CSVs are byte-for-byte md5-identical (C3b). The ONLY textual difference in the .dat files is sqlite float-repr precision on a derived column (max numeric delta 3.6e-15 = bit-identical IEEE doubles). This is a clean reproduction of the data-mining stage. The paper's HEADLINE inverse-docking results (docking Z-scores, preferential targets, 66/75/81/88% validation recovery, AUC 0.95, Apilimod Z=-2.21) are NOT reproducible from the deposit and were not attempted 1:1: the Vinardo-beta scoring binary is unobtainable (on request) and neither the raw docking scores nor the validation benchmark are deposited — an auditability/availability gap, with no fabrication evidence in the reproducible part. Finding: repo checksums.dat pins a WRONG sha256 for ChEMBL v27 (same hash as v26); actual EBI tarball sha256=5abce60d...; tarball integrity OK against canonical EBI source.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.5557825

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 40
    assessed: 2026-06-18 ⛓ b56d1c52d8a7
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Drugs already shown by consensus of multiple high-throughput screens to have anti-SARS-CoV-2 activity act through specific, currently unknown viral or human protein targets that can be identified using an improved multi-scoring-function inverse-docking (INDO) protocol.

Core claims
  • Combining Vinardo, Ledock, and Korp-PL scoring functions (via averaged Z-scores) improves correct target identification over any single scoring function. method
  • TMPRSS2 and PIKfyve are the most common preferential human targets among the repurposed anti-SARS-CoV-2 drugs tested. finding
  • Helicase and PLpro are the most common preferential viral targets among the repurposed drugs tested. finding
  • All compounds that preferentially select TMPRSS2 are known serine protease inhibitors. finding
  • All compounds that preferentially select PIKfyve are known tyrosine kinase inhibitors. finding
  • The INDO protocol was validated using a curated test set of known drug-target crystallographic complexes, including antiviral and non-antiviral compounds. method
  • A curated 'HARD' list of 152 repurposed drugs (from consensus of three HTS studies) was subjected to inverse docking against 18 SARS-CoV-2 and 6 human protein targets. resource
  • Detailed structural analysis of docking poses reveals molecular interaction patterns explaining why TMPRSS2 and PIKfyve arise as preferential targets. mechanism
Experimental setups
Assay System Perturbation Readout Platform
Inverse docking (multi-scoring-function validation) 209 curated protein-ligand crystallographic complexes (PDBBIND refined 2018 set) none top-1/top-N correct target recovery rate, ROC AUC Vinardo (beta), Ledock, Korp-PL
Inverse docking (HARD list target identification) 152 repurposed drugs vs 18 SARS-CoV-2 proteins and 6 human proteins (65 structures, 35 sites) drug (repurposed compound) docking, no biological perturbation consensus Z-score-ranked preferred protein target per drug Vinardo (beta), Ledock, Korp-PL
Molecular dynamics trajectory analysis Glycosylated SARS-CoV-2 spike (S) protein trimer, closed state none atomic density of glycans at interprotomer interfaces Shaw Research 10 μs MD simulation
Exploratory/site-identification docking Interior cavities of SARS-CoV-2 spike protein 300 randomly selected FDA-approved probe drugs (MW 400-600 Da, logP 0-6) clustering of docked ligand centers of mass to reveal druggable internal sites SuperDrug2 database structures; docking software as above
Homology/structure modeling TMPRSS2, Nsp6, ExoN, M protein (targets lacking experimental structures) none predicted three-dimensional protein structure for use as docking target Swiss-Model, AlphaFold, I-TASSER (Zhang lab)
High-throughput screening (source data, not performed in this study) Cultured SARS-CoV-2-infected cells large compound libraries (FDA/EMA-approved and clinical-trial drugs) anti-SARS-CoV-2 activity used to build the HARD drug list
Key results
  • Combined use of Ledock, Korp-PL, and Vinardo recovered the correct crystallographic target as top-1/top-3/top-5/top-10 in the validation test set 66%, 75%, 81%, 88%
  • ROC analysis showed strong discrimination of true from false target predictions AUC = 0.95 (all predictions), AUC = 0.85 (top-1 only)
  • Among preferential targets selected across the HARD drug list, TMPRSS2 and PIKfyve (human enzymes) were most frequently chosen
  • Helicase and PLpro (viral enzymes) were the next most frequently selected preferential targets after TMPRSS2 and PIKfyve
  • HARD list construction from three independent HTS datasets yielded compounds meeting activity and molecular weight criteria 158 drugs reduced to 152 after MW filter
  • Optimal Z-score cutoffs separating true from false positive target assignments were identified from maximal TPR-FPR difference Z = -0.90 (all predictions), Z = -1.50 (top-1 only)
  • Target set for the INDO study comprised multiple structures/binding sites per protein to capture conformational diversity 65 protein structures representing 35 distinct sites
Key statistics
  • count 209 crystallographic protein-ligand complexes (PDBBIND-derived INDO validation test set)
  • count 43,681 docking calculations (209×209) per docking program (total pairwise dockings performed for test set validation)
  • other top-1 66%, top-3 75%, top-5 81%, top-10 88% (correct target recovery rate using combined Ledock/Korp-PL/Vinardo scoring)
  • other AUC = 0.95 (all predictions); AUC = 0.85 (top-1 predictions) (ROC curve discrimination performance of INDO procedure)
  • other Z-score thresholds of -0.90 and -1.50 (average Z-scores giving maximal TPR-FPR difference)
  • count 152 compounds (from an initial 158-drug HARD list) (final repurposed drug set subjected to INDO after molecular weight filtering)
  • count 18 SARS-CoV-2 proteins and 6 human proteins as targets (protein target set used in the INDO study)
  • count 65 protein structures representing 35 distinct binding sites (structural diversity of targets used for INDO)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational cheminformatics study employing inverse docking (INDO) with three docking programs (Vinardo, Ledock, Korp-PL) whose scores are Z-score-normalized and combined via arithmetic mean to rank 35 binding sites for each of 152 repurposed anti-SARS-CoV-2 compounds. Validation was performed against 209 crystallographic protein–ligand complexes from PDBBIND, with performance reported as top-N recovery fractions and ROC/AUC values. No traditional inferential hypothesis tests were applied; the primary analytical outputs are point-estimate recovery fractions and AUC metrics.

Replicationtechnical Sample sizeTest set: 209 crystallographic complexes (PDBBIND 2018 refined set, filtered by MW 200–700 Da, 3–7 rotatable bonds, affinity ≤ 1 µM, Tanimoto-clustered); HARD list: 152 compounds (top-25% actives in ≥2 of 3 HTS studies); target set: 65 protein structures representing 35 distinct binding sites; up to 5 conformers per ligand GroupsIndividual scoring functions vs. combined scoring functions; combined scoring functions via exponential consensus ranking vs. arithmetic mean of Z-scores; 152 HARD compounds ranked against 35 binding sites Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Z-score normalization with arithmetic-mean consensus of docking scores across three scoring functions (Vinardo, Ledock, Korp-PL) Core INDO ranking pipeline applied to both the 209-complex validation set and the 152-compound HARD list against 65 protein structures 209 (test set); 152 compounds × 65 structures (HARD list) not stated
Top-N recovery rate (fraction of ligands correctly selecting their crystallographic target within top 1, 3, 5, or 10) Validation of INDO protocol against the PDBBIND-derived test set (Fig. S1) 209 na
ROC curve analysis with area under the curve (AUC) Discrimination performance evaluation of the INDO procedure on the PDBBIND test set (Fig. S2); two scenarios: all predictions (AUC = 0.95) and top-1 predictions only (AUC = 0.85) 209 not stated
Tanimoto-coefficient binning clustering (threshold = 0.6) via Chemmine tools Removal of molecular redundancy when constructing the 209-compound PDBBIND test set not stated
K-means / center-of-mass clustering of docked probe poses Identification of six internal binding sites in the SARS-CoV-2 spike protein using 300 randomly selected FDA-approved probe molecules (Figs. S8–S9) 300 not stated
Approaches that could also have been used
  • Top-N recovery fractions and AUC values are reported as single point estimates computed over the full 209-complex test set
    Could also: Bootstrap resampling of the 209-complex test set (e.g., 1000 iterations) could provide 95% confidence intervals for both recovery fractions and AUC values — With n = 209, the binomial 95% CI on a 66% top-1 recovery rate spans roughly ±6 percentage points; reporting these intervals would allow readers to judge whether observed differences between scoring-function combinations exceed sampling variability
  • Arithmetic mean of Z-scores was selected over exponential consensus average ranking by comparing both methods on the same full test set
    Could also: k-fold cross-validation within the 209-complex test set would allow both consensus strategies to be compared with uncertainty estimates on their performance difference — Selecting the better of two methods by evaluating both on the same dataset that informs the choice introduces an optimistic bias in the reported performance of the selected method; cross-validation provides a less biased estimate of expected performance on new data
  • Z-score normalization (mean ± SD) was applied to correct scoring-function bias across proteins
    Could also: Robust Z-scores using the median and median absolute deviation (MAD), or percentile-rank normalization, could also be used — Docking score distributions can be skewed or contain extreme values for flexible or unusually large binding sites; rank-based or MAD-based normalizations are less sensitive to such outliers and may yield more stable rankings
  • A single Tanimoto similarity threshold of 0.6 was used to de-replicate the test set
    Could also: Reporting top-N recovery at multiple thresholds (e.g., 0.4, 0.6, 0.8) or using scaffold-based clustering (Bemis–Murcko) would allow assessment of how the diversity of the test set affects apparent difficulty — The chosen threshold determines both test-set size and chemical diversity, which directly affects the top-N recovery metric; a sensitivity analysis across thresholds would show whether the reported 66–88% recovery range is stable across reasonable de-replication choices
  • The HARD list was defined with a fixed top-25% activity cutoff applied uniformly to each of the three HTS datasets
    Could also: A sensitivity analysis at alternative cutoffs (e.g., top-10%, top-33%) or a continuous activity-weighted scheme could also define the compound set — Reporting how the number of included compounds and the resulting target distribution shift under alternative thresholds would help readers assess whether the prioritization of targets such as TMPRSS2 and PIKfyve is robust to the specific cutoff chosen
  • Three hundred FDA-approved probe molecules used to discover internal spike-protein binding sites were selected randomly from a MW/logP-filtered set
    Could also: A diversity-maximizing algorithm (e.g., MaxMin selection using Morgan fingerprint distances) could also be applied instead of random sampling — Random sampling from a filtered pool may by chance over-represent particular chemotypes; a diversity-maximizing selection would provide more uniform coverage of chemical space for the exploratory binding-site identification step, potentially revealing sites with different physicochemical preferences
Software: Vinardo (beta version) · Ledock · Korp-PL · Gypsum-DL · Chemmine tools · ChEMBL database 27

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34825285

Paper: Ribone SR, Paz SA, Abrams CF, Villarreal MA. "Target identification for repurposed drugs active against SARS-CoV-2 via high-throughput inverse docking." J Comput Aided Mol Des 36, 25–37 (2022). DOI 10.1007/s10822-021-00432-3.

Code: https://github.com/alexispaz/HTIDocking-Cov2 (mirror: zenodo 10.5281/zenodo.5557825) Type: data + scripts repository (Shell/SQL). Authors' own repo (not a 3rd-party tool).

What the paper does (pipeline overview)

  1. HARDs list construction (data pipeline, SQL on ChEMBL v27). Three published high-throughput SARS-CoV-2 drug-repurposing cell screens — Ellinger, Heiser, Touret — are deposited in ChEMBL v27 (doc_id 110246/110267/110230). The repo's hards/ SQL scripts rank compounds within each screen, then keep compounds that rank in the top fraction of ≥2 screens ("HARDs"). Output = the candidate drug set that is subsequently docked. THIS IS FULLY DEPOSITED AND DETERMINISTIC.
  2. Inverse docking (compute pipeline). Each HARD drug is docked against 65 protein structures (24 targets / 35 binding sites: 18 viral + 6 human). Docking with Vinardo (beta) and Ledock; top-10 Vinardo poses re-scored with KORP-PL. Ligands prepared with Gypsum-DL. Per-program scores → Z-score normalised per target → averaged → targets with mean Z ≤ −1.0 reported as preferential ("top-1") targets.
  3. Validation (compute). Self-docking benchmark on 209 PDBbind crystallographic complexes (top-N native-target recovery), plus 14 known SARS-CoV-2 inhibitor controls.

In scope (attempted)

# Result Pipeline Reproducibility
S1 HARDs list = 158 compounds (active in ≥2 screens, rank cut<25 & ave<25) hards.shbuild.sql+hards.sql+extract.sql (sqlite3 on ChEMBL v27) HIGH — deterministic SQL, ChEMBL v27 checksum-pinned, reference outputs shipped in output_27/. Compare row-for-row.
S2 11 compounds active in all 3 screens (hards3.dat) same HIGH — same pipeline
S3 Per-screen ranked compound counts (Ellinger/Heiser/Touret CSVs) same HIGH
S4 Docking sanity: re-dock ≥1 control complex (e.g. Apilimod→PIKfyve, GRL-0617→PLpro) smina --scoring vinardo on deposited structures MEDIUM — smina reimplements Vinardo; box specs from Suppl. Table S6

Out of scope / not attempted (with reason)

  • Full inverse-docking Z-score table & preferential-target counts (TMPRSS2 7, PIKfyve 6, Helicase 5, …; Apilimod Z=−2.21). BLOCKERS: (a) Vinardo "beta version" is "available upon request" — not publicly obtainable; (b) the raw docking scores are NOT deposited (repo ships only inputs: structures + drug list, no docking output matrix), so the Z-scores cannot be regenerated 1:1 without re-running the full ~10k×3 docking campaign with the exact (unobtainable) scoring binary and exact box definitions. Partial sanity-docking with smina-vinardo is attempted (S4) but is a different tool, so it cannot reproduce the exact numbers — only test plausibility.
  • Validation top-N recovery (66/75/81/88 %) and AUC 0.95. Requires the 209-complex PDBbind self-docking benchmark with the same (unobtainable) Vinardo-beta + KORP-PL stack. Not deposited as a result file. Not attempted 1:1.
  • Wet-lab / external screen data generation (Ellinger/Heiser/Touret experiments) — external, out of scope.

Datasets profiled (same pass)

  • D1 GitHub/Zenodo deposit HTIDocking-Cov2: structures/ (PDB) + hards/output_27/ (HARDs outputs).
  • D2 ChEMBL v27 SQLite (upstream dependency; checksum 6f23acbf… pinned in hards.sh).
  • D3 The 3 source screens (Ellinger/Heiser/Touret) as embedded ChEMBL documents.

Hard-compute plan («our HPC»/«infra» only)

  • Front1: download chembl_27_sqlite.tar.gz (~3.3 GB) into «infra» work dir, verify checksum, extract.
  • Run build.sql/hards.sql/extract.sql via sqlite3 («infra»); compare regener
Figures / tables: Table
C1
Reported
158 HARDs (compounds in top fraction of >=2 of 3 screens)
Reproduced
158 HARDs, identical compound set/order/values
exact
C2
Reported
152 drugs docked after MW/prep filter
Reproduced
not reproduced (MW/Gypsum-DL prep step not in deposited code; upstream 158 reproduced)
partial
C3
Reported
11 compounds active in all 3 screens
Reproduced
11 compounds, identical (THIOGUANINE, QUINIDINE, VALDECOXIB, NITAZOXANIDE, EPERISONE, CEFACLOR, IMIQUIMOD, PYRIMETHAMINE, RUFLOXACIN, PIDOTIMOD, DROFENINE)
exact
C3b
Reported
per-screen ranked CSVs (Ellinger/Heiser/Touret)
Reproduced
byte-for-byte md5-identical
exact
C4
Reported
65 structures / 35 sites / 24 targets
Reproduced
49 PDB files deposited (multi-site files; S6 box defs absent)
partial
C5
Reported
validation top-1 recovery 66%
Reproduced
not attempted
partial
C6
Reported
validation top-3/5/10 75/81/88%
Reproduced
not attempted
partial
C7
Reported
AUC 0.95 / 0.85
Reproduced
not attempted
partial
C8
Reported
control inhibitors 3/14, 10/14, 12/14
Reproduced
not attempted
partial
C9
Reported
52 preferential hits (Z<=-1.0)
Reproduced
not attempted
partial
C10
Reported
target freq TMPRSS2 7/PIKfyve 6/Helicase 5/S4 5/PLpro 4
Reproduced
not attempted
partial
C11
Reported
Apilimod->PIKfyve avg Z -2.21
Reproduced
not attempted
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 63/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is the authors' own data+SQL repository, and the HARDs-selection pipeline (C1=158, C3=11) is fully deposited, deterministic on ChEMBL v27, and being reproduced 1:1 (pending). The inverse-docking core (validation recovery 66/75/81/88%, AUC 0.95/0.85, 52 preferential HARDs, Apilimod -2.21) is not reproducible: the 'Vinardo beta' engine is request-only and the raw docking scores were never deposited, and the structure deposit is short (49 of 65). The problem sits on the authors'/data-availability side (non-deposited scores + unobtainable tool + under-deposit), not on a computed discrepancy — so the central claims are unverified rather than disproven. Severity is moderate and bounded by availability, hence a yellow/partial overall rather than a fabrication-suspect red.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

72.2 k
tokens (I/O) · 3.2 M incl. cache
28 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.