Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

The electrostatic profile of consecutive Cβ atoms applied to protein structure quality assessment.

F1000Res · 2013
L1 42/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🔴The deviation was non-trivial in magnitude
  • 🔴The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
How its reproducibility compares
42/100
Reproducibility score
1.8 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 4% of all assessed papers rank 1120 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

FINAL: Two independent decoy-set comparisons were run end-to-end via an independent pdb2pqr30+APBS3.4.1 electrostatics pipeline and a from-scratch Python reimplementation of escapist.pl's CB-based PDscore, verified line-by-line against the original Perl. Result is a genuine PARTIAL reproduction overall -- not 1:1, not a full mismatch, with real evidence on both sides. Table 4 (misfold decoy set, 49 structures, 26 correct/incorrect pairings): the paper's specific named claim ('barring 1CBH, 1FDX, 2SSI, incorrect structures score worse than correct') reproduces exactly in direction for all 3 named cases. But the implied aggregate specificity does NOT reproduce (ours: 16/26=61.5% vs paper's implied ~23/26=88%), 7 additional pairings fail that the paper doesn't name, and absolute PDscore magnitudes run systematically ~1.7-1.9x lower than the paper's reported values for the 3 named cases (direction/ranking preserved, scale not) -- cause not isolated (candidate: PDB2PQR/APBS version differences vs. the original 2013-era toolchain). Table 5A (hg_structal decoy set, 29 globin-family proteins, full 29x29 self-inclusive threading matrix = 841 decoys): after self-discovering and fixing a case-sensitivity bug in our own batch driver (initial run silently processed only 14/29 proteins' own decoys; fixed via «job», confirmed complete and bug-free at 870/870 lines, 0 pipeline failures), the corrected result is a clear MISMATCH against the paper -- overall pooled specificity 0.454 vs the paper's reported 0.91, and the specific, directly-checkable 1ASH worked example (paper's own Table 5A: specificity=1; ESCAPIST manual's worked example: native PDscore19.9, cohort minimum) reproduces as specificity=0.621, native PDscore=15.2, and the native is explicitly NOT the minimum in its own 29-decoy cohort (minimum=13.3). Dataset profiling found the paper's formal 'Data' deposit (zenodo:10.5281/zenodo.7134) is actually a code/manual archive identical in substance to the GitHub repo, not a distinct experimental dataset -- the real structural data driving every result comes from the third-party Decoys'R'Us database (dd.compbio.org), which the paper never formally cites as a data accession; that dataset's own hg_structal deposit is itself complete and well-formed (870/870 structures verified), so the Table 5A mismatch traces to the scoring pipeline/comparison, not to missing or malformed decoy data. Per the '80/20 is a floor' mandate, this pass covered both Table 4 and Table 5A rather than stopping at Table 4 alone. NOT attempted: Tables 5B (4state_reduced), 5C (fisa), and Supplementary S2 (ig_structal) -- all substantially larger (thousands of structures each), and further extension was also blocked at the end of this pass by a fleet-wide HPC home-directory quota outage (a foreign infra issue affecting all reproduction rooms on this cluster, not specific to this task, and explicitly not grounds to mark this reproduction as failed) that prevented new hummel_submit real-compute jobs; also not attempted: independent re-derivation of the PISCES training statistics behind pd.CB.score.full (reused as shipped).

💻 Code ↗ 🗄 Data: 10.5281/zenodo.7134

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-31
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper tests whether the electrostatic potential difference (EPD) between Cβ atoms of consecutive residues constitutes an amino-acid-type-specific signature that can serve as a statistical potential for model quality assessment; based on the Boltzmann hypothesis and Anfinsen's thermodynamic hypothesis, the authors hypothesize that deviation of observed EPD values from learned mean values is minimized in the native structure.

Core claims
  • The EPD between Cβ atoms of consecutive residues provides unique signatures of amino acid pair types and can discriminate native from decoy protein structures. finding
  • ESCAPIST, a novel knowledge-based MQAP scoring function (PDScore) derived from Cβ EPD statistics of 1000 non-homologous PISCES structures, ranks native structures better than decoys. method
  • The EPD of the C-N peptide bond and the EPD between consecutive Cα atoms are independent of amino acid type and therefore non-discriminatory, because interatomic distances for these atoms are nearly invariant across structures. finding
  • Amino acid pairs containing cysteine have anomalously high variance in Cβ EPD (SD ~90) and cannot be used for discrimination. finding
  • Proline-containing pairs are outliers in Cα EPD with higher absolute mean values, highlighting the unique nature of proline in protein structures. finding
  • The Cβ EPD signal is non-trivial because the variance of EPD between residue pairs increases with increasing sequence separation, i.e., correlation is specific to consecutive residues. mechanism
  • The electrostatic discriminator fails on the fisa decoy set, which consists of physically nonviable structures for which electrostatic analysis is meaningless (these are better discriminated by a distance-based criterion). finding
  • Source code and manual are publicly available at https://github.com/sanchak/mqap and archived at doi 10.5281/zenodo.7134. resource
Experimental setups
Assay System Perturbation Readout Platform
Finite-difference Poisson-Boltzmann electrostatic potential calculation (EPD of Cβ atoms in consecutive residues) 1000 non-homologous protein crystal structures from the PISCES database (20% sequence identity cutoff, 1.6 Å resolution cutoff, R-factor cutoff 0.25) none Mean and standard deviation of electrostatic potential difference (kT/e units) per amino acid pair; number of samples per pair APBS (Adaptive Poisson-Boltzmann Solver) with PDB2PQR
Electrostatic potential difference profiling of the C-N peptide bond Same set of 1000 non-homologous PISCES protein structures none Probability distribution, mean and SD of EPD across all amino acid pairs APBS / PDB2PQR
Electrostatic potential difference profiling between Cα atoms of consecutive residues Same set of 1000 non-homologous PISCES protein structures none Probability distribution, mean and SD of EPD per amino acid pair (proline pairs analyzed separately) APBS / PDB2PQR
Sequence-distance dependence analysis of Cβ EPD PISCES structure set, amino acid pairs 'DF' (Asp/Phe) and 'HS' (His/Ser), ≥30 samples per sequence distance increasing sequence separation between residues (distance 1 to larger) Standard deviation of EPD as a function of sequence distance
Decoy discrimination benchmark (PDScore) — misfold set ~20 protein structures each with a correct and an incorrect (misfolded) structure, from Decoys 'R' Us deliberately misfolded models PDScore of correct vs incorrect structure; per-PDB specificity (0/1) Decoys 'R' Us database (http://dd.compbio.washington.edu)
Decoy discrimination benchmark (PDScore) — hg_structal set ~30 globin-family proteins, 30 structures each (including native) computationally generated decoys Specificity of native-structure identification per protein and averaged Decoys 'R' Us database
Decoy discrimination benchmark (PDScore) — 4state_reduced set 7 proteins (1CTF, 1R69, 1SN3, 2CRO, 3ICB, 4PTI, 4RXN), ~600–690 structures each computationally generated decoys Specificity of native-structure identification per protein and averaged Decoys 'R' Us database
Decoy discrimination benchmark (PDScore) — ig_structal and fisa sets ig_structal immunoglobulin decoy set; fisa set of 4 proteins (1FC2, 1HDDC, 2CRO, 4ICB) with 501 structures each computationally generated decoys Specificity of native-structure identification Decoys 'R' Us database
Key results
  • Average specificity on the hg_structal, 4state_reduced and ig_structal decoy sets 0.91, 0.94 and 0.93 respectively
  • In the misfold decoy set, the incorrect structure had a higher PDScore than the correct structure for all but three proteins (1CBH, 1FDX, 2SSI) 20 of 23 proteins correctly discriminated
  • EPD of the C-N peptide bond has a Gaussian distribution with identical means for all amino acid pairs and large variance, so amino acids are indistinguishable by this feature mean = 420 EPD units, SD = 55 EPD units
  • Mean EPD values between Cβ atoms of consecutive residues are much more varied across amino acid pairs than for Cα atoms or the C-N bond, with lower variance typical SD ~35 EPD units for non-cysteine Cβ pairs
  • Cysteine-containing Cβ pairs show random means and high standard deviation, precluding discrimination (exception: 'CC' pair with low mean and SD) SD ~90 vs ~35 for other pairs; CC: mean -7.1, SD 30.4
  • Proline-containing Cα pairs are outliers with higher absolute mean EPD while SD magnitude is unchanged e.g. DP mean -184.5 (SD 29.5); PS mean 172.3 (SD 25.9)
  • SD of Cβ EPD increases with sequence separation, showing loss of correlation beyond consecutive residues from 29.8 ('DF') and 31.8 ('HS') EPD units at distance 1 to ~60 EPD units
  • Low specificities obtained on the fisa decoy set 4ICB 0, 1HDDC 0.1, 1FC2 0.4, 2CRO 0.7
Key statistics
  • mean mean = 420 EPD units, SD = 55 EPD units (EPD of the C-N peptide bond, Gaussian distribution across all structures)
  • other specificity 0.91 (Average specificity on hg_structal decoy set (~30 proteins x 30 structures))
  • other specificity 0.94 (Average specificity on 4state_reduced decoy set (~600 structures per protein))
  • other specificity 0.93 (Specificity on ig_structal decoy set)
  • count 1000 proteins; 20% identity cutoff, 1.6 Å resolution cutoff, 0.25 R-factor cutoff (PISCES learning set used to derive mean EPD values)
  • mean DF: -108.9 (SD 29.5, n=481); HS: 93.7 (SD 31.8, n=235); HT: 95.1 (SD 31.5, n=235) (Representative discriminatory Cβ consecutive-pair EPD values (Table 3))
  • mean CS: 106.3 (SD 98.1, n=184); CT: 109.9 (SD 97.5, n=173); CM: 61.9 (SD 100, n=39) (Cysteine-containing Cβ pairs with high SD (Table 2))
  • other SD rises from 29.8 (DF) and 31.8 (HS) at sequence distance 1 to around 60 EPD units (Sequence-distance dependence of Cβ EPD variance; at least 30 samples per distance)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational/structural-bioinformatics study rather than a classical wet-lab experiment: electrostatic potential differences (EPD) between backbone atoms were computed from a training set of 1000 non-homologous PISCES structures using APBS, summarized per amino-acid pair as mean, standard deviation, and sample count, and used to build a knowledge-based scoring function (PDScore). The scoring function was then evaluated by computing 'specificity' (the fraction of cases where the native/correct structure scored better than incorrect/decoy structures) across several standard decoy-set benchmarks (misfold, hg_structal, 4state_reduced, ig_structal, fisa). Results are reported as descriptive summary statistics (mean, SD, n) in tables and as aggregate specificity values per decoy set, without formal inferential hypothesis tests.

Replicationunclear Sample sizeTraining set described as 1000 non-homologous protein structures (PISCES, ≤20% identity, ≤1.6 Å resolution, R-factor ≤0.25); decoy-set sizes stated per benchmark (e.g., ~20 proteins in misfold set, 7 proteins x ~600 structures in 4state_reduced); no power calculation described Groupsnative/correct protein structures vs incorrect/decoy structures, scored via PDScore across multiple decoy-set benchmarks Pairingmixed Randomization/blindingna DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Descriptive summary statistics (mean, SD, sample size) per amino-acid pair, no formal hypothesis test reported Tables 1-3, Figures 1-4 (EPD distributions for C-N bond, Cα, and Cβ atom pairs) varies by pair, as stated per row in Tables 1-3 (e.g., n=36 to n=688) stated (Gaussian/Boltzmann distribution assumed for the underlying statistical potential)
Specificity (fraction of native structures scoring better than decoy structures by PDScore) Decoy-set validation: misfold set (Table 4), hg_structal and 4state_reduced sets (Table 5A/B), ig_structal (Supplementary Table 1), fisa set (Table 5C) as stated per decoy set (e.g., ~20 proteins in misfold; ~30 proteins x 30 structures in hg_structal; 7 proteins x ~600 structures in 4state_reduced) not stated
Approaches that could also have been used
  • Differences in mean EPD between amino-acid pairs (Tables 1-3, Figures 1-3) are presented as descriptive means and SDs compared visually, without a formal statistical test of difference.
    Could also: A formal test such as one-way ANOVA (or Kruskal-Wallis if distributions are non-Gaussian) across amino-acid pair groups, potentially with a post-hoc comparison method. — This would provide a quantitative measure (p-value, effect size) of how distinguishable the EPD distributions are between pairs, complementing the visual scatter/probability-distribution comparisons.
  • The Gaussian nature of EPD distributions (e.g., C-N bond mean=420, SD=55) is described based on visual inspection of the plotted distributions.
    Could also: A formal normality test such as Shapiro-Wilk or Kolmogorov-Smirnov. — This would give a numerical basis for the Gaussian assumption underlying the Boltzmann-based statistical potential, alongside the visual distribution plots.
  • Decoy-set performance is summarized as a single specificity value per set (e.g., 0.91, 0.94, 0.93) without a reported measure of variability.
    Could also: Bootstrap resampling to obtain a confidence interval around specificity, or reporting complementary metrics such as sensitivity, precision, or AUC-ROC across a range of thresholds. — This would convey the precision of the specificity estimate and show how performance varies with the discrimination threshold, beyond a single point value.
  • Sample sizes per amino-acid pair vary considerably (e.g., n=36 for 'CC' vs n=688 for '4PTI'-scale sets), and only SD (not SEM or CI) is reported alongside the mean.
    Could also: Reporting the standard error of the mean or a 95% confidence interval for the mean EPD per pair. — This would make the precision of each mean estimate explicit, which is particularly informative for pairs based on smaller samples.
  • In the misfold decoy set (Table 4), each protein contributes a natural correct/incorrect PDScore pair, and results are reported as per-structure specificity (0 or 1) without an aggregate paired statistical comparison.
    Could also: A paired test such as the Wilcoxon signed-rank test comparing correct vs incorrect PDScores across the set. — This would provide an overall statistical summary of whether correct structures score better than incorrect ones across the paired set, in addition to the per-structure specificity tally.
  • Many amino-acid pairs (up to 190 possible pairs) are evaluated for their discriminatory utility without a stated adjustment for multiple comparisons.
    Could also: A multiple-testing correction procedure such as Benjamini-Hochberg FDR control, if the pairwise evaluations were framed as a family of hypothesis tests. — This would control the false-discovery rate when many pairs are being screened simultaneously for use as discriminators.
Software: APBS (Adaptive Poisson-Boltzmann Solver) · PDB2PQR

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

table4_named_failures_direction
Reported
Paper Table 4 / Discussion (misfold decoy set, Holm & Sander 1992): 'Barring three structures (PDBids: 1CBH, 1FDX and 2SSI), the PDScore of the incorrect structure is higher than that of the correct structures.' I.e. for 1CBH, 1FDX, 2SSI specifically, the incorrect (misfolded) structure scores LOWER (better/more native-like) than the correct native structure -- these are the paper's named failure cases (specificity=0).
Reproduced
Independent reimplementation (pdb2pqr30 v3.7.1 + APBS 3.4.1 electrostatics, pdie=2.0/sdie=78.54/srad=1.4A/T=298.15K/zero ionic strength, matching paper's stated parameters; Python reimplementation of escapist.pl's CB-based PDscore against the shipped pd.CB.score.full lookup table) reproduces this exact qualitative direction for all 3 named cases: 1cbh native=9.5 vs incorrect(1cbhon1ppt)=5.2; 1fdx native=19.3 vs incorrect(1fdxon5rxn)=12.7; 2ssi native=11.9 vs incorrect(2ssion2cdv)=10.4 -- incorrect scores lower than native in all three, matching the paper's specific named-failure claim.
within tolerance
table4_overall_specificity
Reported
Paper implies that across the misfold decoy set (dataset's own doc/NOTES documents 23 unique native PDB entries forming 26 correct-incorrect pairings, three natives -- 1sn3, 2ci2, 2cro -- each paired against two different misfolded decoys), only the 3 named structures (1CBH, 1FDX, 2SSI) fail to discriminate correctly, i.e. roughly 23/26 (~88%) of pairings should have the native structure scoring lower (better) than its misfolded counterpart.
Reproduced
16/26 pairings (61.5%) show the native structure scoring lower than its misfolded counterpart in our reimplementation; 10/26 fail, comprising the 3 paper-named cases (1cbh, 1fdx, 2ssi) plus 7 additional pairings the paper does not report as failures: 1rhd (vs 1rhdon2cyp), 1sn3 (vs 1sn3on2ci2), 2ci2 (vs 2ci2on2cro), 2cro (vs 2croon2ci2), 2i1b (vs 2i1bon1lh1), 2paz (vs 2pazon1bp2), 2tmn (vs 2tmnon2ts1).
did not match
table4_absolute_pdscore_magnitude
Reported
Paper Table 4 / PMC full text, named-example PDscore values: 1CBH correct=18.7, incorrect=12.6; 1FDX correct=33.0, incorrect=30.9; 2SSI correct=22.6, incorrect=20.0.
Reproduced
1cbh=9.5/5.2; 1fdx=19.3/12.7; 2ssi=11.9/10.4 -- reproduced absolute PDscore values are systematically lower than the paper's reported values (roughly 1.7-1.9x lower), though same order of magnitude and relative ranking (native vs misfolded) direction is preserved for these three. Likely sources of the scale offset: differences between the original ~2013-era PDB2PQR/APBS toolchain and the modern pdb2pqr30 v3.7.1 / APBS 3.4.1 used here (e.g. protonation-state assignment, grid/focusing defaults, or minor differences in the PARSE force field implementation over software versions) -- not independently isolated in this pass.
partial
table5a_aggregate_specificity
Reported
Paper Table 5A caption (hg_structal decoy set, homologous globin-family structures, Samudrala et al 1998c): 'each of which has 30 structures. The average specificity obtained for the set is 0.91.' (29 proteins, 1 native + 29 decoys = 30 structures each, including self-threading.)
Reproduced
After self-discovering and fixing a case-sensitivity bug in our own batch driver's decoy-file glob («job» initially processed only 14/29 proteins' own decoys correctly because the glob '${dname}_*_r.pdb' was case-sensitive and silently skipped the 15 proteins whose directory names contain uppercase chain-ID suffixes, e.g. 1bab-B; fixed via case-insensitive `find -iname` in «job», bringing results_hg.tsv from 435 to the full 870 lines = 29 natives + 841 decoys, 0 pipeline failures), the corrected full dataset gives overall_pooled_specificity=0.454 and overall_mean_of_per_protein_specificity=0.454 (equal because every protein has the same n=29 decoy cohort size) -- roughly half the paper's reported 0.91.
did not match
table5a_1ash_worked_example
Reported
Paper Table 5A row: PDB=1ASH, NRes=147, NStructures=30, Specificity=1 (perfect discrimination -- native scores best among its full 30-structure cohort). ESCAPIST manual's separate worked example for 1ash states the native structure's PDscore should be ~19.9, theminimum among a cohort ranging ~19-37.
Reproduced
Our reimplementation (same pipeline as used for Table 4 and the rest of Table 5A, full 29-decoy cohort after the case-sensitivity bugfix): 1ash native PDscore=15.2 (not ~19.9 as the manual's worked example states), n_decoys=29, n_worse=18, specificity=0.621 (not 1). The native's PDscore of 15.2 is explicitly NOT the minimum among its own 29-decoy cohort (cohort minimum=13.3, held by a decoy) -- directly contradicting both the paper's own Table 5A per-protein value for 1ASH (specificity=1) and the manual's separate worked-example claim.
did not match
pipeline_methodology_fidelity
Reported
Paper Methods: PDscore computed via PDB2PQR (PARSE force field) + APBS electrostatics (mg-auto focusing, pdie=2.0, sdie=~78, srad=1.4A, T=298K, zero ionic strength) then escapist.pl scoring of consecutive Cbeta electrostatic-potential differences against a PISCES-trained lookup table (pd.CB.score.full, shipped in the repo).
Reproduced
Successfully reimplemented and ran this full pipeline end-to-end (pdb2pqr30 -> APBS 3.4.1 -> pqr2csv (reimplemented, tool missing from modern APBS distribution) -> multivalue -> Python reimplementation of escapist.pl's CB-scoring logic, verified line-by-line against the original 432-line escapist.pl) for all 49 misfold-set structures with 0 pipeline failures (0/49 FAIL_* entries), and separately for the hg_structal set (see datasets/hummel_jobs). APBS .in parameters generated by pdb2pqr30 match the paper's stated values exactly (pdie=2.0, sdie=78.54, srad=1.40, temp=298.15K, no ionic-strength lines).
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 42/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🔴6. Severity of the deviation
🔴7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

Input identity is essentially perfect here — the exact public Decoys'R'Us misfold and hg_structal sets, the authors' own shipped pd.CB.score.full lookup table, and APBS parameters matching the paper's stated pdie=2.0/sdie=78.54/srad=1.4Å/298.15K — so the deviation cannot be blamed on data availability or cohort choice. The pipeline reproduced cleanly (0/49 misfold and 0/870 hg_structal failures, escapist.pl verified line-by-line), yet average specificity comes out 0.454 versus the Table 5A caption's 0.91, and the paper's own showcase protein 1ASH gives specificity 0.621 with native PDscore 15.2 sitting above a decoy at 13.3 — against a tabulated specificity of 1 and an ESCAPIST-manual worked example of ~19.9-as-cohort-minimum, which the shipped code itself contradicts. Absolute PDscores are additionally ~1.7–1.9x low (1CBH 9.5/5.2 vs 18.7/12.6), plausibly a 2013-vs-modern toolchain scale offset — but specificity is rank-based and scale-invariant, so that cannot excuse the halved discrimination. This lands on the authors' side: the headline performance numbers are not derivable from what was deposited, and the only element that reproduced faithfully was the paper's failure list (1CBH/1FDX/2SSI), tempered only by the fact that Tables 5B/5C/S2 were left unattempted.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.