The electrostatic profile of consecutive Cβ atoms applied to protein structure quality assessment.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🔴The deviation was non-trivial in magnitude
- 🔴The central claim did not (fully) hold under reproduction
- 🔴Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
FINAL: Two independent decoy-set comparisons were run end-to-end via an independent pdb2pqr30+APBS3.4.1 electrostatics pipeline and a from-scratch Python reimplementation of escapist.pl's CB-based PDscore, verified line-by-line against the original Perl. Result is a genuine PARTIAL reproduction overall -- not 1:1, not a full mismatch, with real evidence on both sides. Table 4 (misfold decoy set, 49 structures, 26 correct/incorrect pairings): the paper's specific named claim ('barring 1CBH, 1FDX, 2SSI, incorrect structures score worse than correct') reproduces exactly in direction for all 3 named cases. But the implied aggregate specificity does NOT reproduce (ours: 16/26=61.5% vs paper's implied ~23/26=88%), 7 additional pairings fail that the paper doesn't name, and absolute PDscore magnitudes run systematically ~1.7-1.9x lower than the paper's reported values for the 3 named cases (direction/ranking preserved, scale not) -- cause not isolated (candidate: PDB2PQR/APBS version differences vs. the original 2013-era toolchain). Table 5A (hg_structal decoy set, 29 globin-family proteins, full 29x29 self-inclusive threading matrix = 841 decoys): after self-discovering and fixing a case-sensitivity bug in our own batch driver (initial run silently processed only 14/29 proteins' own decoys; fixed via «job», confirmed complete and bug-free at 870/870 lines, 0 pipeline failures), the corrected result is a clear MISMATCH against the paper -- overall pooled specificity 0.454 vs the paper's reported 0.91, and the specific, directly-checkable 1ASH worked example (paper's own Table 5A: specificity=1; ESCAPIST manual's worked example: native PDscore19.9, cohort minimum) reproduces as specificity=0.621, native PDscore=15.2, and the native is explicitly NOT the minimum in its own 29-decoy cohort (minimum=13.3). Dataset profiling found the paper's formal 'Data' deposit (zenodo:10.5281/zenodo.7134) is actually a code/manual archive identical in substance to the GitHub repo, not a distinct experimental dataset -- the real structural data driving every result comes from the third-party Decoys'R'Us database (dd.compbio.org), which the paper never formally cites as a data accession; that dataset's own hg_structal deposit is itself complete and well-formed (870/870 structures verified), so the Table 5A mismatch traces to the scoring pipeline/comparison, not to missing or malformed decoy data. Per the '80/20 is a floor' mandate, this pass covered both Table 4 and Table 5A rather than stopping at Table 4 alone. NOT attempted: Tables 5B (4state_reduced), 5C (fisa), and Supplementary S2 (ig_structal) -- all substantially larger (thousands of structures each), and further extension was also blocked at the end of this pass by a fleet-wide HPC home-directory quota outage (a foreign infra issue affecting all reproduction rooms on this cluster, not specific to this task, and explicitly not grounds to mark this reproduction as failed) that prevented new hummel_submit real-compute jobs; also not attempted: independent re-derivation of the PISCES training statistics behind pd.CB.score.full (reused as shipped).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-31
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe paper tests whether the electrostatic potential difference (EPD) between Cβ atoms of consecutive residues constitutes an amino-acid-type-specific signature that can serve as a statistical potential for model quality assessment; based on the Boltzmann hypothesis and Anfinsen's thermodynamic hypothesis, the authors hypothesize that deviation of observed EPD values from learned mean values is minimized in the native structure.
- ★ The EPD between Cβ atoms of consecutive residues provides unique signatures of amino acid pair types and can discriminate native from decoy protein structures. finding
- ★ ESCAPIST, a novel knowledge-based MQAP scoring function (PDScore) derived from Cβ EPD statistics of 1000 non-homologous PISCES structures, ranks native structures better than decoys. method
- ★ The EPD of the C-N peptide bond and the EPD between consecutive Cα atoms are independent of amino acid type and therefore non-discriminatory, because interatomic distances for these atoms are nearly invariant across structures. finding
- Amino acid pairs containing cysteine have anomalously high variance in Cβ EPD (SD ~90) and cannot be used for discrimination. finding
- Proline-containing pairs are outliers in Cα EPD with higher absolute mean values, highlighting the unique nature of proline in protein structures. finding
- The Cβ EPD signal is non-trivial because the variance of EPD between residue pairs increases with increasing sequence separation, i.e., correlation is specific to consecutive residues. mechanism
- The electrostatic discriminator fails on the fisa decoy set, which consists of physically nonviable structures for which electrostatic analysis is meaningless (these are better discriminated by a distance-based criterion). finding
- Source code and manual are publicly available at https://github.com/sanchak/mqap and archived at doi 10.5281/zenodo.7134. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Finite-difference Poisson-Boltzmann electrostatic potential calculation (EPD of Cβ atoms in consecutive residues) | 1000 non-homologous protein crystal structures from the PISCES database (20% sequence identity cutoff, 1.6 Å resolution cutoff, R-factor cutoff 0.25) | none | Mean and standard deviation of electrostatic potential difference (kT/e units) per amino acid pair; number of samples per pair | APBS (Adaptive Poisson-Boltzmann Solver) with PDB2PQR |
| Electrostatic potential difference profiling of the C-N peptide bond | Same set of 1000 non-homologous PISCES protein structures | none | Probability distribution, mean and SD of EPD across all amino acid pairs | APBS / PDB2PQR |
| Electrostatic potential difference profiling between Cα atoms of consecutive residues | Same set of 1000 non-homologous PISCES protein structures | none | Probability distribution, mean and SD of EPD per amino acid pair (proline pairs analyzed separately) | APBS / PDB2PQR |
| Sequence-distance dependence analysis of Cβ EPD | PISCES structure set, amino acid pairs 'DF' (Asp/Phe) and 'HS' (His/Ser), ≥30 samples per sequence distance | increasing sequence separation between residues (distance 1 to larger) | Standard deviation of EPD as a function of sequence distance | — |
| Decoy discrimination benchmark (PDScore) — misfold set | ~20 protein structures each with a correct and an incorrect (misfolded) structure, from Decoys 'R' Us | deliberately misfolded models | PDScore of correct vs incorrect structure; per-PDB specificity (0/1) | Decoys 'R' Us database (http://dd.compbio.washington.edu) |
| Decoy discrimination benchmark (PDScore) — hg_structal set | ~30 globin-family proteins, 30 structures each (including native) | computationally generated decoys | Specificity of native-structure identification per protein and averaged | Decoys 'R' Us database |
| Decoy discrimination benchmark (PDScore) — 4state_reduced set | 7 proteins (1CTF, 1R69, 1SN3, 2CRO, 3ICB, 4PTI, 4RXN), ~600–690 structures each | computationally generated decoys | Specificity of native-structure identification per protein and averaged | Decoys 'R' Us database |
| Decoy discrimination benchmark (PDScore) — ig_structal and fisa sets | ig_structal immunoglobulin decoy set; fisa set of 4 proteins (1FC2, 1HDDC, 2CRO, 4ICB) with 501 structures each | computationally generated decoys | Specificity of native-structure identification | Decoys 'R' Us database |
- – Average specificity on the hg_structal, 4state_reduced and ig_structal decoy sets 0.91, 0.94 and 0.93 respectively
- ▲ In the misfold decoy set, the incorrect structure had a higher PDScore than the correct structure for all but three proteins (1CBH, 1FDX, 2SSI) 20 of 23 proteins correctly discriminated
- – EPD of the C-N peptide bond has a Gaussian distribution with identical means for all amino acid pairs and large variance, so amino acids are indistinguishable by this feature mean = 420 EPD units, SD = 55 EPD units
- ▲ Mean EPD values between Cβ atoms of consecutive residues are much more varied across amino acid pairs than for Cα atoms or the C-N bond, with lower variance typical SD ~35 EPD units for non-cysteine Cβ pairs
- ▲ Cysteine-containing Cβ pairs show random means and high standard deviation, precluding discrimination (exception: 'CC' pair with low mean and SD) SD ~90 vs ~35 for other pairs; CC: mean -7.1, SD 30.4
- ▲ Proline-containing Cα pairs are outliers with higher absolute mean EPD while SD magnitude is unchanged e.g. DP mean -184.5 (SD 29.5); PS mean 172.3 (SD 25.9)
- ▲ SD of Cβ EPD increases with sequence separation, showing loss of correlation beyond consecutive residues from 29.8 ('DF') and 31.8 ('HS') EPD units at distance 1 to ~60 EPD units
- ▼ Low specificities obtained on the fisa decoy set 4ICB 0, 1HDDC 0.1, 1FC2 0.4, 2CRO 0.7
- mean mean = 420 EPD units, SD = 55 EPD units (EPD of the C-N peptide bond, Gaussian distribution across all structures)
- other specificity 0.91 (Average specificity on hg_structal decoy set (~30 proteins x 30 structures))
- other specificity 0.94 (Average specificity on 4state_reduced decoy set (~600 structures per protein))
- other specificity 0.93 (Specificity on ig_structal decoy set)
- count 1000 proteins; 20% identity cutoff, 1.6 Å resolution cutoff, 0.25 R-factor cutoff (PISCES learning set used to derive mean EPD values)
- mean DF: -108.9 (SD 29.5, n=481); HS: 93.7 (SD 31.8, n=235); HT: 95.1 (SD 31.5, n=235) (Representative discriminatory Cβ consecutive-pair EPD values (Table 3))
- mean CS: 106.3 (SD 98.1, n=184); CT: 109.9 (SD 97.5, n=173); CM: 61.9 (SD 100, n=39) (Cysteine-containing Cβ pairs with high SD (Table 2))
- other SD rises from 29.8 (DF) and 31.8 (HS) at sequence distance 1 to around 60 EPD units (Sequence-distance dependence of Cβ EPD variance; at least 30 samples per distance)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a computational/structural-bioinformatics study rather than a classical wet-lab experiment: electrostatic potential differences (EPD) between backbone atoms were computed from a training set of 1000 non-homologous PISCES structures using APBS, summarized per amino-acid pair as mean, standard deviation, and sample count, and used to build a knowledge-based scoring function (PDScore). The scoring function was then evaluated by computing 'specificity' (the fraction of cases where the native/correct structure scored better than incorrect/decoy structures) across several standard decoy-set benchmarks (misfold, hg_structal, 4state_reduced, ig_structal, fisa). Results are reported as descriptive summary statistics (mean, SD, n) in tables and as aggregate specificity values per decoy set, without formal inferential hypothesis tests.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Descriptive summary statistics (mean, SD, sample size) per amino-acid pair, no formal hypothesis test reported | Tables 1-3, Figures 1-4 (EPD distributions for C-N bond, Cα, and Cβ atom pairs) | varies by pair, as stated per row in Tables 1-3 (e.g., n=36 to n=688) | stated (Gaussian/Boltzmann distribution assumed for the underlying statistical potential) |
| Specificity (fraction of native structures scoring better than decoy structures by PDScore) | Decoy-set validation: misfold set (Table 4), hg_structal and 4state_reduced sets (Table 5A/B), ig_structal (Supplementary Table 1), fisa set (Table 5C) | as stated per decoy set (e.g., ~20 proteins in misfold; ~30 proteins x 30 structures in hg_structal; 7 proteins x ~600 structures in 4state_reduced) | not stated |
-
Differences in mean EPD between amino-acid pairs (Tables 1-3, Figures 1-3) are presented as descriptive means and SDs compared visually, without a formal statistical test of difference.↳ Could also: A formal test such as one-way ANOVA (or Kruskal-Wallis if distributions are non-Gaussian) across amino-acid pair groups, potentially with a post-hoc comparison method. — This would provide a quantitative measure (p-value, effect size) of how distinguishable the EPD distributions are between pairs, complementing the visual scatter/probability-distribution comparisons.
-
The Gaussian nature of EPD distributions (e.g., C-N bond mean=420, SD=55) is described based on visual inspection of the plotted distributions.↳ Could also: A formal normality test such as Shapiro-Wilk or Kolmogorov-Smirnov. — This would give a numerical basis for the Gaussian assumption underlying the Boltzmann-based statistical potential, alongside the visual distribution plots.
-
Decoy-set performance is summarized as a single specificity value per set (e.g., 0.91, 0.94, 0.93) without a reported measure of variability.↳ Could also: Bootstrap resampling to obtain a confidence interval around specificity, or reporting complementary metrics such as sensitivity, precision, or AUC-ROC across a range of thresholds. — This would convey the precision of the specificity estimate and show how performance varies with the discrimination threshold, beyond a single point value.
-
Sample sizes per amino-acid pair vary considerably (e.g., n=36 for 'CC' vs n=688 for '4PTI'-scale sets), and only SD (not SEM or CI) is reported alongside the mean.↳ Could also: Reporting the standard error of the mean or a 95% confidence interval for the mean EPD per pair. — This would make the precision of each mean estimate explicit, which is particularly informative for pairs based on smaller samples.
-
In the misfold decoy set (Table 4), each protein contributes a natural correct/incorrect PDScore pair, and results are reported as per-structure specificity (0 or 1) without an aggregate paired statistical comparison.↳ Could also: A paired test such as the Wilcoxon signed-rank test comparing correct vs incorrect PDScores across the set. — This would provide an overall statistical summary of whether correct structures score better than incorrect ones across the paired set, in addition to the per-structure specificity tally.
-
Many amino-acid pairs (up to 190 possible pairs) are evaluated for their discriminatory utility without a stated adjustment for multiple comparisons.↳ Could also: A multiple-testing correction procedure such as Benjamini-Hochberg FDR control, if the pairwise evaluations were framed as a family of hypothesis tests. — This would control the false-discovery rate when many pairs are being screened simultaneously for use as discriminators.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Input identity is essentially perfect here — the exact public Decoys'R'Us misfold and hg_structal sets, the authors' own shipped pd.CB.score.full lookup table, and APBS parameters matching the paper's stated pdie=2.0/sdie=78.54/srad=1.4Å/298.15K — so the deviation cannot be blamed on data availability or cohort choice. The pipeline reproduced cleanly (0/49 misfold and 0/870 hg_structal failures, escapist.pl verified line-by-line), yet average specificity comes out 0.454 versus the Table 5A caption's 0.91, and the paper's own showcase protein 1ASH gives specificity 0.621 with native PDscore 15.2 sitting above a decoy at 13.3 — against a tabulated specificity of 1 and an ESCAPIST-manual worked example of ~19.9-as-cohort-minimum, which the shipped code itself contradicts. Absolute PDscores are additionally ~1.7–1.9x low (1CBH 9.5/5.2 vs 18.7/12.6), plausibly a 2013-vs-modern toolchain scale offset — but specificity is rank-based and scale-invariant, so that cannot excuse the halved discrimination. This lands on the authors' side: the headline performance numbers are not derivable from what was deposited, and the only element that reproduced faithfully was the paper's failure list (1CBH/1FDX/2SSI), tempered only by the fact that Tables 5B/5C/S2 were left unattempted.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.