Protein structure quality assessment based on the distance profiles of consecutive backbone Cα atoms.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Of 7 graded claims: 3 exact (1CNV backbone anomalies, 2JLI cleaved-bond distance, Decoys R Us 1FC2-vs-decoy example), 1 within-tolerance (top100H reference state via a documented Top8000 substitute, since the original dataset is dead), 3 partial/mismatch (CASP8 outlier count: 1/121 reproduces vs paper's 5/121, with 2 of the missing 4 explained by post-2014 PDB corrections and 2genuinely unexplained; 1ADS violation only weakly located; PISCES Table 1 resolution-binned trend does not reproduce the paper's own non-monotonic bin pattern, though order-of-magnitude RDCC values match). The core RDCC algorithm itself is faithfully ported from the actual Perl source and behaves correctly end-to-end: native structures cluster far under the 0.012 Å cutoff, decoys/errors are correctly flagged, and the two most dramatic illustrative examples (2JLI cleaved bond, 1FC2 decoy) reproduce exactly. Two legacy reference datasets (top100H, the exact 2013 PISCES snapshot) are confirmed unobtainable 13 years on and were replaced with clearly-labeled current substitutes rather than skipped. Not attempted: nothing was out-of-scope, as this is a fully computational method paper; the T0492/T0419 CASP8 discrepancies remain unexplained rather than force-matched.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe authors ask whether a simple geometric statistical potential — the root-mean-square deviation of consecutive Cα–Cα distances from an ideal 3.8 Å (RDCC) — is minimized in native protein structures and can therefore serve as a model quality assessment discriminator. They hypothesize that CADistScore is minimal for the native structure among a set of candidate/decoy structures.
- ★ The distance between consecutive backbone Cα atoms in high-quality structures is normally distributed with mean 3.8 Å and standard deviation 0.04 Å, justifying a reference state in which all consecutive Cα atoms are 3.8 Å apart. finding
- ★ RDCC (root-mean-square deviation of consecutive Cα distances from 3.8 Å) is minimized in native structures and can be used as a model quality assessment program (PROQUAD). method
- ★ A RDCC cutoff of 0.012 Å (specificity of 1 on top100H) filters out non-native/low-quality structures. method
- ★ All decoy structures of the fisa decoy set from Decoys 'R' Us violate the RDCC criterion, indicating the fisa set is physically non-viable and problematic for benchmarking other methods. finding
- ★ Existing structure-validation tools (MolProbity, ProSA, WHATIF) do not check consecutive Cα distances and failed to report the Cα distance anomaly in the fisa decoys, so RDCC is complementary to them. finding
- ★ The RDCC relationship is not an equivalence relation: high RDCC implies low structure quality, but low quality does not necessitate high RDCC; RDCC should be used as a fast first-pass filter. mechanism
- RDCC values are independent of the crystallographic resolution of the structure. finding
- Source code and manual for the RDCC/PROQUAD discriminator are released at https://github.com/sanchak/mqap and archived at doi 10.5281/zenodo.7134. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Statistical analysis of consecutive Cα–Cα interatomic distances (frequency/probability distribution) | top100H database of ~100 high-quality protein crystal structures | none | Distance between Res_n(Cα) and Res_n+1(Cα); mean and standard deviation of the distance distribution | top100H database (http://kinemage.biochem.duke.edu/databases/top100.php) |
| RDCC/CADistScore computation (in-house algorithm AssessCADist, ignoring IgnoreNTerm/IgnoreCTerm terminal residue pairs) | top100H high-quality structures (~100 proteins; 3 excluded: 2ER7, 1XSO, 4PTP with multiple conformations) | none | RDCC score per structure; specificity as a function of cutoff | Source code at https://github.com/sanchak/mqap (10.5281/zenodo.7134) |
| RDCC discriminator benchmarking on decoy sets | Decoys 'R' Us database — fisa, hg_structal, misfold, 4state_reduced sets (including PDBid 1FC2 and decoy AXPROA00-MIN) | decoy (computationally generated non-native) structures vs native structure | RDCC per structure vs 0.012 Å cutoff; number of decoys discriminated | Decoys 'R' Us database (http://dd.compbio.washington.edu/) |
| RDCC evaluation on predicted-model decoy suite | I-TASSER CASP8 decoy set, 121 test cases | none (predicted models); residue-renumbering correction applied to T0423 and T0470 | RDCC per target vs 0.012 Å threshold | http://zhanglab.ccmb.med.umich.edu/casp8/decoys |
| RDCC versus crystallographic resolution analysis (frequency distributions binned by resolution) | Non-homologous PDB structures (20% sequence identity cutoff) from the PISCES database, binned at 1.0, 1.5, 2.0, 2.5 and 3.0 Å resolution | none (outliers such as PDBid 2JLI removed) | Mean and standard deviation of RDCC per resolution bin | PISCES database (http://dunbrack.fccc.edu/PISCES.php) |
| Comparative structure validation / clash analysis | Native structure PDBid 1FC2 and fisa decoy AXPROA00-MIN | native vs decoy comparison | MolProbity ClashScore and Cβ deviations; ProSA Z-score; WHATIF steric clash report | MolProbity, ProSA, WHATIF web servers |
| Structural superimposition | Native PDBid 1FC2 and fisa decoy AXPROA00-MIN | native vs decoy comparison | Superimposed backbone geometry; Ile12(Cα)–Leu13(Cα) distance | MUSTANG |
| Case-based inspection of Cα distance outliers against PDB annotations | Individual PDB entries — concanavalin B (1CNV), aldose reductase (1ADS), 2JLI | none | Per-residue-pair Cα–Cα distances and their correspondence to cis-peptide/non-trans or cleaved peptide bond annotations | — |
- – In top100H, 88% of consecutive Cα pairs are 3.8 Å apart, 8% 3.9 Å, 3% 3.7 Å, and only 0.1% deviate (range 2.9–4 Å), giving a normal distribution with mean 3.8 Å and SD 0.04 Å. 14,281/16,162 = 88% at 3.8 Å; SD 0.04 Å
- ▼ All top100H structures show low RDCC values (barring three structures excluded for multiple conformations), validating that RDCC is minimized in native structures.
- – A RDCC cutoff of 0.012 Å yields a specificity of 1 on the top100H set. specificity = 1
- ▲ All 500 decoy structures for each protein in the fisa decoy set exceed the 0.012 Å RDCC cutoff and are discriminated, whereas the native structure of every Decoys 'R' Us set tested is below the cutoff. 500/500 decoys per protein discriminated
- ▲ For 1FC2, the Ile12(Cα)–Leu13(Cα) distance is 3.8 Å in the native structure but 4.1 Å in the decoy AXPROA00-MIN. 3.8 Å → 4.1 Å
- – MolProbity (ClashScore, Cβ deviations) and WHATIF flagged steric clashes in the decoy, but ProSA could not discriminate it from the native; none reported the abnormal consecutive Cα distances. ProSA Z-scores −4.12 (decoy) vs −5.28 (native)
- – Only 5 of 121 I-TASSER CASP8 test cases exceeded the 0.012 Å threshold, and at least two of these were due to erroneous residue numbering that, once corrected, gave normal RDCC values. 5/121; T0423 0.09 → 0.002 Å; T0470 0.051 → 0.001 Å
- – Mean RDCC across PISCES resolution bins ranges only from 1.8 to 3.0 × 10⁻³ Å, far below the 0.012 Å cutoff, showing no correlation between RDCC and resolution. 1.8–3.0 × 10⁻³ Å
- – The hg_structal and misfold decoy sets are indistinguishable by RDCC, and only a few structures failed in 4state_reduced, showing RDCC is a one-way (non-equivalence) discriminator.
- count 16,162 pairs of consecutive Cα atom distances; 14,281 (88%) at 3.8 Å, 1297 (8%) at 3.9 Å, 553 (3%) at 3.7 Å, 31 (0.1%) other (Consecutive Cα distance distribution in ~100 top100H proteins)
- mean mean 3.8 Å, standard deviation 0.04 Å (Normal distribution of consecutive Cα distances in top100H)
- other RDCC cutoff 0.012 Å with specificity of 1 (Chosen threshold for filtering non-native structures)
- other T0492 0.013 Å, T0476 0.014 Å, T0419 0.025 Å, T0470 0.051 Å, T0423 0.09 Å (The only 5 of 121 I-TASSER CASP8 decoy test cases above the 0.012 Å threshold)
- other ProSA Z-scores −4.12 (decoy AXPROA00-MIN) and −5.28 (native 1FC2) (ProSA failed to discriminate decoy from native)
- mean Mean RDCC (10⁻³ Å) / SD (10⁻³ Å) by resolution: ≤1.0 Å: 3.0/4.3 (N=165); ≤1.5 Å: 1.9/2.7 (N=682); ≤2.0 Å: 1.8/2.8 (N=753); ≤2.5 Å: 2.2/2.6 (N=422); ≤3.0 Å: 2.2/2.5 (N=269) (PISCES non-homologous structures binned by resolution (Table 1))
- other 3.8 Å (native) vs 4.1 Å (decoy) for Ile12(Cα)–Leu13(Cα) (PDBid 1FC2 native vs fisa decoy AXPROA00-MIN)
- count Four violations of the 3.8 Å constraint in 1CNV: Ile33/Ser34 4 Å, Ser34/Phe35 3 Å, Pro56/Ser57 4 Å, Trp265/Asn266 3.4 Å (cis-peptide-bond-associated outliers in concanavalin B)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper develops a knowledge-based structural quality metric (RDCC, root-mean-square deviation of consecutive Cα distances from 3.8 Å) derived from distance statistics in the top100H structure database, then evaluates a fixed cutoff (0.012 Å) using a specificity calculation and applies this threshold to classify native versus decoy structures across several benchmark decoy sets and a resolution-binned PISCES set. Results are reported primarily as descriptive statistics (means, standard deviations, and specificity values) and visual/tabular comparisons rather than formal inferential hypothesis tests.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Descriptive distribution characterization (mean, SD) of Cα-Cα distances | Figure 1a, top100H database distances | ~100 proteins, 16,162 pairs of consecutive Cα distances | not stated |
| Specificity calculation (TN/(FP+TN)) across candidate cutoff values | Figure 1c, cutoff selection for RDCC threshold | ~100 top100H proteins | not stated |
| Threshold-based classification (fixed RDCC cutoff of 0.012 Å) applied to decoy sets | Figures 1d, 2; Decoys 'R' Us and I-TASSER CASP8 decoy suite | 121 I-TASSER testcases; >500 decoys per protein in fisa set; multiple decoy sets (hg_structal, misfold, 4state_reduced, fisa) | not stated |
| Mean and SD summary by resolution bin | Table 1, PISCES-derived non-homologous structure sets | 165, 682, 422, 753, 269 structures per resolution bin | not stated |
-
A single cutoff value (0.012 Å) was chosen based on a specificity curve (Figure 1c) without accompanying sensitivity trade-off statistics at each point.↳ Could also: A full ROC curve with AUC could also be used — Reporting sensitivity alongside specificity across the full range of thresholds would convey the trade-off more completely and allow readers to see how cutoff choice affects both false negatives and false positives.
-
Native versus decoy RDCC values are compared visually/qualitatively (Figures 1d, 2) rather than via a formal statistical test of difference.↳ Could also: A nonparametric test such as the Mann-Whitney U test, or a t-test if distributional assumptions are met, could also be applied between native and decoy RDCC distributions — A formal test would provide a quantitative measure (e.g., a p-value or effect size) of how distinguishable the two groups are, complementing the threshold-based classification already used.
-
The claim that Cα-Cα distances are normally distributed (Figure 1a) is presented as a visual/qualitative observation.↳ Could also: A formal normality test (e.g., Shapiro-Wilk or Kolmogorov-Smirnov) could also be applied — A formal test would give a statistical basis for the normality claim underlying the choice of mean and SD as the reference-state parameters.
-
Table 1 reports mean and SD of RDCC across resolution bins and concludes RDCC is largely independent of resolution.↳ Could also: A correlation coefficient (e.g., Pearson or Spearman) between resolution and RDCC, or a regression analysis, could also be computed — A quantitative correlation statistic would formally support the stated absence of a relationship between resolution and RDCC, beyond the qualitative comparison of binned means.
-
Sample sizes for the top100H and decoy analyses are given as raw counts without an accompanying power or precision analysis.↳ Could also: A confidence interval around the estimated mean and SD of the reference distance (3.8 Å, 0.04 Å) could also be reported — Confidence intervals would communicate the precision of these reference parameters, which underlie the downstream cutoff selection.
-
The RDCC threshold is applied uniformly across several distinct decoy sets (fisa, hg_structal, misfold, 4state_reduced, I-TASSER CASP8) as separate evaluations.↳ Could also: A pooled or meta-analytic comparison with a multiple-testing correction (e.g., Bonferroni or Benjamini-Hochberg) across decoy sets could also be used — This would formally account for testing the same threshold across multiple independent datasets when summarizing overall discriminative performance.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The RDCC/CADistScore method itself reproduces convincingly: ported line-for-line from the actual Perl source, it hits 1CNV (4.0/3.0/4.0/3.4 Å), 2JLI (Asn263-Pro264 = 9.4 Å) and the 1FC2 native-vs-decoy pair (3.775→3.8 Å vs 4.061→4.1 Å, RDCC 0.00597 vs 0.02339) exactly, and 121/121 CASP8 natives cluster far below the 0.012 Å cutoff (mean 0.00203). The deviations sit almost entirely on the input/availability side: top100H is dead and the 2013 PISCES snapshot unrecoverable, so both had to be replaced by labelled substitutes (140 Top8000 chains; 2026 PISCES, 742/2291 chains), and post-2014 PDB remediation explains why T0470/T0423 now return the paper's own corrected 0.001/0.002. Two residuals remain on the authors' side: T0492 and T0419 are reported as outliers (0.013, 0.025) but score 0.00143/0.00187 with no correction noted, and the paper's Table 1 shows the worst mean+SD at the best resolution bin — a counterintuitive pattern that reproduction inverts. Severity is moderate (headline count 5/121 → 1/121) but the central conclusion is unaffected, hence yellow overall rather than red.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.