Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Protein structure quality assessment based on the distance profiles of consecutive backbone Cα atoms.

F1000Res · 2013
L1 71/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
71/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 38% of all assessed papers rank 694 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Of 7 graded claims: 3 exact (1CNV backbone anomalies, 2JLI cleaved-bond distance, Decoys R Us 1FC2-vs-decoy example), 1 within-tolerance (top100H reference state via a documented Top8000 substitute, since the original dataset is dead), 3 partial/mismatch (CASP8 outlier count: 1/121 reproduces vs paper's 5/121, with 2 of the missing 4 explained by post-2014 PDB corrections and 2genuinely unexplained; 1ADS violation only weakly located; PISCES Table 1 resolution-binned trend does not reproduce the paper's own non-monotonic bin pattern, though order-of-magnitude RDCC values match). The core RDCC algorithm itself is faithfully ported from the actual Perl source and behaves correctly end-to-end: native structures cluster far under the 0.012 Å cutoff, decoys/errors are correctly flagged, and the two most dramatic illustrative examples (2JLI cleaved bond, 1FC2 decoy) reproduce exactly. Two legacy reference datasets (top100H, the exact 2013 PISCES snapshot) are confirmed unobtainable 13 years on and were replaced with clearly-labeled current substitutes rather than skipped. Not attempted: nothing was out-of-scope, as this is a fully computational method paper; the T0492/T0419 CASP8 discrepancies remain unexplained rather than force-matched.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.7134

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The authors ask whether a simple geometric statistical potential — the root-mean-square deviation of consecutive Cα–Cα distances from an ideal 3.8 Å (RDCC) — is minimized in native protein structures and can therefore serve as a model quality assessment discriminator. They hypothesize that CADistScore is minimal for the native structure among a set of candidate/decoy structures.

Core claims
  • The distance between consecutive backbone Cα atoms in high-quality structures is normally distributed with mean 3.8 Å and standard deviation 0.04 Å, justifying a reference state in which all consecutive Cα atoms are 3.8 Å apart. finding
  • RDCC (root-mean-square deviation of consecutive Cα distances from 3.8 Å) is minimized in native structures and can be used as a model quality assessment program (PROQUAD). method
  • A RDCC cutoff of 0.012 Å (specificity of 1 on top100H) filters out non-native/low-quality structures. method
  • All decoy structures of the fisa decoy set from Decoys 'R' Us violate the RDCC criterion, indicating the fisa set is physically non-viable and problematic for benchmarking other methods. finding
  • Existing structure-validation tools (MolProbity, ProSA, WHATIF) do not check consecutive Cα distances and failed to report the Cα distance anomaly in the fisa decoys, so RDCC is complementary to them. finding
  • The RDCC relationship is not an equivalence relation: high RDCC implies low structure quality, but low quality does not necessitate high RDCC; RDCC should be used as a fast first-pass filter. mechanism
  • RDCC values are independent of the crystallographic resolution of the structure. finding
  • Source code and manual for the RDCC/PROQUAD discriminator are released at https://github.com/sanchak/mqap and archived at doi 10.5281/zenodo.7134. resource
Experimental setups
Assay System Perturbation Readout Platform
Statistical analysis of consecutive Cα–Cα interatomic distances (frequency/probability distribution) top100H database of ~100 high-quality protein crystal structures none Distance between Res_n(Cα) and Res_n+1(Cα); mean and standard deviation of the distance distribution top100H database (http://kinemage.biochem.duke.edu/databases/top100.php)
RDCC/CADistScore computation (in-house algorithm AssessCADist, ignoring IgnoreNTerm/IgnoreCTerm terminal residue pairs) top100H high-quality structures (~100 proteins; 3 excluded: 2ER7, 1XSO, 4PTP with multiple conformations) none RDCC score per structure; specificity as a function of cutoff Source code at https://github.com/sanchak/mqap (10.5281/zenodo.7134)
RDCC discriminator benchmarking on decoy sets Decoys 'R' Us database — fisa, hg_structal, misfold, 4state_reduced sets (including PDBid 1FC2 and decoy AXPROA00-MIN) decoy (computationally generated non-native) structures vs native structure RDCC per structure vs 0.012 Å cutoff; number of decoys discriminated Decoys 'R' Us database (http://dd.compbio.washington.edu/)
RDCC evaluation on predicted-model decoy suite I-TASSER CASP8 decoy set, 121 test cases none (predicted models); residue-renumbering correction applied to T0423 and T0470 RDCC per target vs 0.012 Å threshold http://zhanglab.ccmb.med.umich.edu/casp8/decoys
RDCC versus crystallographic resolution analysis (frequency distributions binned by resolution) Non-homologous PDB structures (20% sequence identity cutoff) from the PISCES database, binned at 1.0, 1.5, 2.0, 2.5 and 3.0 Å resolution none (outliers such as PDBid 2JLI removed) Mean and standard deviation of RDCC per resolution bin PISCES database (http://dunbrack.fccc.edu/PISCES.php)
Comparative structure validation / clash analysis Native structure PDBid 1FC2 and fisa decoy AXPROA00-MIN native vs decoy comparison MolProbity ClashScore and Cβ deviations; ProSA Z-score; WHATIF steric clash report MolProbity, ProSA, WHATIF web servers
Structural superimposition Native PDBid 1FC2 and fisa decoy AXPROA00-MIN native vs decoy comparison Superimposed backbone geometry; Ile12(Cα)–Leu13(Cα) distance MUSTANG
Case-based inspection of Cα distance outliers against PDB annotations Individual PDB entries — concanavalin B (1CNV), aldose reductase (1ADS), 2JLI none Per-residue-pair Cα–Cα distances and their correspondence to cis-peptide/non-trans or cleaved peptide bond annotations
Key results
  • In top100H, 88% of consecutive Cα pairs are 3.8 Å apart, 8% 3.9 Å, 3% 3.7 Å, and only 0.1% deviate (range 2.9–4 Å), giving a normal distribution with mean 3.8 Å and SD 0.04 Å. 14,281/16,162 = 88% at 3.8 Å; SD 0.04 Å
  • All top100H structures show low RDCC values (barring three structures excluded for multiple conformations), validating that RDCC is minimized in native structures.
  • A RDCC cutoff of 0.012 Å yields a specificity of 1 on the top100H set. specificity = 1
  • All 500 decoy structures for each protein in the fisa decoy set exceed the 0.012 Å RDCC cutoff and are discriminated, whereas the native structure of every Decoys 'R' Us set tested is below the cutoff. 500/500 decoys per protein discriminated
  • For 1FC2, the Ile12(Cα)–Leu13(Cα) distance is 3.8 Å in the native structure but 4.1 Å in the decoy AXPROA00-MIN. 3.8 Å → 4.1 Å
  • MolProbity (ClashScore, Cβ deviations) and WHATIF flagged steric clashes in the decoy, but ProSA could not discriminate it from the native; none reported the abnormal consecutive Cα distances. ProSA Z-scores −4.12 (decoy) vs −5.28 (native)
  • Only 5 of 121 I-TASSER CASP8 test cases exceeded the 0.012 Å threshold, and at least two of these were due to erroneous residue numbering that, once corrected, gave normal RDCC values. 5/121; T0423 0.09 → 0.002 Å; T0470 0.051 → 0.001 Å
  • Mean RDCC across PISCES resolution bins ranges only from 1.8 to 3.0 × 10⁻³ Å, far below the 0.012 Å cutoff, showing no correlation between RDCC and resolution. 1.8–3.0 × 10⁻³ Å
  • The hg_structal and misfold decoy sets are indistinguishable by RDCC, and only a few structures failed in 4state_reduced, showing RDCC is a one-way (non-equivalence) discriminator.
Key statistics
  • count 16,162 pairs of consecutive Cα atom distances; 14,281 (88%) at 3.8 Å, 1297 (8%) at 3.9 Å, 553 (3%) at 3.7 Å, 31 (0.1%) other (Consecutive Cα distance distribution in ~100 top100H proteins)
  • mean mean 3.8 Å, standard deviation 0.04 Å (Normal distribution of consecutive Cα distances in top100H)
  • other RDCC cutoff 0.012 Å with specificity of 1 (Chosen threshold for filtering non-native structures)
  • other T0492 0.013 Å, T0476 0.014 Å, T0419 0.025 Å, T0470 0.051 Å, T0423 0.09 Å (The only 5 of 121 I-TASSER CASP8 decoy test cases above the 0.012 Å threshold)
  • other ProSA Z-scores −4.12 (decoy AXPROA00-MIN) and −5.28 (native 1FC2) (ProSA failed to discriminate decoy from native)
  • mean Mean RDCC (10⁻³ Å) / SD (10⁻³ Å) by resolution: ≤1.0 Å: 3.0/4.3 (N=165); ≤1.5 Å: 1.9/2.7 (N=682); ≤2.0 Å: 1.8/2.8 (N=753); ≤2.5 Å: 2.2/2.6 (N=422); ≤3.0 Å: 2.2/2.5 (N=269) (PISCES non-homologous structures binned by resolution (Table 1))
  • other 3.8 Å (native) vs 4.1 Å (decoy) for Ile12(Cα)–Leu13(Cα) (PDBid 1FC2 native vs fisa decoy AXPROA00-MIN)
  • count Four violations of the 3.8 Å constraint in 1CNV: Ile33/Ser34 4 Å, Ser34/Phe35 3 Å, Pro56/Ser57 4 Å, Trp265/Asn266 3.4 Å (cis-peptide-bond-associated outliers in concanavalin B)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper develops a knowledge-based structural quality metric (RDCC, root-mean-square deviation of consecutive Cα distances from 3.8 Å) derived from distance statistics in the top100H structure database, then evaluates a fixed cutoff (0.012 Å) using a specificity calculation and applies this threshold to classify native versus decoy structures across several benchmark decoy sets and a resolution-binned PISCES set. Results are reported primarily as descriptive statistics (means, standard deviations, and specificity values) and visual/tabular comparisons rather than formal inferential hypothesis tests.

Replicationunclear Sample sizeSample sizes given as counts of structures/pairs analyzed per database or decoy set (e.g., ~100 proteins, 16,162 residue pairs, 121 decoy testcases, resolution-binned counts in Table 1); no power analysis or sample-size justification described Groupsnative (high-quality/top100H) protein structures vs. decoy (non-native) structures across multiple decoy sets; also structures grouped by resolution bin Pairingna Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionno
Statistical tests used
Test Applied to n Assumptions
Descriptive distribution characterization (mean, SD) of Cα-Cα distances Figure 1a, top100H database distances ~100 proteins, 16,162 pairs of consecutive Cα distances not stated
Specificity calculation (TN/(FP+TN)) across candidate cutoff values Figure 1c, cutoff selection for RDCC threshold ~100 top100H proteins not stated
Threshold-based classification (fixed RDCC cutoff of 0.012 Å) applied to decoy sets Figures 1d, 2; Decoys 'R' Us and I-TASSER CASP8 decoy suite 121 I-TASSER testcases; >500 decoys per protein in fisa set; multiple decoy sets (hg_structal, misfold, 4state_reduced, fisa) not stated
Mean and SD summary by resolution bin Table 1, PISCES-derived non-homologous structure sets 165, 682, 422, 753, 269 structures per resolution bin not stated
Approaches that could also have been used
  • A single cutoff value (0.012 Å) was chosen based on a specificity curve (Figure 1c) without accompanying sensitivity trade-off statistics at each point.
    Could also: A full ROC curve with AUC could also be used — Reporting sensitivity alongside specificity across the full range of thresholds would convey the trade-off more completely and allow readers to see how cutoff choice affects both false negatives and false positives.
  • Native versus decoy RDCC values are compared visually/qualitatively (Figures 1d, 2) rather than via a formal statistical test of difference.
    Could also: A nonparametric test such as the Mann-Whitney U test, or a t-test if distributional assumptions are met, could also be applied between native and decoy RDCC distributions — A formal test would provide a quantitative measure (e.g., a p-value or effect size) of how distinguishable the two groups are, complementing the threshold-based classification already used.
  • The claim that Cα-Cα distances are normally distributed (Figure 1a) is presented as a visual/qualitative observation.
    Could also: A formal normality test (e.g., Shapiro-Wilk or Kolmogorov-Smirnov) could also be applied — A formal test would give a statistical basis for the normality claim underlying the choice of mean and SD as the reference-state parameters.
  • Table 1 reports mean and SD of RDCC across resolution bins and concludes RDCC is largely independent of resolution.
    Could also: A correlation coefficient (e.g., Pearson or Spearman) between resolution and RDCC, or a regression analysis, could also be computed — A quantitative correlation statistic would formally support the stated absence of a relationship between resolution and RDCC, beyond the qualitative comparison of binned means.
  • Sample sizes for the top100H and decoy analyses are given as raw counts without an accompanying power or precision analysis.
    Could also: A confidence interval around the estimated mean and SD of the reference distance (3.8 Å, 0.04 Å) could also be reported — Confidence intervals would communicate the precision of these reference parameters, which underlie the downstream cutoff selection.
  • The RDCC threshold is applied uniformly across several distinct decoy sets (fisa, hg_structal, misfold, 4state_reduced, I-TASSER CASP8) as separate evaluations.
    Could also: A pooled or meta-analytic comparison with a multiple-testing correction (e.g., Bonferroni or Benjamini-Hochberg) across decoy sets could also be used — This would formally account for testing the same threshold across multiple independent datasets when summarizing overall discriminative performance.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

casp8_outlier_count
Reported
5/121 CASP8 I-TASSER native structures exceed RDCC cutoff 0.012 A: T0492=0.013, T0476=0.014, T0419=0.025, T0470=0.051, T0423=0.09 (paper notes T0470/T0423 are numbering errors, corrected to 0.001/0.002)
Reproduced
1/121 exceeds cutoff: T0476=0.01372 (matches). T0470=0.00098, T0423=0.00193 match paper's corrected values. T0492=0.00143, T0419=0.00187 do not reproduce as outliers (unexplained). mean=0.00203, median=0.00186 across all 121.
partial
1cnv_anomalies
Reported
Ile33/Ser34=4 A, Ser34/Phe35=3 A, Pro56/Ser57=4 A, Trp265/Asn266=3.4 A
Reproduced
4.0 A, 3.0 A, 4.0 A, 3.4 A -- exact match on all four pairs
exact
1ads_violation
Reported
"violation between two cis-prolines", Glu223-Asp24 (no numeric distance given in paper)
Reproduced
Glu223-Asp224 = 3.5 A (deviation 0.3 A, at hard-rejection boundary; weak signal, no paper value to confirm against)
partial
2jli_cleaved_bond
Reported
Asn263-Pro264 = 9.4 A
Reproduced
9.4 A exact, largest deviation in structure by far (next largest 0.5 A)
exact
decoysrus_fisa_1fc2
Reported
1FC2 native Ile12/Leu13 = 3.8 A; decoy AXPROA00-MIN = 4.1 A
Reproduced
native = 3.775 A (rounds to 3.8), decoy = 4.061 A (rounds to 4.1); RDCC native=0.00597 (passes), decoy=0.02339 (flagged, ~2x cutoff)
exact
top100h_reference_state
Reported
mean=3.8 A, SD=0.04 A; 88%/8%/3%/0.1% of 16162 pairs at 3.8/3.9/3.7/outlier A
Reproduced
[Top8000 substitute, top100H unobtainable] mean=3.809 A, SD=0.0236 A; 94.59%/4.00%/1.11%/0.28% at 3.8/3.9/3.7/rejected of 33838 pairs, n=140/140 structures
within tolerance
pisces_table1_resolution_bins
Reported
(mean/SD x1e-3 A) 1.0A: N165 3.0/4.3; 1.5A: N682 1.9/2.7; 2.0A: N753 1.8/2.8; 2.5A: N422 2.2/2.6; 3.0A: N269 2.2/2.5 (paper's own trend is non-monotonic, highest mean+SD at BEST resolution bin)
Reproduced
[2026 PISCES substitute, 2013 snapshot unrecoverable] 1.0A: N147 1.89/1.32; 1.5A: N147 1.54/1.48; 2.0A: N150 1.63/1.61; 2.5A: N149 1.84/2.73; 3.0A: N149 1.75/3.24 (x1e-3 A) -- SD increases monotonically with worse resolution, opposite direction from paper's own table; order of magnitude matches
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 71/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

The RDCC/CADistScore method itself reproduces convincingly: ported line-for-line from the actual Perl source, it hits 1CNV (4.0/3.0/4.0/3.4 Å), 2JLI (Asn263-Pro264 = 9.4 Å) and the 1FC2 native-vs-decoy pair (3.775→3.8 Å vs 4.061→4.1 Å, RDCC 0.00597 vs 0.02339) exactly, and 121/121 CASP8 natives cluster far below the 0.012 Å cutoff (mean 0.00203). The deviations sit almost entirely on the input/availability side: top100H is dead and the 2013 PISCES snapshot unrecoverable, so both had to be replaced by labelled substitutes (140 Top8000 chains; 2026 PISCES, 742/2291 chains), and post-2014 PDB remediation explains why T0470/T0423 now return the paper's own corrected 0.001/0.002. Two residuals remain on the authors' side: T0492 and T0419 are reported as outliers (0.013, 0.025) but score 0.00143/0.00187 with no correction noted, and the paper's Table 1 shows the worst mean+SD at the best resolution bin — a counterintuitive pattern that reproduction inverts. Severity is moderate (headline count 5/121 → 1/121) but the central conclusion is unaffected, hence yellow overall rather than red.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.