DNA binding analysis of rare variants in homeodomains reveals homeodomain specificity-determining residues.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Same input data as the authors
- ✓No relevant deviation in data/preprocessing
- ✓Any deviation was negligible
- 🟡Reported values were only indirectly comparable
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (1:1, directional + relative). The authors' own tool, the upbm R/Bioconductor package (github.com/pkimes/upbm @019977a), was run end-to-end on the authors' own HOXC9 uPBM data shipped by the companion upbmData package (hoxc9alexa/hoxc9cy3 = WT HOXC9 + the three rare allelic variants R193K/K195R/R222W; GSE233827) -- P16-valid. Pipeline: upbmPreprocess (Cy3 norm + bg subtract + spatial adjust + filter) -> probeFit(stratify=condition) -> kmerFit(8-mer, baseline HOXC9-REF) -> kmerTestAffinity/Contrast/Specificity. Deterministic (fixed .rda inputs, no RNG). RESULT: on the 4772 REF-preferential 8-mers (affinityQ<1e-6), R222W has mean contrastDifference -0.819 with 100% of 8-mers reduced and 4857/32896 8-mers significantly differential (contrastQ<1e-6; 83.0% of REF-preferential), whereas R193K and K195R each have ZERO significant differential 8-mers. This reproduces both paper claims: (C1) R222W strongly reduces affinity, and (C2) R222W is overwhelmingly the most affinity-altering of the three substitutions while the others are mild -- exactly the paper's >15%-of-REF-preferential affinity-altering criterion (only R222W qualifies). Env note: had to pin R 4.0 + Bioconductor 3.12 (upbm's 2020 era) -- modern Bioc 3.18 broke upbm's assay()/DataFrameList API -- and install CRAN-archived NormalGamma 1.1 from the CRAN Archive. Ran on «our HPC» compute node n093 (SLURM «job», prior «job» failed on those two env issues). NOT attempted: whole-study reanalysis of all GSE233827 TFs from raw GPR, and wet-lab/structural/conservation analyses (out of scope -- not pipeline-derived). Grades provisional; a human reviewer signs off.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-15 ⛓ 24e145a96f75
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether rare and disease-associated human homeodomain (HD) missense variants alter intrinsic DNA binding affinity and/or specificity relative to their reference alleles, and whether systematically surveying such variants can reveal previously unknown HD specificity-determining residues.
- ★ Many of the 92 assayed HD missense variants alter DNA binding affinity and/or specificity compared to their corresponding reference alleles finding
- ★ Detailed biochemical analysis and structural modeling identify 14 previously unknown specificity-determining positions in HDs, 5 of which do not contact DNA finding
- ★ The same missense substitution at analogous positions within different HDs often has different effects on DNA binding activity finding
- ★ Variant effect prediction tools perform moderately well at distinguishing variants with altered DNA binding affinity but perform poorly at distinguishing variants with altered binding specificity finding
- ★ The authors developed 'upbm,' a parametric statistical method for analyzing universal PBM data that identifies altered DNA binding affinity and specificity, validated against published assays and HOXD13 ChIP-seq data method
- ★ Variants with altered DNA binding activity are enriched for pathogenicity annotations in the ClinVar and ADDRESS databases finding
- Lower-affinity HD binding sites whose binding is altered by coding variants are occupied in vivo finding
- ★ The study generated a PBM dataset of nearly 8 million unique HD-8mer binding evaluations across 122 HD alleles (92 variants plus reference alleles) resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| universal protein binding microarray (PBM) | 92 human homeodomain missense variants and corresponding reference ('wildtype') HD alleles, recombinant protein in vitro | missense variant vs. reference allele | relative DNA binding affinity and specificity across all 8-bp sequences (8-mers) | universal PBM (custom array with all 8-bp sequences represented 32x/16x) |
| ChIP-seq | HOXD13 wildtype, Q325R, and Q325K mutant alleles (HD50 position) | HD50 missense mutation (Q325R/Q325K) | genomic occupancy sites | — |
| electrophoretic mobility shift assay (EMSA) | S. cerevisiae paralogous zinc finger TF DNA-binding domains: Msn2, Msn4, Com2, Usv1, Rgm1 | paralog comparison / DBD mutations | alternate DNA binding sequence preference | — |
| in silico variant effect prediction | amino acid sequences of assayed HD variants (computational) | missense variant | predicted deleteriousness/pathogenicity score, evaluated by AUROC for distinguishing variants with altered affinity or specificity | 42 variant interpretation tools (e.g., MutationAssessor, MetaLR, ClinPred, SIFT, AlphaMissense, CADD, MutationTaster, PhastCons, MetaSVM) |
| population/clinical variant database survey | human HD-containing transcription factors (221 HDs in gnomAD; HD Pfam domain variants in ClinVar) | none | count of unique missense variants per HD position | gnomAD; ClinVar |
- – Of 92 assayed HD variants, 51 showed altered DNA binding affinity and 28 showed altered specificity (17 overlapping) 51/92 affinity; 28/92 specificity
- – Identified 14 previously unknown specificity-determining positions, 5 of which do not contact DNA 14 positions (5 non-DNA-contacting)
- ▲ Variants with altered DNA binding activity were enriched for pathogenicity in ClinVar/ADDRESS databases P = 0.0203 (two-tailed Fisher's exact test)
- – MutationAssessor was the top-performing tool distinguishing variants with altered DNA binding affinity AUROC = 0.86
- ▼ The best-performing tool (MetaSVM) for distinguishing specificity-altering variants performed far worse than for affinity AUROC = 0.66
- ▼ 22 of 29 assayed known pathogenic/disease-associated HD mutations showed reduced DNA binding affinity 22/29
- ▲ A single rare variant, NKX2-4 R246Q, showed increased DNA binding affinity rather than loss
- ▼ 19 of 24 published loss-of-binding assays were reproduced by upbm as highly significant negative contrast differences, validating the method 19/24, P < 10^-4
- pvalue P = 0.0203 (enrichment of pathogenicity among variants with altered DNA binding activity (Fisher's exact test))
- other AUROC = 0.86 (MutationAssessor performance distinguishing altered-affinity variants (top performer))
- other AUROC = 0.66 (MetaSVM performance distinguishing altered-specificity variants (best of 42 tools))
- count 4,719 unique variants across 221 HDs (gnomAD survey of HD missense variation across 141,456 individuals)
- count 1232 missense variants (HD missense variants identified in ClinVar database)
- pvalue P < 10^-4 (19 of 24 published loss-of-binding experiments showed significant negative contrast differences by upbm)
- count 51 variants with altered affinity; 28 with altered specificity; 17 in both (summary of DNA binding activity changes across all assayed HD variants)
- other MAF 3.2 x 10^-5 (allele frequency of rare variant NKX2-6 R150C showing diminished binding affinity)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study assayed 122 homeodomain (HD) alleles (92 missense variants and 30 reference alleles) in technical duplicate on universal protein binding microarrays (PBMs), generating ~8 million unique HD–8-mer binding evaluations. A newly developed parametric statistical pipeline ('upbm') assigned three Q-values (affinityQ, contrastQ, specificityQ) per 8-mer per variant–reference comparison, using Q < 10⁻⁶ per-8-mer and 5% FDR variant-level thresholds to classify variants as having altered affinity or specificity. Enrichment of clinical pathogenicity annotations among altered-binding variants was tested with a two-tailed Fisher's exact test, and the discriminative ability of 42 variant effect prediction tools was evaluated by AUROC.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Parametric Q-value scoring (upbm): three Q-values per 8-mer (affinityQ, contrastQ, specificityQ); Q < 10⁻⁶ threshold per 8-mer; 5% FDR threshold for variant-level classification | All 92 variant vs. corresponding reference HD allele comparisons across ~8 million 8-mer binding evaluations | ~8 million unique HD–8-mer evaluations; 92 variants across 30 allelic series | not stated |
| Two-tailed Fisher's exact test | Enrichment of ClinVar/ADDRESS pathogenicity annotations among variants with altered DNA binding activity | 92 variants total (51 with altered affinity, 28 with altered specificity) | not stated |
| AUROC (area under receiver operating characteristic curve) | Performance of 42 variant effect prediction tools in discriminating variants with altered affinity, altered specificity, or either | 46 variants with altered affinity, 23 with altered specificity (17 in both sets), from 92 total variants | na |
| B-spline trendline fit | Specificity plots: fitting a trendline to contrast difference vs. reference affinity to compute specificityQ deviations | — | not stated |
-
Technical replicates (at least two PBMs per allele) were used to estimate measurement variability for the ~8 million 8-mer binding evaluations↳ Could also: Biological replicates (independent protein preparations) could also be included alongside technical replicates — Biological replicates capture protein preparation and folding variability in addition to measurement noise, providing a broader estimate of reproducibility and potentially more generalizable binding estimates across independently produced protein batches
-
Individual 8-mer binding changes were assessed with a parametric Q-value approach (upbm), with variant-level calls based on the count of significantly altered 8-mers meeting a 5% FDR threshold calibrated from reference-vs-reference controls↳ Could also: A hierarchical or mixed-effects model fit across all 8-mers simultaneously could also be applied, borrowing strength across groups of related sequences — A hierarchical approach would explicitly account for the correlation structure among overlapping 8-mer sequences, potentially improving sensitivity for low-affinity 8-mers and providing a single unified test per variant rather than a two-stage count-then-FDR procedure
-
Enrichment of pathogenicity annotations among variants with altered DNA binding was assessed with a single two-tailed Fisher's exact test (P = 0.0203)↳ Could also: Logistic regression could also model pathogenicity as a function of binding alteration type (affinity vs. specificity vs. both) while adjusting for variant source (ClinVar, ADDRESS, gnomAD) — Logistic regression would allow simultaneous adjustment for multiple variant characteristics and could distinguish whether affinity changes and specificity changes independently associate with pathogenicity annotation, rather than collapsing both into a single enrichment test
-
Prediction tool performance was compared across 42 tools using AUROC as the primary metric↳ Could also: Area under the precision-recall curve (AUPRC) could serve as an additional or co-primary performance metric (precision-recall curves were reported in supplementary figures) — AUPRC is often preferred when positive cases (variants with altered binding) are a minority of the tested set, as it is more sensitive to performance on the positive class than AUROC under class imbalance
-
The upbm specificityQ score is based on deviation from a B-spline trendline fitted to contrast difference vs. reference affinity↳ Could also: Locally weighted smoothing (LOESS) or a generalized additive model (GAM) could also be used to estimate the trendline — LOESS and GAMs offer alternative flexibility-bias tradeoffs with automatic smoothing parameter selection, and their residual distributions may be amenable to different parametric Q-value formulations
-
Variant-level false discovery rate was calibrated empirically using negative control reference-vs-reference replicate comparisons to set the 5% FDR threshold↳ Could also: A permutation-based null distribution constructed by randomly re-pairing variant and reference allele PBM datasets could also calibrate the variant-level threshold — Permutation nulls do not rely on reference-vs-reference replicates as a surrogate for variant-vs-reference noise, and may better reflect the empirical null when variant effects are sparse and the reference-vs-reference noise structure differs from the variant-vs-reference comparison
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
3 of 8 ClinVar variants of uncertain significance in homeodomains showed reduced DNA binding affinity by PBM.microarray human homeodomain down 2024×1papers★ This paper is the founder (earliest)
-
16 of 32 homeodomain missense variants present only in gnomAD showed reduced DNA binding affinity by PBM.microarray human homeodomain down 2024×1papers★ This paper is the founder (earliest)
-
19 of 24 published homeodomain loss-of-binding experiments were confirmed as significant negative contrast differences by universal PBM (upbm), validating the assay.microarray human homeodomain down 2024×1papers★ This paper is the founder (earliest)
-
PBM screening of homeodomain variants identified 14 previously unknown specificity-determining positions, 5 of which make no direct DNA contact.microarray human homeodomain 2024×1papers★ This paper is the founder (earliest)
-
92 homeodomain missense variants tested by PBM revealed 51 with altered DNA binding affinity and 28 with altered specificity.microarray human homeodomain mixed 2024×1papers★ This paper is the founder (earliest)
-
NKX2-4 R246Q, a rare gnomAD variant at a non-DNA-contacting position, increased DNA binding affinity relative to reference by PBM.microarray human homeodomain up 2024×1papers★ This paper is the founder (earliest)
-
22 of 29 pathogenic/disease-associated homeodomain missense variants showed reduced DNA binding affinity by PBM.microarray human homeodomain down 2024×1papers★ This paper is the founder (earliest)
-
MutationAssessor achieved the highest AUROC (0.86) for discriminating affinity-altered homeodomain variants; no tool exceeded AUROC 0.66 for specificity-altered variants.other human homeodomain mixed 2024×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-38600112
Paper: Kock et al. 2024, Nat Commun 15:3110. "DNA binding analysis of rare variants in homeodomains reveals homeodomain specificity-determining residues." PMID 38600112 · PMC11006913 · DOI 10.1038/s41467-024-47396-0
Authors' code: https://github.com/pkimes/upbm (R/Bioconductor package upbm)
- pinned: upbm @
019977afd48c75534e5ce5e87e9d2fcfe46da53e - data: upbmData @
d2f4225db9860c2d71fb9cb6284eec4527049b0e - arrays: upbmAux @
31b7599e9e9d851d8029d1ed5597f56b68206e3e
Data: GEO GSE233827 (full study, raw GPR scans). The companion upbmData
package ships the paper's own HOXC9 allelic-variant uPBM data as analysis-ready
SummarizedExperiment objects: hoxc9alexa (Alexa488 scans, multiple PMT gains),
hoxc9cy3 (Cy3 scans), for wild-type HOXC9 and three rare allelic variants
R193K, K195R, R222W. Source: Dropbox PBM-HOXC9.zip, loaded via
upbm::gpr2PBMExperiment. This is identical-provenance data behind the paper's
HOXC9 result — using it satisfies HARD RULE 2 (P16: the authors' own tool on the
paper's own data).
In scope (pipeline-derived, deterministic)
The full upbm analysis pipeline, end to end, on the HOXC9 data:
upbmPreprocess— Cy3 normalization, background subtract, spatial adjust, filter/trim probes (probe-level).probeFit— probe-level replicate model, stratified by condition.kmerFit— 8-mer (k=8) affinity + variance estimates, baseline = HOXC9-REF.- Inference:
kmerTestAffinity(preferential),kmerTestContrast(differential affinity vs REF),kmerTestSpecificity(differential specificity vs REF).
This pipeline IS the method the paper applies to every TF/variant. It is fully deterministic (fixed input .rda, no RNG seeds needed in the core estimators).
Specific reported claim targeted
Discussion + Supplementary Fig. 6c,d: "Arg-to-Trp substitution at canonical
position 31 resulted in strongly reduced affinity in HOXC9" (R222W), in
contrast to milder effects of other substitutions. The paper's significance
thresholds (Methods): differential affinity contrastQ < 1e-6; preferential
affinity affinityQ < 1e-6; an allele is "affinity-altering" if >15% of the
REF-preferential 8-mers are differentially bound.
Reproduction question: Does the upbm pipeline, run on the shipped HOXC9 data, show R222W with strongly reduced 8-mer affinity relative to HOXC9-REF, and is R222W the most severe of the three variants? (directional + relative-magnitude 1:1 check, plus the quickstart's deterministic top differential 8-mers.)
Out of scope (not attempted)
- Whole-study reanalysis of all «path» of homeodomain TFs in GSE233827 (raw GPR → the same pipeline at scale) — the hard 80→100% tail; data volume + per-TF curation. We reproduce the representative HOXC9 unit that backs the headline rare-variant claim.
- Wet-lab / structural / evolutionary-conservation analyses (out of scope by definition — not pipeline-derived).
- Exact byte match of Supplementary Fig. 6 panels (figure rendering, not a value).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.