Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

DNA binding analysis of rare variants in homeodomains reveals homeodomain specificity-determining residues.

Nat Commun · 2024
L1 93/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
93/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 85% of all assessed papers rank 154 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (1:1, directional + relative). The authors' own tool, the upbm R/Bioconductor package (github.com/pkimes/upbm @019977a), was run end-to-end on the authors' own HOXC9 uPBM data shipped by the companion upbmData package (hoxc9alexa/hoxc9cy3 = WT HOXC9 + the three rare allelic variants R193K/K195R/R222W; GSE233827) -- P16-valid. Pipeline: upbmPreprocess (Cy3 norm + bg subtract + spatial adjust + filter) -> probeFit(stratify=condition) -> kmerFit(8-mer, baseline HOXC9-REF) -> kmerTestAffinity/Contrast/Specificity. Deterministic (fixed .rda inputs, no RNG). RESULT: on the 4772 REF-preferential 8-mers (affinityQ<1e-6), R222W has mean contrastDifference -0.819 with 100% of 8-mers reduced and 4857/32896 8-mers significantly differential (contrastQ<1e-6; 83.0% of REF-preferential), whereas R193K and K195R each have ZERO significant differential 8-mers. This reproduces both paper claims: (C1) R222W strongly reduces affinity, and (C2) R222W is overwhelmingly the most affinity-altering of the three substitutions while the others are mild -- exactly the paper's >15%-of-REF-preferential affinity-altering criterion (only R222W qualifies). Env note: had to pin R 4.0 + Bioconductor 3.12 (upbm's 2020 era) -- modern Bioc 3.18 broke upbm's assay()/DataFrameList API -- and install CRAN-archived NormalGamma 1.1 from the CRAN Archive. Ran on «our HPC» compute node n093 (SLURM «job», prior «job» failed on those two env issues). NOT attempted: whole-study reanalysis of all GSE233827 TFs from raw GPR, and wet-lab/structural/conservation analyses (out of scope -- not pipeline-derived). Grades provisional; a human reviewer signs off.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-15 ⛓ 24e145a96f75
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether rare and disease-associated human homeodomain (HD) missense variants alter intrinsic DNA binding affinity and/or specificity relative to their reference alleles, and whether systematically surveying such variants can reveal previously unknown HD specificity-determining residues.

Core claims
  • Many of the 92 assayed HD missense variants alter DNA binding affinity and/or specificity compared to their corresponding reference alleles finding
  • Detailed biochemical analysis and structural modeling identify 14 previously unknown specificity-determining positions in HDs, 5 of which do not contact DNA finding
  • The same missense substitution at analogous positions within different HDs often has different effects on DNA binding activity finding
  • Variant effect prediction tools perform moderately well at distinguishing variants with altered DNA binding affinity but perform poorly at distinguishing variants with altered binding specificity finding
  • The authors developed 'upbm,' a parametric statistical method for analyzing universal PBM data that identifies altered DNA binding affinity and specificity, validated against published assays and HOXD13 ChIP-seq data method
  • Variants with altered DNA binding activity are enriched for pathogenicity annotations in the ClinVar and ADDRESS databases finding
  • Lower-affinity HD binding sites whose binding is altered by coding variants are occupied in vivo finding
  • The study generated a PBM dataset of nearly 8 million unique HD-8mer binding evaluations across 122 HD alleles (92 variants plus reference alleles) resource
Experimental setups
Assay System Perturbation Readout Platform
universal protein binding microarray (PBM) 92 human homeodomain missense variants and corresponding reference ('wildtype') HD alleles, recombinant protein in vitro missense variant vs. reference allele relative DNA binding affinity and specificity across all 8-bp sequences (8-mers) universal PBM (custom array with all 8-bp sequences represented 32x/16x)
ChIP-seq HOXD13 wildtype, Q325R, and Q325K mutant alleles (HD50 position) HD50 missense mutation (Q325R/Q325K) genomic occupancy sites
electrophoretic mobility shift assay (EMSA) S. cerevisiae paralogous zinc finger TF DNA-binding domains: Msn2, Msn4, Com2, Usv1, Rgm1 paralog comparison / DBD mutations alternate DNA binding sequence preference
in silico variant effect prediction amino acid sequences of assayed HD variants (computational) missense variant predicted deleteriousness/pathogenicity score, evaluated by AUROC for distinguishing variants with altered affinity or specificity 42 variant interpretation tools (e.g., MutationAssessor, MetaLR, ClinPred, SIFT, AlphaMissense, CADD, MutationTaster, PhastCons, MetaSVM)
population/clinical variant database survey human HD-containing transcription factors (221 HDs in gnomAD; HD Pfam domain variants in ClinVar) none count of unique missense variants per HD position gnomAD; ClinVar
Key results
  • Of 92 assayed HD variants, 51 showed altered DNA binding affinity and 28 showed altered specificity (17 overlapping) 51/92 affinity; 28/92 specificity
  • Identified 14 previously unknown specificity-determining positions, 5 of which do not contact DNA 14 positions (5 non-DNA-contacting)
  • Variants with altered DNA binding activity were enriched for pathogenicity in ClinVar/ADDRESS databases P = 0.0203 (two-tailed Fisher's exact test)
  • MutationAssessor was the top-performing tool distinguishing variants with altered DNA binding affinity AUROC = 0.86
  • The best-performing tool (MetaSVM) for distinguishing specificity-altering variants performed far worse than for affinity AUROC = 0.66
  • 22 of 29 assayed known pathogenic/disease-associated HD mutations showed reduced DNA binding affinity 22/29
  • A single rare variant, NKX2-4 R246Q, showed increased DNA binding affinity rather than loss
  • 19 of 24 published loss-of-binding assays were reproduced by upbm as highly significant negative contrast differences, validating the method 19/24, P < 10^-4
Key statistics
  • pvalue P = 0.0203 (enrichment of pathogenicity among variants with altered DNA binding activity (Fisher's exact test))
  • other AUROC = 0.86 (MutationAssessor performance distinguishing altered-affinity variants (top performer))
  • other AUROC = 0.66 (MetaSVM performance distinguishing altered-specificity variants (best of 42 tools))
  • count 4,719 unique variants across 221 HDs (gnomAD survey of HD missense variation across 141,456 individuals)
  • count 1232 missense variants (HD missense variants identified in ClinVar database)
  • pvalue P < 10^-4 (19 of 24 published loss-of-binding experiments showed significant negative contrast differences by upbm)
  • count 51 variants with altered affinity; 28 with altered specificity; 17 in both (summary of DNA binding activity changes across all assayed HD variants)
  • other MAF 3.2 x 10^-5 (allele frequency of rare variant NKX2-6 R150C showing diminished binding affinity)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study assayed 122 homeodomain (HD) alleles (92 missense variants and 30 reference alleles) in technical duplicate on universal protein binding microarrays (PBMs), generating ~8 million unique HD–8-mer binding evaluations. A newly developed parametric statistical pipeline ('upbm') assigned three Q-values (affinityQ, contrastQ, specificityQ) per 8-mer per variant–reference comparison, using Q < 10⁻⁶ per-8-mer and 5% FDR variant-level thresholds to classify variants as having altered affinity or specificity. Enrichment of clinical pathogenicity annotations among altered-binding variants was tested with a two-tailed Fisher's exact test, and the discriminative ability of 42 variant effect prediction tools was evaluated by AUROC.

Replicationtechnical Sample size92 missense variants across 30 HD allelic series; 122 HD alleles total; each assayed in at least duplicate on PBMs GroupsVariant HD alleles vs. corresponding reference (wildtype) HD allele; pathogenic/disease-associated vs. VUS/gnomAD-only variants Pairingpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionFDR (Q-values); Q < 10⁻⁶ per 8-mer; 5% FDR for variant-level calls
Statistical tests used
Test Applied to n Assumptions
Parametric Q-value scoring (upbm): three Q-values per 8-mer (affinityQ, contrastQ, specificityQ); Q < 10⁻⁶ threshold per 8-mer; 5% FDR threshold for variant-level classification All 92 variant vs. corresponding reference HD allele comparisons across ~8 million 8-mer binding evaluations ~8 million unique HD–8-mer evaluations; 92 variants across 30 allelic series not stated
Two-tailed Fisher's exact test Enrichment of ClinVar/ADDRESS pathogenicity annotations among variants with altered DNA binding activity 92 variants total (51 with altered affinity, 28 with altered specificity) not stated
AUROC (area under receiver operating characteristic curve) Performance of 42 variant effect prediction tools in discriminating variants with altered affinity, altered specificity, or either 46 variants with altered affinity, 23 with altered specificity (17 in both sets), from 92 total variants na
B-spline trendline fit Specificity plots: fitting a trendline to contrast difference vs. reference affinity to compute specificityQ deviations not stated
Approaches that could also have been used
  • Technical replicates (at least two PBMs per allele) were used to estimate measurement variability for the ~8 million 8-mer binding evaluations
    Could also: Biological replicates (independent protein preparations) could also be included alongside technical replicates — Biological replicates capture protein preparation and folding variability in addition to measurement noise, providing a broader estimate of reproducibility and potentially more generalizable binding estimates across independently produced protein batches
  • Individual 8-mer binding changes were assessed with a parametric Q-value approach (upbm), with variant-level calls based on the count of significantly altered 8-mers meeting a 5% FDR threshold calibrated from reference-vs-reference controls
    Could also: A hierarchical or mixed-effects model fit across all 8-mers simultaneously could also be applied, borrowing strength across groups of related sequences — A hierarchical approach would explicitly account for the correlation structure among overlapping 8-mer sequences, potentially improving sensitivity for low-affinity 8-mers and providing a single unified test per variant rather than a two-stage count-then-FDR procedure
  • Enrichment of pathogenicity annotations among variants with altered DNA binding was assessed with a single two-tailed Fisher's exact test (P = 0.0203)
    Could also: Logistic regression could also model pathogenicity as a function of binding alteration type (affinity vs. specificity vs. both) while adjusting for variant source (ClinVar, ADDRESS, gnomAD) — Logistic regression would allow simultaneous adjustment for multiple variant characteristics and could distinguish whether affinity changes and specificity changes independently associate with pathogenicity annotation, rather than collapsing both into a single enrichment test
  • Prediction tool performance was compared across 42 tools using AUROC as the primary metric
    Could also: Area under the precision-recall curve (AUPRC) could serve as an additional or co-primary performance metric (precision-recall curves were reported in supplementary figures) — AUPRC is often preferred when positive cases (variants with altered binding) are a minority of the tested set, as it is more sensitive to performance on the positive class than AUROC under class imbalance
  • The upbm specificityQ score is based on deviation from a B-spline trendline fitted to contrast difference vs. reference affinity
    Could also: Locally weighted smoothing (LOESS) or a generalized additive model (GAM) could also be used to estimate the trendline — LOESS and GAMs offer alternative flexibility-bias tradeoffs with automatic smoothing parameter selection, and their residual distributions may be amenable to different parametric Q-value formulations
  • Variant-level false discovery rate was calibrated empirically using negative control reference-vs-reference replicate comparisons to set the 5% FDR threshold
    Could also: A permutation-based null distribution constructed by randomly re-pairing variant and reference allele PBM datasets could also calibrate the variant-level threshold — Permutation nulls do not rely on reference-vs-reference replicates as a surrogate for variant-vs-reference noise, and may better reflect the empirical null when variant effects are sparse and the reference-vs-reference noise structure differs from the variant-vs-reference comparison
Software: upbm (custom package developed in this study)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
25
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

3RKQ PDBe in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
9ANT PDBe in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
PF00046 Pfam in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-38600112

Paper: Kock et al. 2024, Nat Commun 15:3110. "DNA binding analysis of rare variants in homeodomains reveals homeodomain specificity-determining residues." PMID 38600112 · PMC11006913 · DOI 10.1038/s41467-024-47396-0

Authors' code: https://github.com/pkimes/upbm (R/Bioconductor package upbm)

  • pinned: upbm @ 019977afd48c75534e5ce5e87e9d2fcfe46da53e
  • data: upbmData @ d2f4225db9860c2d71fb9cb6284eec4527049b0e
  • arrays: upbmAux @ 31b7599e9e9d851d8029d1ed5597f56b68206e3e

Data: GEO GSE233827 (full study, raw GPR scans). The companion upbmData package ships the paper's own HOXC9 allelic-variant uPBM data as analysis-ready SummarizedExperiment objects: hoxc9alexa (Alexa488 scans, multiple PMT gains), hoxc9cy3 (Cy3 scans), for wild-type HOXC9 and three rare allelic variants R193K, K195R, R222W. Source: Dropbox PBM-HOXC9.zip, loaded via upbm::gpr2PBMExperiment. This is identical-provenance data behind the paper's HOXC9 result — using it satisfies HARD RULE 2 (P16: the authors' own tool on the paper's own data).

In scope (pipeline-derived, deterministic)

The full upbm analysis pipeline, end to end, on the HOXC9 data:

  1. upbmPreprocess — Cy3 normalization, background subtract, spatial adjust, filter/trim probes (probe-level).
  2. probeFit — probe-level replicate model, stratified by condition.
  3. kmerFit — 8-mer (k=8) affinity + variance estimates, baseline = HOXC9-REF.
  4. Inference: kmerTestAffinity (preferential), kmerTestContrast (differential affinity vs REF), kmerTestSpecificity (differential specificity vs REF).

This pipeline IS the method the paper applies to every TF/variant. It is fully deterministic (fixed input .rda, no RNG seeds needed in the core estimators).

Specific reported claim targeted

Discussion + Supplementary Fig. 6c,d: "Arg-to-Trp substitution at canonical position 31 resulted in strongly reduced affinity in HOXC9" (R222W), in contrast to milder effects of other substitutions. The paper's significance thresholds (Methods): differential affinity contrastQ < 1e-6; preferential affinity affinityQ < 1e-6; an allele is "affinity-altering" if >15% of the REF-preferential 8-mers are differentially bound.

Reproduction question: Does the upbm pipeline, run on the shipped HOXC9 data, show R222W with strongly reduced 8-mer affinity relative to HOXC9-REF, and is R222W the most severe of the three variants? (directional + relative-magnitude 1:1 check, plus the quickstart's deterministic top differential 8-mers.)

Out of scope (not attempted)

  • Whole-study reanalysis of all «path» of homeodomain TFs in GSE233827 (raw GPR → the same pipeline at scale) — the hard 80→100% tail; data volume + per-TF curation. We reproduce the representative HOXC9 unit that backs the headline rare-variant claim.
  • Wet-lab / structural / evolutionary-conservation analyses (out of scope by definition — not pipeline-derived).
  • Exact byte match of Supplementary Fig. 6 panels (figure rendering, not a value).
Figures / tables: Fig. 6cFig. 6
C1
Reported
HOXC9 R222W (canonical HD position 31, Arg->Trp) shows strongly reduced DNA-binding affinity (Discussion + Supplementary Fig. 6c,d)
Reproduced
R222W mean contrastDifference = -0.819 (log2) on 4772 REF-preferential 8-mers; 100% of them reduced; 4857/32896 8-mers differentially bound (contrastQ<1e-6); 83.0% of REF-preferential 8-mers significant -> strongly reduced affinity
within tolerance
C2
Reported
R222W is the most affinity-altering of the three HOXC9 substitutions (R193K, K195R, R222W); the others are milder
Reproduced
R222W = 4857 differential 8-mers and 83.0% of REF-preferential 8-mers altered; R193K = 0 (0.0%); K195R = 0 (0.0%). Under the paper's >15%-of-REF-preferential rule only R222W is affinity-altering. Ordering reproduced exactly.
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 93/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

247.7 k
tokens (I/O) · 26.6 M incl. cache
129 min
runtime · 0.11 CPU-h
2.3 GB
peak RAM
3 (2 failed)
HPC jobs
hummel
machine