Missense variants in human forkhead transcription factors reveal determinants of forkhead DNA bispecificity.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡Reported values were only indirectly comparable
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (1:1 on structure, direction-matched on biology). The paper reports NO original code, so per P16 the described PBM pipeline was re-implemented from Methods and run on the paper's OWN deposited 8-mer E-score matrix (GEO GSE297784) via a 31s «our HPC» SLURM job (2208157). Dataset structure matches the paper EXACTLY: 260 PBM experiments, 32896 non-redundant 8-mers, 12 reference forkhead proteins, 83 variants (95 alleles - 12 refs). Bispecificity analysis (FKH=RYAAAYA vs FHL=GACGC, E>=0.4) classifies most references as monospecific and identifies FOXN2/FOXN3/FOXN4 as bispecific - exactly the FOXN subfamily the paper's variant case-studies center on. All four headline variant effects reproduce in the correct direction with high significance: FOXF1 Y89C and FOXB2 E52G strongly reduce FKH affinity (dE=-0.15/-0.34, p<=6e-8); FOXN2 T154A shifts toward FHL and FOXN3 L199V toward FKH (Mann-Whitney p<=7e-12 / 2e-9). Biological-effect claims are reported qualitatively in the paper so they are graded within-tol (claim reproduced, not an exact author number); structural counts are exact. NOT attempted: raw-.gpr->E-score re-derivation (proprietary MATLAB suite), gnomAD/ClinVar screening (external DBs), and ChIP-seq (not deposited in this series). All provisional - a human reviewer signs off.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 92assessed: 2026-06-19 ⛓ 1ca8e460aada
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-19
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests what determines whether a forkhead (FH) transcription factor domain binds only the FKH motif, only the FHL motif, or both (bispecificity), hypothesizing that non-DNA-contacting residues—particularly in the helix2-3 loop and wing 2—control this specificity.
- ★ Non-DNA-contacting residues, especially in the loop between helices 2 and 3 and in wing 2, control mono- vs. bispecificity of FH domains for the FKH and FHL motifs mechanism
- ★ PBM screening of 12 reference FH proteins, 61 naturally occurring missense variants, and 22 designed mutant/chimeric FH proteins reveals effects of coding variation on DNA-binding activity method
- ★ No tested missense variant switched a monospecific FH domain to the opposite motif specificity or converted it to bispecificity, consistent with a distributed model requiring both subdomains to be permissive finding
- ★ Subdomain-swap experiments show the 2-3 loop and wing 2 each independently influence FKH/FHL preference, but neither alone confers bispecificity mechanism
- ★ FOXN3 ChIP-seq peaks in HepG2 and MCF-7 cells containing FHL motifs are specifically enriched for cilium-related genes/phenotypes, while FKH-motif peaks show no consistent enrichment finding
- ★ 19 naturally occurring FH missense variants in disease-associated genes altered DNA-binding activity, 16 not previously reported finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Universal (all 10-mer) protein-binding microarray (PBM) | 12 reference human FH domains + 61 natural missense variants + 22 designed mutants (recombinant protein) | missense mutation / designed mutation | 8-mer E-scores reflecting DNA-binding specificity for FKH vs. FHL motifs | universal all 10-mer PBM |
| PBM on subdomain-swap chimeric constructs | Chimeric FH domains (FOXN3, FOXA1, FOXN1, FOXJ3 combinations; 2-3 loop, wing, and helix 1 swaps) | subdomain swap (2-3 loop, wings, helix 1 residues) | shift in FKH vs. FHL binding preference (E-scores) | universal PBM |
| ChIP-seq (reanalysis of published data) | FOXN3 in HepG2 and MCF-7 human cell lines | none (endogenous) | FKH/FHL motif enrichment at peaks; GO and human phenotype term enrichment via GREAT | GREAT tool |
| Structural/bioinformatic analysis of protein-DNA cocrystal structures | FOXK1, FOXC2, FOXG1, FOXN1, FOXN3 FH domains | none | intradomain wing2-helix1 contacts; recognition-helix asparagine rotamer position | — |
| Prior in vivo knockout phenotyping (cited) | Embryonic mouse retina, Foxn3 gene | Foxn3 knockout | dysregulation of cilia-associated genes; ciliary organization/morphology | — |
- ▼ FOXN2 L119P (proline in helix 1) causes near-complete loss of DNA binding, almost no 8-mers with E>0.4
- – FOXJ3 V162A negative control (allele frequency 0.806) shows no significant difference from reference in binding
- – 0/6 negative control variants showed unambiguous binding effects; 7/7 positive controls compromised binding
- ▼ FOXF1 Y89C (likely pathogenic, alveolar capillary dysplasia) shows severely reduced FKH-motif binding specificity
- ▲ FOXN2 T154A shows increased preference for FHL vs. FKH 8-mers
- ▼ FOXN2 P153L largely eliminates FHL-preferential binding while sparing FKH binding
- ▼ FOXN3 2-3 loop replaced with FOXA1's abrogates FHL recognition without compromising FKH binding
- ▲ FHL-motif-containing FOXN3 ChIP-seq peaks show 31 GO terms and 8 phenotype terms significantly enriched in common between HepG2 and MCF-7, all related to cilium function 31 GO terms; 8 phenotype terms
- count 2,402 (gnomAD variants identified within FH domain or flanking 15 positions)
- count 245 total (12 benign/likely benign, 130 pathogenic/likely pathogenic/risk factor, 103 VUS) (ClinVar variants within FH domain or flanking 15 positions)
- count 19 variants (16 novel) (naturally occurring variants in disease-associated genes showing altered DNA binding)
- other 0.806 (population allele frequency of FOXJ3 V162A negative control variant)
- count 0/6 (negative control variants showing unambiguous binding effects)
- count 31 (GO Biological Process terms significantly enriched in both HepG2 and MCF-7 FOXN3 FHL-motif ChIP-seq peaks)
- count 8 (human phenotype associations significantly enriched in common, all ciliary dysfunction-related)
- count 7 of top 10 (top enriched GO terms in MCF-7 FHL peaks shared with top 10 in HepG2 FHL peaks)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper describes a variant-effect screening study using protein-binding microarrays (PBMs) to compare DNA-binding enrichment (E-score) profiles of reference versus variant/mutant forkhead transcription factor domains, with results interpreted largely through scatterplot comparisons and qualitative descriptions of consistency across replicates (e.g., '2 independent replicates') rather than through named inferential statistical tests. A separate analysis examined enrichment of FKH/FHL DNA motifs within published ChIP-seq peaks relative to a matched genomic background, and used the GREAT tool to test for enrichment of Gene Ontology and human phenotype terms among peak-associated genes, reporting results as 'significantly enriched' without stating the specific test, correction method, or exact p-values in the provided text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Not explicitly named; comparison of PBM E-scores for reference vs. variant/mutant alleles via scatterplots and replicate consistency | Figures 2D, 2E, 3A-3E, 5A-5G, S1-S14 (variant and subdomain-swap binding comparisons) | 2 independent replicates stated for FOXN2 L119P; replicate numbers not consistently stated for other variants | not stated |
| Not explicitly named; motif enrichment in ChIP-seq peaks vs. matched genomic background | Figures 4A and 4B (FOXN3 ChIP-seq peaks, HepG2 and MCF-7 cell lines) | — | not stated |
| GREAT tool enrichment analysis (Gene Ontology / human phenotype term enrichment) | Figures 4C and 4D (GO Biological Process and phenotype terms among FKH- vs. FHL-motif-containing ChIP-seq peaks) | — | not stated |
-
Differences in PBM E-scores between reference and variant/mutant alleles are described qualitatively (e.g., 'consistently increased preference,' 'subtle but highly consistent shift') based on scatterplot comparisons across a small number of replicates.↳ Could also: A paired statistical test comparing matched k-mer E-scores between reference and variant alleles (e.g., a paired t-test or Wilcoxon signed-rank test) could also be used. — This would provide a quantitative statistic and p-value summarizing the magnitude and consistency of binding shifts, complementing the visual/qualitative comparison of scatterplots.
-
GO and human phenotype term enrichment among ChIP-seq peaks was assessed with the GREAT tool and described as 'significantly enriched' without stating the specific correction method applied across the many terms tested.↳ Could also: Explicitly reporting the multiple-testing correction (e.g., Benjamini-Hochberg FDR) and the significance thresholds used for GO/phenotype enrichment could also be included. — Stating the correction method clarifies how the false-discovery rate was controlled given the large number of GO/phenotype terms tested simultaneously.
-
Motif enrichment within ChIP-seq peaks was assessed relative to a matched genomic background without naming a specific statistical test in the provided text.↳ Could also: A hypergeometric test or Fisher's exact test (as implemented in tools such as HOMER or AME/MEME) with reported enrichment p- or q-values could also be used. — This would give a standard, reproducible quantitative measure of motif enrichment strength alongside the qualitative description.
-
Reproducibility of binding measurements was described through a small number of independent replicates (e.g., 2 replicates for FOXN2 L119P) assessed largely by visual consistency.↳ Could also: Reporting a replicate correlation statistic (e.g., Pearson or Spearman correlation coefficient with a confidence interval) between replicate E-score sets could also be used. — This would formally quantify reproducibility across replicates rather than relying on qualitative visual assessment of scatterplots.
-
Negative and positive control variant effects were summarized as simple counts (e.g., '0/6 negative control variants showed unambiguous effects on binding').↳ Could also: A binomial or exact test comparing the observed proportion of affected controls to an expected null rate could also be used. — This would provide a formal statistical basis for the control validation claim in addition to the reported counts.
-
No dispersion measures (e.g., SD, SEM, or confidence intervals) around E-scores or enrichment estimates are reported in the provided text.↳ Could also: Reporting a dispersion or interval estimate (e.g., SD across replicates or a bootstrap confidence interval on E-scores) could also be included alongside point estimates. — This would convey the variability of binding measurements, which is useful for readers assessing the reliability of small-magnitude specificity shifts.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 41124077 (Cell Reports 2025)
Title: Missense variants in human forkhead transcription factors reveal determinants of forkhead DNA bispecificity. DOI: 10.1016/j.celrep.2025.116427 · PMCID PMC12795473 · GEO GSE297784 · BioProject PRJNA1266190
Code/data availability (verbatim from paper)
"PBM data have been deposited in the GEO database under accession number GSE297784 and are publicly available as of the date of publication. This paper does not report original code. Any additional information required to reanalyze the data ... is available from the lead contact upon request."
=> The BRIEF's "Code: github.com/BulykLab/MEDEA" is a MIS-HARVESTED lab repo (MEDEA is the lab's 2020 chromatin-accessibility tool, Genome Research; it is NOT this paper's pipeline). Per P16, reproduction is done by re-running the described pipeline (third-party / standard tools + a from-description reimplementation of the bespoke analysis) on the paper's OWN deposited data. This is equally valid.
Data
GSE297784 = universal protein binding microarray (uPBM), Agilent 8x60K "all 10-mer" arrays (AMADID 030236), GST-fusion forkhead DBDs, in-vitro expressed, 750 nM. Deposited PROCESSED data (GSE297784_processed_data_Matrix.txt.gz) = the 8-mer E-score matrix: 32896 non-redundant 8-mers x 260 experiments. Sample titles encode GENE.allele_arrayID (e.g. FOXA1.ref_388, FOXB2.E52G_414).
IN SCOPE (pipeline-derived, reproducible from deposited E-score matrix)
S1. Dataset structure: N experiments=260; 8-mers=32896; reference proteins=12; variants=83 (95 alleles-12 refs). S2. Motif definitions: 8-mers matching FKH (RYAAAYA) vs FHL (GACGC), E>=0.4 threshold (paper Methods). S3. Bispecificity classification of the 12 reference FH proteins (binds FKH only / FHL only / both=bispecific). S4. Affinity-altering variant effects (paired/signed-rank Wilcoxon on dE = variant-ref over motif 8-mers): FOXF1 Y89C -> severe FKH reduction; FOXB2 E52G -> reduced FKH binding. S5. Specificity-altering variant effects (FKH-dE vs FHL-dE distributions): FOXN2 T154A -> increased FHL preference; FOXN3 L199V -> shifted toward FKH preference.
OUT OF SCOPE (wet-lab / external / not pipeline-derivable here)
- Site-directed mutagenesis, protein expression, the PBM wet-lab assay itself.
- gnomAD/ClinVar variant screening counts (2402/245) — external database queries, not from GSE297784.
- ChIP-seq (MACS2/Bowtie2/GREAT) — no ChIP accession in this GEO series; not deposited here.
- Re-deriving E-scores from raw .gpr (Universal PBM Analysis Suite, MATLAB-proprietary) — deposited E-scores are the canonical processed product; we start from them (re-derivation is a deeper stretch goal).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Clean reproduction. All structural counts (260 PBM experiments, 32896 8-mers, 12 references, 83 variants) match the paper exactly off its own deposited matrix (GSE297784), and all four headline variant effects reproduce in the correct direction with high significance (p from 6e-8 to 7e-12). The only caveats are on our/data side, not the authors': the biological claims are qualitative in the paper (graded on direction, not a number), the pipeline had to be re-implemented because no original code was shared, and some claims (gnomAD/ClinVar, ChIP-seq) are external/undeposited and out of scope. Severity is negligible and the central bispecificity conclusion holds fully.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.