Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Missense variants in human forkhead transcription factors reveal determinants of forkhead DNA bispecificity.

Cell Rep · 2025
L1 92/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
How its reproducibility compares
92/100
Reproducibility score
1.0 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 83% of all assessed papers rank 179 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (1:1 on structure, direction-matched on biology). The paper reports NO original code, so per P16 the described PBM pipeline was re-implemented from Methods and run on the paper's OWN deposited 8-mer E-score matrix (GEO GSE297784) via a 31s «our HPC» SLURM job (2208157). Dataset structure matches the paper EXACTLY: 260 PBM experiments, 32896 non-redundant 8-mers, 12 reference forkhead proteins, 83 variants (95 alleles - 12 refs). Bispecificity analysis (FKH=RYAAAYA vs FHL=GACGC, E>=0.4) classifies most references as monospecific and identifies FOXN2/FOXN3/FOXN4 as bispecific - exactly the FOXN subfamily the paper's variant case-studies center on. All four headline variant effects reproduce in the correct direction with high significance: FOXF1 Y89C and FOXB2 E52G strongly reduce FKH affinity (dE=-0.15/-0.34, p<=6e-8); FOXN2 T154A shifts toward FHL and FOXN3 L199V toward FKH (Mann-Whitney p<=7e-12 / 2e-9). Biological-effect claims are reported qualitatively in the paper so they are graded within-tol (claim reproduced, not an exact author number); structural counts are exact. NOT attempted: raw-.gpr->E-score re-derivation (proprietary MATLAB suite), gnomAD/ClinVar screening (external DBs), and ChIP-seq (not deposited in this series). All provisional - a human reviewer signs off.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 92
    assessed: 2026-06-19 ⛓ 1ca8e460aada
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests what determines whether a forkhead (FH) transcription factor domain binds only the FKH motif, only the FHL motif, or both (bispecificity), hypothesizing that non-DNA-contacting residues—particularly in the helix2-3 loop and wing 2—control this specificity.

Core claims
  • Non-DNA-contacting residues, especially in the loop between helices 2 and 3 and in wing 2, control mono- vs. bispecificity of FH domains for the FKH and FHL motifs mechanism
  • PBM screening of 12 reference FH proteins, 61 naturally occurring missense variants, and 22 designed mutant/chimeric FH proteins reveals effects of coding variation on DNA-binding activity method
  • No tested missense variant switched a monospecific FH domain to the opposite motif specificity or converted it to bispecificity, consistent with a distributed model requiring both subdomains to be permissive finding
  • Subdomain-swap experiments show the 2-3 loop and wing 2 each independently influence FKH/FHL preference, but neither alone confers bispecificity mechanism
  • FOXN3 ChIP-seq peaks in HepG2 and MCF-7 cells containing FHL motifs are specifically enriched for cilium-related genes/phenotypes, while FKH-motif peaks show no consistent enrichment finding
  • 19 naturally occurring FH missense variants in disease-associated genes altered DNA-binding activity, 16 not previously reported finding
Experimental setups
Assay System Perturbation Readout Platform
Universal (all 10-mer) protein-binding microarray (PBM) 12 reference human FH domains + 61 natural missense variants + 22 designed mutants (recombinant protein) missense mutation / designed mutation 8-mer E-scores reflecting DNA-binding specificity for FKH vs. FHL motifs universal all 10-mer PBM
PBM on subdomain-swap chimeric constructs Chimeric FH domains (FOXN3, FOXA1, FOXN1, FOXJ3 combinations; 2-3 loop, wing, and helix 1 swaps) subdomain swap (2-3 loop, wings, helix 1 residues) shift in FKH vs. FHL binding preference (E-scores) universal PBM
ChIP-seq (reanalysis of published data) FOXN3 in HepG2 and MCF-7 human cell lines none (endogenous) FKH/FHL motif enrichment at peaks; GO and human phenotype term enrichment via GREAT GREAT tool
Structural/bioinformatic analysis of protein-DNA cocrystal structures FOXK1, FOXC2, FOXG1, FOXN1, FOXN3 FH domains none intradomain wing2-helix1 contacts; recognition-helix asparagine rotamer position
Prior in vivo knockout phenotyping (cited) Embryonic mouse retina, Foxn3 gene Foxn3 knockout dysregulation of cilia-associated genes; ciliary organization/morphology
Key results
  • FOXN2 L119P (proline in helix 1) causes near-complete loss of DNA binding, almost no 8-mers with E>0.4
  • FOXJ3 V162A negative control (allele frequency 0.806) shows no significant difference from reference in binding
  • 0/6 negative control variants showed unambiguous binding effects; 7/7 positive controls compromised binding
  • FOXF1 Y89C (likely pathogenic, alveolar capillary dysplasia) shows severely reduced FKH-motif binding specificity
  • FOXN2 T154A shows increased preference for FHL vs. FKH 8-mers
  • FOXN2 P153L largely eliminates FHL-preferential binding while sparing FKH binding
  • FOXN3 2-3 loop replaced with FOXA1's abrogates FHL recognition without compromising FKH binding
  • FHL-motif-containing FOXN3 ChIP-seq peaks show 31 GO terms and 8 phenotype terms significantly enriched in common between HepG2 and MCF-7, all related to cilium function 31 GO terms; 8 phenotype terms
Key statistics
  • count 2,402 (gnomAD variants identified within FH domain or flanking 15 positions)
  • count 245 total (12 benign/likely benign, 130 pathogenic/likely pathogenic/risk factor, 103 VUS) (ClinVar variants within FH domain or flanking 15 positions)
  • count 19 variants (16 novel) (naturally occurring variants in disease-associated genes showing altered DNA binding)
  • other 0.806 (population allele frequency of FOXJ3 V162A negative control variant)
  • count 0/6 (negative control variants showing unambiguous binding effects)
  • count 31 (GO Biological Process terms significantly enriched in both HepG2 and MCF-7 FOXN3 FHL-motif ChIP-seq peaks)
  • count 8 (human phenotype associations significantly enriched in common, all ciliary dysfunction-related)
  • count 7 of top 10 (top enriched GO terms in MCF-7 FHL peaks shared with top 10 in HepG2 FHL peaks)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper describes a variant-effect screening study using protein-binding microarrays (PBMs) to compare DNA-binding enrichment (E-score) profiles of reference versus variant/mutant forkhead transcription factor domains, with results interpreted largely through scatterplot comparisons and qualitative descriptions of consistency across replicates (e.g., '2 independent replicates') rather than through named inferential statistical tests. A separate analysis examined enrichment of FKH/FHL DNA motifs within published ChIP-seq peaks relative to a matched genomic background, and used the GREAT tool to test for enrichment of Gene Ontology and human phenotype terms among peak-associated genes, reporting results as 'significantly enriched' without stating the specific test, correction method, or exact p-values in the provided text.

Replicationunclear Groupsreference vs. variant/mutant FH protein alleles for DNA-binding activity; FKH-motif vs. FHL-motif ChIP-seq peaks vs. genomic background; two cell lines (HepG2 vs. MCF-7) for GO/phenotype term overlap Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Not explicitly named; comparison of PBM E-scores for reference vs. variant/mutant alleles via scatterplots and replicate consistency Figures 2D, 2E, 3A-3E, 5A-5G, S1-S14 (variant and subdomain-swap binding comparisons) 2 independent replicates stated for FOXN2 L119P; replicate numbers not consistently stated for other variants not stated
Not explicitly named; motif enrichment in ChIP-seq peaks vs. matched genomic background Figures 4A and 4B (FOXN3 ChIP-seq peaks, HepG2 and MCF-7 cell lines) not stated
GREAT tool enrichment analysis (Gene Ontology / human phenotype term enrichment) Figures 4C and 4D (GO Biological Process and phenotype terms among FKH- vs. FHL-motif-containing ChIP-seq peaks) not stated
Approaches that could also have been used
  • Differences in PBM E-scores between reference and variant/mutant alleles are described qualitatively (e.g., 'consistently increased preference,' 'subtle but highly consistent shift') based on scatterplot comparisons across a small number of replicates.
    Could also: A paired statistical test comparing matched k-mer E-scores between reference and variant alleles (e.g., a paired t-test or Wilcoxon signed-rank test) could also be used. — This would provide a quantitative statistic and p-value summarizing the magnitude and consistency of binding shifts, complementing the visual/qualitative comparison of scatterplots.
  • GO and human phenotype term enrichment among ChIP-seq peaks was assessed with the GREAT tool and described as 'significantly enriched' without stating the specific correction method applied across the many terms tested.
    Could also: Explicitly reporting the multiple-testing correction (e.g., Benjamini-Hochberg FDR) and the significance thresholds used for GO/phenotype enrichment could also be included. — Stating the correction method clarifies how the false-discovery rate was controlled given the large number of GO/phenotype terms tested simultaneously.
  • Motif enrichment within ChIP-seq peaks was assessed relative to a matched genomic background without naming a specific statistical test in the provided text.
    Could also: A hypergeometric test or Fisher's exact test (as implemented in tools such as HOMER or AME/MEME) with reported enrichment p- or q-values could also be used. — This would give a standard, reproducible quantitative measure of motif enrichment strength alongside the qualitative description.
  • Reproducibility of binding measurements was described through a small number of independent replicates (e.g., 2 replicates for FOXN2 L119P) assessed largely by visual consistency.
    Could also: Reporting a replicate correlation statistic (e.g., Pearson or Spearman correlation coefficient with a confidence interval) between replicate E-score sets could also be used. — This would formally quantify reproducibility across replicates rather than relying on qualitative visual assessment of scatterplots.
  • Negative and positive control variant effects were summarized as simple counts (e.g., '0/6 negative control variants showed unambiguous effects on binding').
    Could also: A binomial or exact test comparing the observed proportion of affected controls to an expected null rate could also be used. — This would provide a formal statistical basis for the control validation claim in addition to the reported counts.
  • No dispersion measures (e.g., SD, SEM, or confidence intervals) around E-scores or enrichment estimates are reported in the provided text.
    Could also: Reporting a dispersion or interval estimate (e.g., SD across replicates or a bootstrap confidence interval on E-scores) could also be included alongside point estimates. — This would convey the variability of binding measurements, which is useful for readers assessing the reliability of small-magnitude specificity shifts.
Software: GREAT · SIFT · Mutation Assessor · PROVEAN · PrimateAI

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 41124077 (Cell Reports 2025)

Title: Missense variants in human forkhead transcription factors reveal determinants of forkhead DNA bispecificity. DOI: 10.1016/j.celrep.2025.116427 · PMCID PMC12795473 · GEO GSE297784 · BioProject PRJNA1266190

Code/data availability (verbatim from paper)

"PBM data have been deposited in the GEO database under accession number GSE297784 and are publicly available as of the date of publication. This paper does not report original code. Any additional information required to reanalyze the data ... is available from the lead contact upon request."

=> The BRIEF's "Code: github.com/BulykLab/MEDEA" is a MIS-HARVESTED lab repo (MEDEA is the lab's 2020 chromatin-accessibility tool, Genome Research; it is NOT this paper's pipeline). Per P16, reproduction is done by re-running the described pipeline (third-party / standard tools + a from-description reimplementation of the bespoke analysis) on the paper's OWN deposited data. This is equally valid.

Data

GSE297784 = universal protein binding microarray (uPBM), Agilent 8x60K "all 10-mer" arrays (AMADID 030236), GST-fusion forkhead DBDs, in-vitro expressed, 750 nM. Deposited PROCESSED data (GSE297784_processed_data_Matrix.txt.gz) = the 8-mer E-score matrix: 32896 non-redundant 8-mers x 260 experiments. Sample titles encode GENE.allele_arrayID (e.g. FOXA1.ref_388, FOXB2.E52G_414).

IN SCOPE (pipeline-derived, reproducible from deposited E-score matrix)

S1. Dataset structure: N experiments=260; 8-mers=32896; reference proteins=12; variants=83 (95 alleles-12 refs). S2. Motif definitions: 8-mers matching FKH (RYAAAYA) vs FHL (GACGC), E>=0.4 threshold (paper Methods). S3. Bispecificity classification of the 12 reference FH proteins (binds FKH only / FHL only / both=bispecific). S4. Affinity-altering variant effects (paired/signed-rank Wilcoxon on dE = variant-ref over motif 8-mers): FOXF1 Y89C -> severe FKH reduction; FOXB2 E52G -> reduced FKH binding. S5. Specificity-altering variant effects (FKH-dE vs FHL-dE distributions): FOXN2 T154A -> increased FHL preference; FOXN3 L199V -> shifted toward FKH preference.

OUT OF SCOPE (wet-lab / external / not pipeline-derivable here)

  • Site-directed mutagenesis, protein expression, the PBM wet-lab assay itself.
  • gnomAD/ClinVar variant screening counts (2402/245) — external database queries, not from GSE297784.
  • ChIP-seq (MACS2/Bowtie2/GREAT) — no ChIP accession in this GEO series; not deposited here.
  • Re-deriving E-scores from raw .gpr (Universal PBM Analysis Suite, MATLAB-proprietary) — deposited E-scores are the canonical processed product; we start from them (re-derivation is a deeper stretch goal).
S1_n_experiments
Reported
260 PBM experiments
Reproduced
260
exact
S1_n_8mers
Reported
every possible 8-mer E-score
Reproduced
32896 non-redundant 8-mers
exact
S1_n_reference_proteins
Reported
12 reference FH proteins
Reproduced
12
exact
S1_n_variants
Reported
83 variants
Reproduced
95 alleles - 12 refs = 83
exact
S3_bispecificity
Reported
most FH bind FKH or FHL; some bind both (bispecific)
Reproduced
8 FKH-only, 1 FHL-only, 3 bispecific (FOXN2/FOXN3/FOXN4)
within tolerance
S4_FOXF1_Y89C
Reported
severe reduction of FKH binding
Reproduced
FKH dE -0.145 (0.46->0.31), 100% reduced, p=6e-8
within tolerance
S4_FOXB2_E52G
Reported
reduced FKH binding
Reproduced
FKH dE -0.340 (0.46->0.11), p=3e-11
within tolerance
S5_FOXN2_T154A
Reported
increased FHL preference
Reproduced
FHL-favored vs FKH, Mann-Whitney p=7e-12
within tolerance
S5_FOXN3_L199V
Reported
shifted toward FKH preference
Reproduced
FKH-favored vs FHL, Mann-Whitney p=2e-9
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 92/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5

Clean reproduction. All structural counts (260 PBM experiments, 32896 8-mers, 12 references, 83 variants) match the paper exactly off its own deposited matrix (GSE297784), and all four headline variant effects reproduce in the correct direction with high significance (p from 6e-8 to 7e-12). The only caveats are on our/data side, not the authors': the biological claims are qualitative in the paper (graded on direction, not a number), the pipeline had to be re-implemented because no original code was shared, and some claims (gnomAD/ClinVar, ChIP-seq) are external/undeposited and out of scope. Severity is negligible and the central bispecificity conclusion holds fully.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

147.6 k
tokens (I/O) · 9.8 M incl. cache
52 min
runtime · 0 CPU-h
0.5 GB
peak RAM
1
HPC jobs
hummel
machine