Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Mutations in CHD7, encoding a chromatin-remodeling protein, cause idiopathic hypogonadotropic hypogonadism and Kallmann syndrome.

· 2008
PubMed 18834967 ↗ pmid-18834967
L1 70/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Total score +10
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
70/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 37% of all assessed papers rank 732 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PMID 18834967 (Kim 2008, AJHG) is a wet-lab candidate-gene paper: Sanger sequencing of CHD7 in 197 IHH/KS patients found 7 mutations, validated by RT-PCR and in situ hybridisation. It deposited NO data, ships NO code, and has NO bioinformatic pipeline -> the bulk is out of scope and not reproducible by design. The only two in-silico claims were reproduced on «our HPC» with standard third-party tools: (R1) modern SIFT (Ensembl VEP r110) reproduces 4/5 of the paper's per-variant pathogenicity calls (S834F,H55R,P2880L deleterious; A2789T tolerated); only K2948E flips (paper deleterious -> modern tolerated_low_confidence), consistent with 18 yr of database growth and K2948 being the paper's least-conserved residue. (R2) MAFFT alignment of 119 CHD7 orthologs reproduces the two 'fully conserved' residues exactly (H55 99.1%, S834 99.1%); A2789/P2880/K2948 are conserved in the majority but the fine 'highly vs relative' ranking does not fully reproduce across a broad phylogeny. Overall: an honest PARTIAL reproduction of the paper's small computational footprint; the wet-lab core is not attemptable. NOT attempted: mutation discovery, RT-PCR, in situ, 3D modelling.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 70
    assessed: 2026-06-19 ⛓ 6fef82d59738
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

We hypothesized that CHD7 would be involved in the pathogenesis of idiopathic hypogonadotropic hypogonadism and Kallmann syndrome (IHH/KS) without the CHARGE phenotype, and that IHH/KS represents a milder allelic variant of CHARGE syndrome.

Core claims
  • Heterozygous CHD7 mutations cause idiopathic hypogonadotropic hypogonadism and Kallmann syndrome without the CHARGE phenotype. finding
  • Both normosmic IHH and KS are mild allelic variants of CHARGE syndrome caused by CHD7 mutations. mechanism
  • Sporadic CHD7 mutations occur in approximately 6% of IHH/KS patients. finding
  • CHD7 is the first identified chromatin-remodeling protein with a role in human puberty and the second gene to cause both normosmic IHH and KS in humans. finding
  • Three of the identified mutations affect chromodomains critical for CHD7 function in chromatin remodeling and transcriptional regulation. mechanism
  • CHD7's role is corroborated by specific expression in IHH/KS-relevant tissues and appropriate developmental expression. finding
  • SIFT and protein-structure analysis support that the missense mutations affecting conserved residues are deleterious. method
Experimental setups
Assay System Perturbation Readout Platform
Sanger mutation screening / DNA sequencing 101 IHH/KS patients without CHARGE phenotype (human) none sequence variants in 37 protein-coding exons of CHD7
Targeted exon sequencing additional 96 IHH/KS patients (human) none sequence variants in exons 6–10 encoding conserved chromodomains
RT-PCR patient samples (human) none splice variant / transcript analysis
In situ hybridization IHH/KS-relevant tissues (developmental expression) none CHD7 spatial and developmental expression
In silico SIFT analysis CHD7 protein sequence none predicted deleteriousness of missense variants SIFT
Protein-structure analysis CHD7 chromodomains none structural impact of mutations on chromodomain function
Key results
  • Seven heterozygous mutations (two splice, five missense), absent in ≥180 controls, were identified in three sporadic KS and four sporadic normosmic IHH patients. 7 mutations
  • Three mutations affect chromodomains critical for proper CHD7 function; the other four affect conserved residues suggesting deleterious effect.
  • Sporadic CHD7 mutations occur in 6% of IHH/KS patients. 6%
  • CHD7 shows specific expression in IHH/KS-relevant tissues with appropriate developmental expression.
Key statistics
  • count 101 IHH/KS patients (37 exons screened) (primary mutation screening cohort)
  • count 96 IHH/KS patients (exons 6–10 screened) (additional chromodomain-focused cohort)
  • count 7 heterozygous mutations (2 splice, 5 missense) (mutations identified across patients)
  • count ≥180 controls (mutations absent in controls)
  • count 7 patients (3 sporadic KS, 4 sporadic normosmic IHH) (patients carrying mutations)
  • other 6% (frequency of sporadic CHD7 mutations in IHH/KS patients)
  • count 37 protein-coding exons (exons of CHD7 sequenced)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This report describes a genetic mutation-screening study rather than a hypothesis-testing statistical analysis: CHD7 exons were sequenced in cohorts of IHH/Kallmann syndrome patients (101 patients for all 37 exons, plus 96 additional patients for exons 6-10) and candidate variants were checked for absence in a control panel of at least 180 controls. Supportive evidence for pathogenicity came from in silico/functional approaches (RT-PCR, SIFT, protein-structure analysis, in situ hybridization) rather than inferential statistical tests. The text provided does not report p-values, effect sizes, or formal statistical comparisons.

Replicationunclear Sample size101 IHH/KS patients screened for all 37 CHD7 exons; an additional 96 IHH/KS patients screened for exons 6-10; variants checked against ≥180 controls GroupsIHH/Kallmann syndrome patients vs. controls, for presence/absence of CHD7 variants Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Approaches that could also have been used
  • Pathogenicity of identified variants was supported by absence in ≥180 controls, described as a simple presence/absence comparison.
    Could also: A formal case-control allele/genotype frequency comparison (e.g., Fisher's exact test) or population allele-frequency lookup in a reference database — This would let readers see a quantified estimate (e.g., an odds ratio or exact p-value) alongside the qualitative absence-in-controls observation, which can be useful when control panel sizes are modest.
  • Variant deleteriousness was assessed using SIFT and protein-structure analysis as supportive in silico evidence.
    Could also: Combining multiple in silico predictors (e.g., PolyPhen-2, CADD, or later-developed ensemble scores) alongside SIFT — Using several independent prediction algorithms can provide converging or complementary evidence about a variant's likely functional impact, which some readers find informative alongside a single predictor.
  • The overall mutation frequency was reported as a single percentage (6% of IHH/KS patients) without an accompanying interval estimate.
    Could also: Reporting a confidence interval around the proportion (e.g., a Wilson or Clopper-Pearson interval) — A confidence interval would convey the precision of the frequency estimate given the cohort size, which can be helpful for readers comparing this rate to other cohorts or genes.
Software: SIFT

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

R1_SIFT_S834F
Reported
deleterious
Reproduced
deleterious (0)
within tolerance
R1_SIFT_H55R
Reported
deleterious
Reproduced
deleterious_low_confidence (0.01)
within tolerance
R1_SIFT_P2880L
Reported
deleterious
Reproduced
deleterious (0.01)
within tolerance
R1_SIFT_A2789T
Reported
tolerated
Reproduced
tolerated (0.16)
within tolerance
R1_SIFT_K2948E
Reported
deleterious
Reproduced
tolerated_low_confidence (0.67)
did not match
R2_conserv_H55
Reported
fully conserved
Reproduced
99.1% identity (109/110 orthologs)
exact
R2_conserv_S834
Reported
fully conserved
Reproduced
99.1% identity (116/117)
exact
R2_conserv_A2789
Reported
highly conserved
Reproduced
63.6% identity (70/110)
partial
R2_conserv_P2880
Reported
highly conserved
Reproduced
63.6% identity (70/110)
partial
R2_conserv_K2948
Reported
relative conservation
Reproduced
73.6% identity (78/106)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 70/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Total score +10

This is a 2008 wet-lab candidate-gene paper (CHD7 in IHH/Kallmann) that deposited no data and no code, so its central claim — discovery of 7 CHD7 mutations in 197 patients, with RT-PCR and in situ support — is out of scope and not computationally reproducible. The only reproducible footprint is two in-silico analyses, and these hold well: 4/5 SIFT pathogenicity calls are concordant and the two 'fully conserved' residues reproduce exactly (H55 99.1%, S834 99.1%). The deviations are not the authors' fault — K2948E flips deleterious->tolerated_low_conf(0.67) purely from ~18 yrs of SIFT database drift, and the C-terminal conservation tiers shift because of our self-chosen broad 119-ortholog alignment. Overall a solid, honest partial reproduction with explainable deviations and no fabrication concern.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

147.6 k
tokens (I/O) · 8.1 M incl. cache
16 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.