Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Next-generation phenotyping integrated in a national framework for patients with ultrarare disorders improves genetic diagnostics and yields new molecular findi

Nat Genet · 2024
L1 87/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
How its reproducibility compares
87/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 72% of all assessed papers rank 301 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED. The paper IS described well enough and the deposited authors' code (TNAMSE_geno_pheno Snakemake+R) reproduces its headline numbers 1:1 from the OPEN medRxiv supplement (media-3.xlsx, 1577-case cohort). 11 of 12 in-scope claims reproduce exact or within-tolerance: C1 1577/268/1309, C3 499(31.6%), C4 child 32.5%/adult 27.6% Fisher P=0.1301, C5 510, C6 228 de novo, C7 11 dual, C8 n=375, C9 VUS 80.2%/solving 44.1% missense, C10 34 novel, C11 23 candidate -- all match. Only C12 is a near-miss (363 distinct disease genes vs 370 reported, off by 7, likely SV/CNV inclusion). Figures 2,3,4 (gene-discovery vs year-of-first-report) and 5 (LASSO) reproduced at pipeline level; supplementary VUS/autozygosity/population figures too. 3 small faithful patches were needed (skip=2->skip=0 header-offset for the medRxiv-vs-published supplement layout x2; n()/count->n()/count[1] for dplyr>=1.1). NOT attempted/out of scope: C2 enrollment total (not deposited), C13 PEDIA/NGP (restricted facial images, separate repo), figure1 donuts (viz dep webr/PieDonut won't build), OMIM-gated figure1_part2/CCDS (license-gated, authors also leave commented out), VEP supporting (online). No fabrication concern -- every reproduced value is derivable from the shipped open data+code.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.10964188

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 87
    assessed: 2026-06-19 ⛓ 6c239fe1fbca
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether a structured, multidisciplinary national diagnostic framework (TRANSLATE NAMSE) using exome sequencing—combined with next-generation phenotyping tools such as AI-driven facial image analysis (GestaltMatcher/PEDIA)—improves molecular diagnosis of ultrarare disorders and enables discovery of novel gene-disease associations compared to conventional phenotype-only approaches.

Core claims
  • A structured multidisciplinary exome sequencing framework established molecular genetic diagnoses in 32% of patients with suspected ultrarare disorders, comprising 370 distinct molecular causes. finding
  • The diagnostic process identified 34 novel and 23 candidate genotype-phenotype associations, mainly in neurodevelopmental disorders. finding
  • Computer-assisted facial image analysis (GestaltMatcher) enabled more efficient prioritization of exome sequencing data than approaches based solely on clinical features and molecular scores. finding
  • YieldPred, a phenotype-based model, was developed to estimate the probability of establishing a molecular diagnosis via exome sequencing. method
  • Grouping patients into single major disease categories is overly simplistic, as HPO-term-based phenotype clusters for different disease categories show partial overlap. finding
  • Most patients in the exome sequencing cohort were children with neurodevelopmental disorders, while adults were predominantly categorized with neurological/neuromuscular disorders. finding
Experimental setups
Assay System Perturbation Readout Platform
exome sequencing human patients (rare/ultrarare disease cohort, TRANSLATE NAMSE) none (diagnostic testing) molecular genetic diagnosis / causative variants
deep phenotyping with Human Phenotype Ontology (HPO) annotation human patients (n=1,577 exome cohort) none HPO terms per patient, disease category assignment Human Phenotype Ontology
computer-assisted facial image analysis (GestaltMatcher, PEDIA) human patients who consented to facial image analysis none prioritization/ranking of candidate genes from exome data based on facial dysmorphism GestaltMatcher
phenotype-based diagnostic yield prediction (YieldPred) human patients (exome sequencing cohort) none predicted probability of establishing a molecular diagnosis with exome sequencing
dimensionality reduction/clustering of HPO term profiles human patients (exome cohort) none two-dimensional projection of patient phenotypic similarity
Key results
  • 1,577 of 5,652 enrolled patients underwent exome sequencing after MDT evaluation
  • Molecular genetic diagnosis established in 32% of exome-sequenced patients 32%
  • 370 distinct molecular genetic causes identified, most with prevalence below 1:50,000 370
  • 34 novel and 23 candidate genotype-phenotype associations identified 34 novel; 23 candidate
  • Facial-image-based prioritization (GestaltMatcher) improved efficiency of candidate gene prioritization versus clinical-feature/molecular-score-only approaches
  • Average of five HPO terms specified per individual 5 HPO terms/patient
  • 54% of children (n=702) assigned to 'neurodevelopmental disorders' category; 44% of adults (n=117) assigned to 'neurological or neuromuscular disorders' 54%; 44%
  • Phenotype clusters by disease category showed partial overlap in 2D HPO-term projection
Key statistics
  • count 5,652 total enrolled (2,033 adults, 3,619 children) (TRANSLATE NAMSE overall cohort)
  • count 1,577 exome sequencing cohort (268 adults, 1,309 children) (subset selected for exome sequencing)
  • other 32% diagnostic yield (molecular diagnosis rate in exome cohort)
  • count 370 distinct molecular genetic causes (diversity of causes identified)
  • count 34 novel and 23 candidate genotype-phenotype associations (new disease gene discoveries)
  • mean average of 5 HPO terms per individual (phenotype annotation density)
  • count n=702 (54%) children with neurodevelopmental disorders (largest pediatric disease category)
  • count n=117 (44%) adults with neurological/neuromuscular disorders (largest adult disease category)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The excerpt describes a prospective, multicenter observational cohort study (TRANSLATE NAMSE) in which 1,577 patients with suspected rare diseases underwent exome sequencing; results are reported mainly as descriptive counts, percentages, and one stated average (mean HPO terms per patient), with patient phenotype similarity visualized via dimensionality reduction of HPO term vectors. No inferential statistical tests, p-values, or formal group comparisons are described in the provided text.

Replicationunclear Sample sizeCohort sizes are stated directly (5,652 enrolled; 1,577 in the exome sequencing subcohort, comprising 268 adults and 1,309 children); no formal power or sample-size calculation is described in this excerpt. GroupsDescriptive comparison of adults vs. children and across six disease categories (e.g., neurodevelopmental, neurological/neuromuscular, organ malformation, endocrine/metabolic, immune/hematologic, cardiovascular) Pairingna Randomization/blindingnot stated Dispersionnone
Approaches that could also have been used
  • The average number of HPO terms per patient is reported as a single mean value without an accompanying measure of spread.
    Could also: Reporting the mean alongside SD or IQR (or presenting the full distribution, e.g., as a histogram or boxplot) — A dispersion measure would convey how much HPO-term counts varied across patients, which a lone mean does not capture.
  • Diagnostic and phenotype-category findings (e.g., percentages of patients by disease category, by age group) are presented as point-estimate proportions.
    Could also: Accompanying key proportions with binomial or Wilson confidence intervals — A CI on a proportion communicates the precision of that estimate, which is particularly informative when comparing subgroup sizes that differ (e.g., 268 adults vs. 1,309 children).
  • Patient phenotype similarity across disease categories was visualized by projecting HPO terms into a two-dimensional space, with qualitative description of cluster overlap.
    Could also: Complementing the visualization with a quantitative cluster-separation or overlap metric (e.g., silhouette score, permutation-based test of cluster separation) — A quantitative metric would let readers gauge the degree of category overlap numerically alongside the visual impression.
  • Differences in disease-category distribution between adults and children are described narratively (e.g., most children assigned to neurodevelopmental disorders, most adults to neurological/neuromuscular disorders) without a formal statistical comparison in this excerpt.
    Could also: A chi-square or Fisher's exact test comparing categorical distributions between the adult and pediatric subgroups — Such a test would provide a formal statistical basis for the observed difference in category distribution between age groups, in addition to the descriptive comparison.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39039281 (TRANSLATE-NAMSE geno/pheno)

Paper: Schmidt A et al. Next-generation phenotyping integrated in a national framework for patients with ultrarare disorders improves genetic diagnostics and yields new molecular findings. Nat Genet 2024. DOI 10.1038/s41588-024-01836-1.

Code: https://github.com/Ax-Sch/TNAMSE_geno_pheno (commit 6877589, cloned on «infra»). Data deposit: Zenodo 10.5281/zenodo.10964188 = code archive only (TRANSLATE_NAMSE_code.zip).

Data flow (from README + Snakefile)

  • The cohort table is downloaded by rule download_supplement from the medRxiv preprint supplement (config["supp_table"] -> medrxiv .../2023.04.19.23288824/DC3/embed/media-3.xlsx). OPEN, public. This single xlsx is the patient-level table the whole pipeline parses (parse_table.Rmd -> supp_solved_cases.Rmd). No restricted patient genomes are needed.
  • Reference resources are shipped in the repo (resources/): HGNC, HPO obo + categorization, ClinVar (variant_summary.txt.gz 77 MB + clinvar_20171203.vcf.gz), CCDS, Turro 2020 supplement (41586_2020_2434_MOESM5_ESM.xlsx), ACMG SF genes, gene year-of-first-report. All real files (no LFS).
  • OMIM files (genemap2, genes_to_phenotype, mimTitles) are license-gated (download requires an OMIM licence + a password handed out on request).

IN SCOPE (pipeline-derived, reproducible from open data + shipped resources)

Pipeline = Snakemake + R (snakemake 6, R/tidyverse). Targets:

  • parse_table -> cohort counts -> C1 (1,577; 268 adults / 1,309 children)
  • create_solved_gene_table (supp_solved_cases) -> C3,C5,C7,C10,C11,C12 (499 solved, 510 dx, 11 dual, 34 novel, 23 candidate, 370 distinct)
  • figure2_part1 -> diagnostic yield by age/category -> C3,C4,C12
  • figure3 -> mode-of-inheritance, de novo, autozygosity -> C5,C6,C8
  • figure4 (prep ClinVar+Turro+TNAMSE, combine, plot) -> gene counts vs year of first report (Fig 4)
  • figure5 -> LASSO HPO-based prioritization (Fig 5)
  • figure1_part1 -> HPO category phenotype overview (Fig 1) [needs webr/moonBook from GitHub]
  • supporting_missense (supp_VUS_vs_solving_vars) -> C9 (80% vs 45% missense)
  • supporting_pops, supporting_autozygosity -> supplementary

OUT OF SCOPE

  • figure1_part2 and supporting_ccds_length: require OMIM license-gated files (genemap2_15_07_2021.txt, genes_to_phenotype.txt, mimTitles.txt). Not attempted (no OMIM licence). Authors' own rule all also leaves these commented out.
  • supporting2_annotate_w_VEP: requires Ensembl VEP with online --database mode + network; heavy/fragile. Attempt only opportunistically; not core.
  • PEDIA / GestaltMatcher / NGP (C13): separate repo igsb/PEDIA-TNAMSE and requires restricted patient facial images (not deposited). Out of scope — wet/NGP, not reproducible from public data. (P16 third-party-tool rule does not apply: data is restricted.)

Datasets to profile

  1. medRxiv supplement media-3.xlsx (cohort table) — the de-facto primary dataset.
  2. Zenodo 10964188 — code archive (not a data dataset per se; profiled as code deposit).
  3. Shipped reference resources (ClinVar variant_summary, Turro 2020 supp) — profiled as bundled inputs.
Figures / tables: Fig 2aFig 3tableFig 3bFig 2
C1
Reported
1577 (268 adults, 1309 children)
Reproduced
1577 (268 adults, 1309 children)
exact
C2
Reported
5652 enrolled
Reproduced
not in deposited data
partial
C3
Reported
499 solved (32%)
Reproduced
499 (31.6%)
exact
C4
Reported
children 32% vs adults 28%, P=0.13
Reproduced
child 32.5% vs adult 27.6%, Fisher P=0.1301
exact
C5
Reported
510 diagnoses
Reproduced
510
exact
C6
Reported
228 de novo (45%)
Reproduced
228 (44.7%)
exact
C7
Reported
11 dual diagnoses
Reproduced
11
exact
C8
Reported
n=375 (Fig 3b autozygosity per MOI)
Reproduced
375
exact
C9
Reported
VUS 80% vs solving 45% missense
Reproduced
VUS 80.2% vs solving 44.1%
within tolerance
C10
Reported
34 novel associations
Reproduced
34
exact
C11
Reported
23 candidate associations
Reproduced
23
exact
C12
Reported
370 distinct molecular causes
Reproduced
363 distinct genes (+15 SV/CNV)
partial
C13
Reported
PEDIA/GestaltMatcher top-10 accuracy
Reproduced
not attempted (out of scope)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 87/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5

This is essentially a 1:1 reproduction: the authors' own deposited Snakemake+R pipeline run on the OPEN medRxiv supplement reproduces 11 of 12 in-scope headline numbers exactly or within rounding tolerance (only 3 faithful version/format patches needed), and all central conclusions hold. The single substantive deviation is C12 (363 distinct disease genes vs reported 370, ~2%), an inclusion-rule ambiguity around SV/CNV counting that is mildly underspecified in the shipped code but does not affect any conclusion. The two unverified claims (C2 enrollment total, C13 PEDIA) are data-availability limitations (not deposited / restricted facial images), not author defects, and there is no fabrication concern.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

382.5 k
tokens (I/O) · 28.8 M incl. cache
104 min
runtime · 0.03 CPU-h
2.6 GB
peak RAM
2
HPC jobs
hummel
machine