Next-generation phenotyping integrated in a national framework for patients with ultrarare disorders improves genetic diagnostics and yields new molecular findi
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡A deviation arose in the data or preprocessing
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED. The paper IS described well enough and the deposited authors' code (TNAMSE_geno_pheno Snakemake+R) reproduces its headline numbers 1:1 from the OPEN medRxiv supplement (media-3.xlsx, 1577-case cohort). 11 of 12 in-scope claims reproduce exact or within-tolerance: C1 1577/268/1309, C3 499(31.6%), C4 child 32.5%/adult 27.6% Fisher P=0.1301, C5 510, C6 228 de novo, C7 11 dual, C8 n=375, C9 VUS 80.2%/solving 44.1% missense, C10 34 novel, C11 23 candidate -- all match. Only C12 is a near-miss (363 distinct disease genes vs 370 reported, off by 7, likely SV/CNV inclusion). Figures 2,3,4 (gene-discovery vs year-of-first-report) and 5 (LASSO) reproduced at pipeline level; supplementary VUS/autozygosity/population figures too. 3 small faithful patches were needed (skip=2->skip=0 header-offset for the medRxiv-vs-published supplement layout x2; n()/count->n()/count[1] for dplyr>=1.1). NOT attempted/out of scope: C2 enrollment total (not deposited), C13 PEDIA/NGP (restricted facial images, separate repo), figure1 donuts (viz dep webr/PieDonut won't build), OMIM-gated figure1_part2/CCDS (license-gated, authors also leave commented out), VEP supporting (online). No fabrication concern -- every reproduced value is derivable from the shipped open data+code.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 87assessed: 2026-06-19 ⛓ 6c239fe1fbca
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-19
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether a structured, multidisciplinary national diagnostic framework (TRANSLATE NAMSE) using exome sequencing—combined with next-generation phenotyping tools such as AI-driven facial image analysis (GestaltMatcher/PEDIA)—improves molecular diagnosis of ultrarare disorders and enables discovery of novel gene-disease associations compared to conventional phenotype-only approaches.
- ★ A structured multidisciplinary exome sequencing framework established molecular genetic diagnoses in 32% of patients with suspected ultrarare disorders, comprising 370 distinct molecular causes. finding
- ★ The diagnostic process identified 34 novel and 23 candidate genotype-phenotype associations, mainly in neurodevelopmental disorders. finding
- ★ Computer-assisted facial image analysis (GestaltMatcher) enabled more efficient prioritization of exome sequencing data than approaches based solely on clinical features and molecular scores. finding
- ★ YieldPred, a phenotype-based model, was developed to estimate the probability of establishing a molecular diagnosis via exome sequencing. method
- Grouping patients into single major disease categories is overly simplistic, as HPO-term-based phenotype clusters for different disease categories show partial overlap. finding
- Most patients in the exome sequencing cohort were children with neurodevelopmental disorders, while adults were predominantly categorized with neurological/neuromuscular disorders. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| exome sequencing | human patients (rare/ultrarare disease cohort, TRANSLATE NAMSE) | none (diagnostic testing) | molecular genetic diagnosis / causative variants | — |
| deep phenotyping with Human Phenotype Ontology (HPO) annotation | human patients (n=1,577 exome cohort) | none | HPO terms per patient, disease category assignment | Human Phenotype Ontology |
| computer-assisted facial image analysis (GestaltMatcher, PEDIA) | human patients who consented to facial image analysis | none | prioritization/ranking of candidate genes from exome data based on facial dysmorphism | GestaltMatcher |
| phenotype-based diagnostic yield prediction (YieldPred) | human patients (exome sequencing cohort) | none | predicted probability of establishing a molecular diagnosis with exome sequencing | — |
| dimensionality reduction/clustering of HPO term profiles | human patients (exome cohort) | none | two-dimensional projection of patient phenotypic similarity | — |
- – 1,577 of 5,652 enrolled patients underwent exome sequencing after MDT evaluation
- – Molecular genetic diagnosis established in 32% of exome-sequenced patients 32%
- – 370 distinct molecular genetic causes identified, most with prevalence below 1:50,000 370
- – 34 novel and 23 candidate genotype-phenotype associations identified 34 novel; 23 candidate
- ▲ Facial-image-based prioritization (GestaltMatcher) improved efficiency of candidate gene prioritization versus clinical-feature/molecular-score-only approaches
- – Average of five HPO terms specified per individual 5 HPO terms/patient
- – 54% of children (n=702) assigned to 'neurodevelopmental disorders' category; 44% of adults (n=117) assigned to 'neurological or neuromuscular disorders' 54%; 44%
- – Phenotype clusters by disease category showed partial overlap in 2D HPO-term projection
- count 5,652 total enrolled (2,033 adults, 3,619 children) (TRANSLATE NAMSE overall cohort)
- count 1,577 exome sequencing cohort (268 adults, 1,309 children) (subset selected for exome sequencing)
- other 32% diagnostic yield (molecular diagnosis rate in exome cohort)
- count 370 distinct molecular genetic causes (diversity of causes identified)
- count 34 novel and 23 candidate genotype-phenotype associations (new disease gene discoveries)
- mean average of 5 HPO terms per individual (phenotype annotation density)
- count n=702 (54%) children with neurodevelopmental disorders (largest pediatric disease category)
- count n=117 (44%) adults with neurological/neuromuscular disorders (largest adult disease category)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The excerpt describes a prospective, multicenter observational cohort study (TRANSLATE NAMSE) in which 1,577 patients with suspected rare diseases underwent exome sequencing; results are reported mainly as descriptive counts, percentages, and one stated average (mean HPO terms per patient), with patient phenotype similarity visualized via dimensionality reduction of HPO term vectors. No inferential statistical tests, p-values, or formal group comparisons are described in the provided text.
-
The average number of HPO terms per patient is reported as a single mean value without an accompanying measure of spread.↳ Could also: Reporting the mean alongside SD or IQR (or presenting the full distribution, e.g., as a histogram or boxplot) — A dispersion measure would convey how much HPO-term counts varied across patients, which a lone mean does not capture.
-
Diagnostic and phenotype-category findings (e.g., percentages of patients by disease category, by age group) are presented as point-estimate proportions.↳ Could also: Accompanying key proportions with binomial or Wilson confidence intervals — A CI on a proportion communicates the precision of that estimate, which is particularly informative when comparing subgroup sizes that differ (e.g., 268 adults vs. 1,309 children).
-
Patient phenotype similarity across disease categories was visualized by projecting HPO terms into a two-dimensional space, with qualitative description of cluster overlap.↳ Could also: Complementing the visualization with a quantitative cluster-separation or overlap metric (e.g., silhouette score, permutation-based test of cluster separation) — A quantitative metric would let readers gauge the degree of category overlap numerically alongside the visual impression.
-
Differences in disease-category distribution between adults and children are described narratively (e.g., most children assigned to neurodevelopmental disorders, most adults to neurological/neuromuscular disorders) without a formal statistical comparison in this excerpt.↳ Could also: A chi-square or Fisher's exact test comparing categorical distributions between the adult and pediatric subgroups — Such a test would provide a formal statistical basis for the observed difference in category distribution between age groups, in addition to the descriptive comparison.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-39039281 (TRANSLATE-NAMSE geno/pheno)
Paper: Schmidt A et al. Next-generation phenotyping integrated in a national framework for patients with ultrarare disorders improves genetic diagnostics and yields new molecular findings. Nat Genet 2024. DOI 10.1038/s41588-024-01836-1.
Code: https://github.com/Ax-Sch/TNAMSE_geno_pheno (commit 6877589, cloned on «infra»). Data deposit: Zenodo 10.5281/zenodo.10964188 = code archive only (TRANSLATE_NAMSE_code.zip).
Data flow (from README + Snakefile)
- The cohort table is downloaded by rule
download_supplementfrom the medRxiv preprint supplement (config["supp_table"]-> medrxiv .../2023.04.19.23288824/DC3/embed/media-3.xlsx). OPEN, public. This single xlsx is the patient-level table the whole pipeline parses (parse_table.Rmd->supp_solved_cases.Rmd). No restricted patient genomes are needed. - Reference resources are shipped in the repo (
resources/): HGNC, HPO obo + categorization, ClinVar (variant_summary.txt.gz 77 MB + clinvar_20171203.vcf.gz), CCDS, Turro 2020 supplement (41586_2020_2434_MOESM5_ESM.xlsx), ACMG SF genes, gene year-of-first-report. All real files (no LFS). - OMIM files (genemap2, genes_to_phenotype, mimTitles) are license-gated (download requires an OMIM licence + a password handed out on request).
IN SCOPE (pipeline-derived, reproducible from open data + shipped resources)
Pipeline = Snakemake + R (snakemake 6, R/tidyverse). Targets:
parse_table-> cohort counts -> C1 (1,577; 268 adults / 1,309 children)create_solved_gene_table(supp_solved_cases) -> C3,C5,C7,C10,C11,C12 (499 solved, 510 dx, 11 dual, 34 novel, 23 candidate, 370 distinct)figure2_part1-> diagnostic yield by age/category -> C3,C4,C12figure3-> mode-of-inheritance, de novo, autozygosity -> C5,C6,C8figure4(prep ClinVar+Turro+TNAMSE, combine, plot) -> gene counts vs year of first report (Fig 4)figure5-> LASSO HPO-based prioritization (Fig 5)figure1_part1-> HPO category phenotype overview (Fig 1) [needs webr/moonBook from GitHub]supporting_missense(supp_VUS_vs_solving_vars) -> C9 (80% vs 45% missense)supporting_pops,supporting_autozygosity-> supplementary
OUT OF SCOPE
- figure1_part2 and supporting_ccds_length: require OMIM license-gated files
(genemap2_15_07_2021.txt, genes_to_phenotype.txt, mimTitles.txt). Not attempted
(no OMIM licence). Authors' own
rule allalso leaves these commented out. - supporting2_annotate_w_VEP: requires Ensembl VEP with online --database mode + network; heavy/fragile. Attempt only opportunistically; not core.
- PEDIA / GestaltMatcher / NGP (C13): separate repo igsb/PEDIA-TNAMSE and requires restricted patient facial images (not deposited). Out of scope — wet/NGP, not reproducible from public data. (P16 third-party-tool rule does not apply: data is restricted.)
Datasets to profile
- medRxiv supplement media-3.xlsx (cohort table) — the de-facto primary dataset.
- Zenodo 10964188 — code archive (not a data dataset per se; profiled as code deposit).
- Shipped reference resources (ClinVar variant_summary, Turro 2020 supp) — profiled as bundled inputs.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is essentially a 1:1 reproduction: the authors' own deposited Snakemake+R pipeline run on the OPEN medRxiv supplement reproduces 11 of 12 in-scope headline numbers exactly or within rounding tolerance (only 3 faithful version/format patches needed), and all central conclusions hold. The single substantive deviation is C12 (363 distinct disease genes vs reported 370, ~2%), an inclusion-rule ambiguity around SV/CNV counting that is mildly underspecified in the shipped code but does not affect any conclusion. The two unverified claims (C2 enrollment total, C13 PEDIA) are data-availability limitations (not deposited / restricted facial images), not author defects, and there is no fabrication concern.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.