Genome scans of facial features in East Africans and cross-population comparisons reveal novel associations.
Part of the results reproduced; minor but material deviations remained.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓No authors-side cause for any deviation
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.
▸Reproduction agent’s raw note
DROP (data_restricted). Described well enough to reproduce in principle: code (MeshMonk v0.0.6, github.com/TheWebMonks/meshmonk) is public and alive, and Methods (registration, hierarchical segmentation, per-segment multivariate GWAS, enrichment/co-localization with locuscompare on R 3.6.1) are clearly written. BUT every reported pipeline output depends on individual-level facial + genomic data that the authors deliberately placed behind controlled access due to its highly identifiable nature: Tanzanian phenotypes FaceBase FB00000667.01; Tanzanian genotypes dbGaP phs000622.v1.p1; ALSPAC by application; 3DFN dbGaP phs000949.v1.p1 / FaceBase FB00000491.01; PSU & IUPUI never deposited (no broad-sharing consent). The GEO accession text-mined for this RU (GSE70751) is a CITED CNCC functional-annotation dataset (Prescott/Wysocka 2015, PMID 26365491), NOT the paper's primary data. The only public-data secondary analysis (enrichment of GWAS hits in public CNCC chromatin marks) cannot be reproduced 1:1 because the full GWAS summary statistics are not deposited (GWAS Catalog GCST90044778 fullPvalueSet=false); only the 20 lead loci are reported, and a crude 20-locus overlap would be a different/weaker analysis, so it was NOT attempted. No «our HPC» compute submitted (no obtainable input data => nothing to compute). NOT attempted: controlled-access data applications (dbGaP/FaceBase/ALSPAC require IRB+DUC). Fabrication concern: none; values are consistent with a controlled-access study using a known tool, simply not publicly checkable.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-15 ⛓ 549486041ebf
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusBecause most facial-morphology GWAS loci have been discovered in Europeans, the degree to which the genetic basis of facial features is shared across diverse human populations is unknown; the authors test this by performing a genome scan of 3D facial shape in an East African (Tanzanian) cohort and comparing results to Europeans.
- ★ Genome scans of multivariate 3D facial shape phenotypes in Tanzanian children revealed significant associations at 20 loci (p < 2.5 × 10^-8). finding
- ★ Six of the 20 loci surpassed a stricter study-wide significance threshold accounting for multiple phenotypes (p < 6.25 × 10^-10). finding
- ★ Ten of the association signals were shared with Europeans (seven sharing the same associated SNP), indicating a partly shared genetic basis for facial shape across populations. finding
- ★ The associated loci were enriched for active chromatin elements in human cranial neural crest cells and embryonic craniofacial tissue, consistent with an early developmental origin of facial variation. mechanism
- ★ Two associations (5q31.1 and 12q21.31) lie in highly conserved regions showing craniofacial-specific enhancer activity during embryological development. finding
- ★ Cross-population comparison facilitated fine-mapping of causal variants at previously reported loci. finding
- ★ An open-ended, data-driven global-to-local phenotyping approach segments the face into hierarchically arranged multivariate features after adjusting for age, sex, height, weight, facial size, and population stratification. method
- ★ About half of the associations observed in Tanzanians were not present in Europeans, pointing to population-specific components of facial genetic architecture. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| GWAS (genome scan of multivariate 3D facial shape phenotypes) | Tanzanian children (East African cohort) | none | hierarchical multivariate facial shape phenotypes from 3D images, adjusted for age, sex, height, weight, facial size, and population stratification | 3D facial images; MeshMonk v0.0.6 spatially dense facial mapping; genotype data (dbGaP phs000622.v1.p1) |
| Cross-population comparison / replication of GWAS signals | European comparison datasets (ALSPAC UK; 3DFN, PSU, IUPUI USA) | none | shared vs population-specific association signals; fine-mapping of causal variants | — |
| Functional enrichment / colocalization analysis | human cranial neural crest cells and embryonic craniofacial tissue (ChIP-seq and epigenomic reference datasets) | none | enrichment of associated loci in active chromatin/enhancer elements | locuscompare in R v3.6.1; GTEx v7; 1000G Phase 3; Roadmap Epigenomics; ChIP-seq datasets (GSE70751, GSE82295, GSE89179, GSE119997, GSE97752) |
- – 20 genetic loci significantly associated with facial shape in Tanzanians 20 loci at p < 2.5 × 10^-8
- – 6 loci passed study-wide significance accounting for multiple phenotypes 6 loci at p < 6.25 × 10^-10
- – 10 of 20 signals shared with Europeans, 7 sharing the same associated SNP 10 of 20 (7 same SNP)
- ▲ Associated loci enriched for active chromatin in cranial neural crest cells and embryonic craniofacial tissue
- – Two loci (5q31.1 and 12q21.31) in conserved regions with craniofacial-specific enhancer activity 2 loci
- pvalue p < 2.5 × 10^-8 (genome-wide significance threshold for the 20 loci)
- pvalue p < 6.25 × 10^-10 (stricter study-wide threshold accounting for multiple phenotypes; 6 loci passed)
- count 2,595 (Tanzanian children with 3D facial images in the sample)
- count 20 (loci with significant facial shape associations in Tanzanians)
- count 10 (association signals shared with Europeans (7 sharing same SNP))
- other approximately 40% to 60% (narrow-sense heritability of facial traits from twin and family studies)
- count 203 signals across 138 genetic loci (prior European GWAS meta-analysis of data-driven 3D facial phenotypes)
- count 3,505 (Bantu children and Mwanza adolescents in prior African facial GWAS (Cole et al.))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study performed GWAS of data-driven, hierarchically arranged multivariate 3D facial shape phenotypes in 2,595 Tanzanian children, adjusting for age, sex, height, weight, facial size, and population stratification. Association signals were evaluated at a genome-wide significance threshold (p < 2.5×10⁻⁸) and a stricter study-wide threshold (p < 6.25×10⁻¹⁰) to account for multiple phenotypes tested. Cross-population comparisons with European cohorts and co-localization analyses (locuscompare in R v3.6.1) were used to characterize shared versus population-specific signals, and associated loci were tested for enrichment in craniofacial developmental chromatin marks.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Genome-wide association test (specific model—e.g., linear mixed model or linear regression—not stated in available text) | Association of genome-wide SNPs with multivariate 3D facial shape phenotypes in the Tanzanian cohort | 2595 | not stated |
| Co-localization analysis (locuscompare function in R) | Cross-population comparison of GWAS signals between Tanzanian and European cohorts to identify shared causal variants | — | not stated |
| Chromatin enrichment analysis (specific statistical method not stated in available text) | Testing whether GWAS-significant loci were enriched for active chromatin marks in cranial neural crest cells and embryonic craniofacial tissue | — | not stated |
-
A study-wide significance threshold (p < 6.25×10⁻¹⁰) was derived by applying a correction factor on top of the standard genome-wide threshold to account for multiple phenotypes↳ Could also: Permutation-based or simulation-based significance thresholds could also be derived empirically to account for the correlation structure among the hierarchical phenotypes — When phenotypes are correlated (as hierarchically nested shape features likely are), a simple Bonferroni-style divisor over-corrects; empirical permutation thresholds account for the actual effective number of independent tests
-
Cross-population signal sharing was assessed using co-localization via the locuscompare function↳ Could also: Formal Bayesian colocalization methods such as coloc or eCAVIAR could also be used to compare cross-population signals — Bayesian colocalization frameworks explicitly model posterior probabilities for distinct hypotheses (same causal variant, different causal variants, no association) and quantify uncertainty in signal sharing more formally
-
Enrichment of GWAS loci in craniofacial chromatin marks was assessed by overlapping significant hits with ChIP-seq data↳ Could also: Stratified LD score regression (S-LDSC) could also partition heritability across tissue-specific regulatory annotations genome-wide — S-LDSC uses the full distribution of GWAS summary statistics rather than only genome-wide-significant loci, providing an unbiased estimate of heritability enrichment that is not conditioned on significance thresholds
-
Multivariate 3D shape phenotypes were derived from the hierarchical data-driven approach and genome-scanned↳ Could also: A multivariate GWAS framework (e.g., MANOVA-based or canonical correlation analysis-based GWAS) could also directly model the full covariance structure of shape features in a single test per SNP — Joint multivariate models exploit correlations among shape dimensions and can increase power to detect pleiotropic variants affecting multiple facial features simultaneously
-
Population stratification was adjusted for as a covariate during phenotype derivation↳ Could also: A linear mixed model incorporating a genomic relatedness matrix (GRM) — as implemented in tools such as BOLT-LMM or SAIGE — could also be applied to jointly control for stratification and cryptic relatedness — GRM-based mixed models have been shown to control genomic inflation more robustly in samples with complex structure, as is common in admixed or geographically diverse cohorts such as the Tanzanian sample
-
Fine-mapping of causal variants at previously reported loci was facilitated by cross-population comparison↳ Could also: Formal statistical fine-mapping methods such as SuSiE, FINEMAP, or CAVIARBF could also be applied, leveraging the LD differences between populations to shrink credible sets — Probabilistic fine-mapping tools produce credible sets of putative causal variants with explicit posterior inclusion probabilities, enabling more systematic prioritization of candidates than LD-pattern comparison alone
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Facial shape GWAS loci are enriched for active chromatin marks in cranial neural crest cells and embryonic craniofacial tissue.other human cranial-neural-crest-cell up 2021×1papers★ This paper is the founder (earliest)
-
Facial shape association loci at 5q31.1 and 12q21.31 overlap conserved regions with craniofacial-specific enhancer activity.other human craniofacial-tissue 2021×1papers★ This paper is the founder (earliest)
-
10 of 20 East African facial shape GWAS loci replicate in Europeans; 7 share the same index SNP, indicating partial cross-population genetic overlap.other tanzanian-children-european 2021×1papers★ This paper is the founder (earliest)
-
20 genetic loci significantly associated with 3D facial shape in Tanzanian children at p < 2.5×10^-8.other tanzanian children 2021×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean data_restricted drop: none of the seven claims (e.g. 20 genome-wide loci at p<2.5e-8, novel SCHIP1/PDE8A associations) could be reproduced because every primary input — Tanzanian 3D faces and genotypes plus the European comparison cohorts — is controlled-access or never deposited, and full GWAS summary statistics are withheld (GCST90044778 fullPvalueSet=false), closing even the one public-annotation enrichment path. The limitation is on the data-availability side, which is legitimate for highly identifiable face+genome data, not an authors' defect or a methodological choice of ours. There is no fabrication concern — reported values are consistent with a controlled-access study using a public, living tool (MeshMonk v0.0.6); they are simply not independently checkable from public artifacts, so q1/q2 are red while q5/q7/q8 stay yellow per the restriction-is-not-a-defect principle.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.