Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genome scans of facial features in East Africans and cross-population comparisons reveal novel associations.

PLoS Genet · 2021
L1 No data access 2/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • No authors-side cause for any deviation
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
No data access Data access not granted

This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.

Reproduction agent’s raw note

DROP (data_restricted). Described well enough to reproduce in principle: code (MeshMonk v0.0.6, github.com/TheWebMonks/meshmonk) is public and alive, and Methods (registration, hierarchical segmentation, per-segment multivariate GWAS, enrichment/co-localization with locuscompare on R 3.6.1) are clearly written. BUT every reported pipeline output depends on individual-level facial + genomic data that the authors deliberately placed behind controlled access due to its highly identifiable nature: Tanzanian phenotypes FaceBase FB00000667.01; Tanzanian genotypes dbGaP phs000622.v1.p1; ALSPAC by application; 3DFN dbGaP phs000949.v1.p1 / FaceBase FB00000491.01; PSU & IUPUI never deposited (no broad-sharing consent). The GEO accession text-mined for this RU (GSE70751) is a CITED CNCC functional-annotation dataset (Prescott/Wysocka 2015, PMID 26365491), NOT the paper's primary data. The only public-data secondary analysis (enrichment of GWAS hits in public CNCC chromatin marks) cannot be reproduced 1:1 because the full GWAS summary statistics are not deposited (GWAS Catalog GCST90044778 fullPvalueSet=false); only the 20 lead loci are reported, and a crude 20-locus overlap would be a different/weaker analysis, so it was NOT attempted. No «our HPC» compute submitted (no obtainable input data => nothing to compute). NOT attempted: controlled-access data applications (dbGaP/FaceBase/ALSPAC require IRB+DUC). Fabrication concern: none; values are consistent with a controlled-access study using a known tool, simply not publicly checkable.

💻 Code ↗ 🗄 Data: GSE70751

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-15 ⛓ 549486041ebf
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Because most facial-morphology GWAS loci have been discovered in Europeans, the degree to which the genetic basis of facial features is shared across diverse human populations is unknown; the authors test this by performing a genome scan of 3D facial shape in an East African (Tanzanian) cohort and comparing results to Europeans.

Core claims
  • Genome scans of multivariate 3D facial shape phenotypes in Tanzanian children revealed significant associations at 20 loci (p < 2.5 × 10^-8). finding
  • Six of the 20 loci surpassed a stricter study-wide significance threshold accounting for multiple phenotypes (p < 6.25 × 10^-10). finding
  • Ten of the association signals were shared with Europeans (seven sharing the same associated SNP), indicating a partly shared genetic basis for facial shape across populations. finding
  • The associated loci were enriched for active chromatin elements in human cranial neural crest cells and embryonic craniofacial tissue, consistent with an early developmental origin of facial variation. mechanism
  • Two associations (5q31.1 and 12q21.31) lie in highly conserved regions showing craniofacial-specific enhancer activity during embryological development. finding
  • Cross-population comparison facilitated fine-mapping of causal variants at previously reported loci. finding
  • An open-ended, data-driven global-to-local phenotyping approach segments the face into hierarchically arranged multivariate features after adjusting for age, sex, height, weight, facial size, and population stratification. method
  • About half of the associations observed in Tanzanians were not present in Europeans, pointing to population-specific components of facial genetic architecture. finding
Experimental setups
Assay System Perturbation Readout Platform
GWAS (genome scan of multivariate 3D facial shape phenotypes) Tanzanian children (East African cohort) none hierarchical multivariate facial shape phenotypes from 3D images, adjusted for age, sex, height, weight, facial size, and population stratification 3D facial images; MeshMonk v0.0.6 spatially dense facial mapping; genotype data (dbGaP phs000622.v1.p1)
Cross-population comparison / replication of GWAS signals European comparison datasets (ALSPAC UK; 3DFN, PSU, IUPUI USA) none shared vs population-specific association signals; fine-mapping of causal variants
Functional enrichment / colocalization analysis human cranial neural crest cells and embryonic craniofacial tissue (ChIP-seq and epigenomic reference datasets) none enrichment of associated loci in active chromatin/enhancer elements locuscompare in R v3.6.1; GTEx v7; 1000G Phase 3; Roadmap Epigenomics; ChIP-seq datasets (GSE70751, GSE82295, GSE89179, GSE119997, GSE97752)
Key results
  • 20 genetic loci significantly associated with facial shape in Tanzanians 20 loci at p < 2.5 × 10^-8
  • 6 loci passed study-wide significance accounting for multiple phenotypes 6 loci at p < 6.25 × 10^-10
  • 10 of 20 signals shared with Europeans, 7 sharing the same associated SNP 10 of 20 (7 same SNP)
  • Associated loci enriched for active chromatin in cranial neural crest cells and embryonic craniofacial tissue
  • Two loci (5q31.1 and 12q21.31) in conserved regions with craniofacial-specific enhancer activity 2 loci
Key statistics
  • pvalue p < 2.5 × 10^-8 (genome-wide significance threshold for the 20 loci)
  • pvalue p < 6.25 × 10^-10 (stricter study-wide threshold accounting for multiple phenotypes; 6 loci passed)
  • count 2,595 (Tanzanian children with 3D facial images in the sample)
  • count 20 (loci with significant facial shape associations in Tanzanians)
  • count 10 (association signals shared with Europeans (7 sharing same SNP))
  • other approximately 40% to 60% (narrow-sense heritability of facial traits from twin and family studies)
  • count 203 signals across 138 genetic loci (prior European GWAS meta-analysis of data-driven 3D facial phenotypes)
  • count 3,505 (Bantu children and Mwanza adolescents in prior African facial GWAS (Cole et al.))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study performed GWAS of data-driven, hierarchically arranged multivariate 3D facial shape phenotypes in 2,595 Tanzanian children, adjusting for age, sex, height, weight, facial size, and population stratification. Association signals were evaluated at a genome-wide significance threshold (p < 2.5×10⁻⁸) and a stricter study-wide threshold (p < 6.25×10⁻¹⁰) to account for multiple phenotypes tested. Cross-population comparisons with European cohorts and co-localization analyses (locuscompare in R v3.6.1) were used to characterize shared versus population-specific signals, and associated loci were tested for enrichment in craniofacial developmental chromatin marks.

Replicationbiological Sample size2,595 Tanzanian children for primary GWAS; European comparison cohort (ALSPAC, 3DFN, PSU, IUPUI) sizes not stated in available text GroupsTanzanian children (primary GWAS); cross-population comparison against European cohorts for signal sharing Pairingunpaired Randomization/blindingnot stated Dispersionnone Multiplicity correctionTwo-stage threshold: genome-wide significance (p < 2.5×10⁻⁸) per phenotype; study-wide significance (p < 6.25×10⁻¹⁰) accounting for multiple phenotypes (specific derivation of study-wide threshold not detailed in available text)
Statistical tests used
Test Applied to n Assumptions
Genome-wide association test (specific model—e.g., linear mixed model or linear regression—not stated in available text) Association of genome-wide SNPs with multivariate 3D facial shape phenotypes in the Tanzanian cohort 2595 not stated
Co-localization analysis (locuscompare function in R) Cross-population comparison of GWAS signals between Tanzanian and European cohorts to identify shared causal variants not stated
Chromatin enrichment analysis (specific statistical method not stated in available text) Testing whether GWAS-significant loci were enriched for active chromatin marks in cranial neural crest cells and embryonic craniofacial tissue not stated
Approaches that could also have been used
  • A study-wide significance threshold (p < 6.25×10⁻¹⁰) was derived by applying a correction factor on top of the standard genome-wide threshold to account for multiple phenotypes
    Could also: Permutation-based or simulation-based significance thresholds could also be derived empirically to account for the correlation structure among the hierarchical phenotypes — When phenotypes are correlated (as hierarchically nested shape features likely are), a simple Bonferroni-style divisor over-corrects; empirical permutation thresholds account for the actual effective number of independent tests
  • Cross-population signal sharing was assessed using co-localization via the locuscompare function
    Could also: Formal Bayesian colocalization methods such as coloc or eCAVIAR could also be used to compare cross-population signals — Bayesian colocalization frameworks explicitly model posterior probabilities for distinct hypotheses (same causal variant, different causal variants, no association) and quantify uncertainty in signal sharing more formally
  • Enrichment of GWAS loci in craniofacial chromatin marks was assessed by overlapping significant hits with ChIP-seq data
    Could also: Stratified LD score regression (S-LDSC) could also partition heritability across tissue-specific regulatory annotations genome-wide — S-LDSC uses the full distribution of GWAS summary statistics rather than only genome-wide-significant loci, providing an unbiased estimate of heritability enrichment that is not conditioned on significance thresholds
  • Multivariate 3D shape phenotypes were derived from the hierarchical data-driven approach and genome-scanned
    Could also: A multivariate GWAS framework (e.g., MANOVA-based or canonical correlation analysis-based GWAS) could also directly model the full covariance structure of shape features in a single test per SNP — Joint multivariate models exploit correlations among shape dimensions and can increase power to detect pleiotropic variants affecting multiple facial features simultaneously
  • Population stratification was adjusted for as a covariate during phenotype derivation
    Could also: A linear mixed model incorporating a genomic relatedness matrix (GRM) — as implemented in tools such as BOLT-LMM or SAIGE — could also be applied to jointly control for stratification and cryptic relatedness — GRM-based mixed models have been shown to control genomic inflation more robustly in samples with complex structure, as is common in admixed or geographically diverse cohorts such as the Tanzanian sample
  • Fine-mapping of causal variants at previously reported loci was facilitated by cross-population comparison
    Could also: Formal statistical fine-mapping methods such as SuSiE, FINEMAP, or CAVIARBF could also be applied, leveraging the LD differences between populations to shrink credible sets — Probabilistic fine-mapping tools produce credible sets of putative causal variants with explicit posterior inclusion probabilities, enabling more systematic prioritization of candidates than LD-pattern comparison alone
Software: R (locuscompare function for co-localization) 3.6.1 · MeshMonk (spatially dense 3D facial mapping) 0.0.6

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
26
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

1q22 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
3p14 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
3q28 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
4q31 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
5q14 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
7q22 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
rs10122939 RefSNP in Discussion (http://purl.org/orb/Discussion)
no other assessed paper uses this yet
rs10878346 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs112643361 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs113199279 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs114777090 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs11959408 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs148390647 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs16983329 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs188502472 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs242980 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs56063440 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs56850662 RefSNP in Discussion (http://purl.org/orb/Discussion)
no other assessed paper uses this yet
rs58409393 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs74112009 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs77926594 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs80243479 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs9603276 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
rs9995821 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

C1
Reported
2,595 Tanzanian 3D facial images
Reproduced
NA
m.public.grade.error
C2
Reported
20 genome-wide significant loci (p<2.5e-8)
Reproduced
NA
m.public.grade.error
C3
Reported
6 loci at study-wide significance (p<6.25e-10)
Reproduced
NA
m.public.grade.error
C4
Reported
hits enriched in CNCC active chromatin / embryonic craniofacial enhancers
Reproduced
NA
m.public.grade.error
C5
Reported
conserved craniofacial-enhancer loci 5q31.1 and 12q21.31
Reproduced
NA
m.public.grade.error
C6
Reported
10 signals shared with Europeans (7 same SNP)
Reproduced
NA
m.public.grade.error
C7
Reported
novel SCHIP1 (centroid size) and PDE8A (allometric) associations
Reproduced
NA
m.public.grade.error

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 44/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is a clean data_restricted drop: none of the seven claims (e.g. 20 genome-wide loci at p<2.5e-8, novel SCHIP1/PDE8A associations) could be reproduced because every primary input — Tanzanian 3D faces and genotypes plus the European comparison cohorts — is controlled-access or never deposited, and full GWAS summary statistics are withheld (GCST90044778 fullPvalueSet=false), closing even the one public-annotation enrichment path. The limitation is on the data-availability side, which is legitimate for highly identifiable face+genome data, not an authors' defect or a methodological choice of ours. There is no fabrication concern — reported values are consistent with a controlled-access study using a public, living tool (MeshMonk v0.0.6); they are simply not independently checkable from public artifacts, so q1/q2 are red while q5/q7/q8 stay yellow per the restriction-is-not-a-defect principle.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

55.3 k
tokens (I/O) · 2.1 M incl. cache
5 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.