Tracing human genetic histories and natural selection with precise local ancestry inference.
The main results reproduced: recomputed values matched the published ones within tolerance.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced (partial, leaning positive). Paper introduces Orchestra, a deep-learning local ancestry inference (LAI) method. NOTE two BRIEF template-default errors corrected: real code repo is github.com/omicsedge/orchestra-paper (NOT shapeit4, a phasing tool); data is the bundled toy panel + public 1000G/HGDP/SGDP (NOT GSE80534, an unrelated SNP-array series the paper never cites). C1 environment reproduced via shipped CI container. C2 toy LAI pipeline (SLiM simulation -> base+smooth training -> inference on 20 admixed Mexicans) ran fully end-to-end on «our HPC» «job» (COMPLETED, exit 0:0) producing ancestry.tsv. C3: the BASE-layer global ancestry (NAM 49.2/EUR 44.4/AFR 6.4) matches the literature MXL proportions (47/48/5) within 3.6% — a strong quantitative hit; the SMOOTH (final) layer is directionally correct but EUR-skewed, an expected artifact of the TOY model (100 epochs, 30 reference samples, not the production model). C4: the Step3 LAD-score machinery reproduces structurally end-to-end on the toy output (1992 windows, canonical two-panel layout), but the canonical selection-scan FIGURES (Fadm/Manhattan) are out of reach from the GitHub-shipped toy artifacts alone — they need MAF-filtered full 1KGP/HGDP genotype panels hosted on CodeOcean (no DOI given in the repo) plus the production model. NOT attempted (out of scope): the headline Orchestra-vs-RFMix/FLARE/Gnomix accuracy benchmark and the Ashkenazi/Latin-American/Viking downstream analyses, which require the full production training panel + weights that are not shipped. Honest verdict: the paper is described well enough that its toy LAI pipeline reproduces 1:1 and the central MXL-ancestry direction is recovered; the production-scale benchmarks and CodeOcean-gated selection figures are not reproducible from the public toy artifacts.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 76assessed: 2026-06-20 ⛓ 7b74443948c3
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-20
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetCurrent local ancestry inference (LAI) methods struggle with closely related reference populations and large reference panels, so the authors hypothesize that a new, more granular LAI method (Orchestra) can achieve higher-resolution, within-continent ancestry resolution than existing tools and thereby reveal finer demographic history and selection signals.
- ★ Orchestra, a two-stage LAI method combining a recombination-distance base layer with a deep learning (convolutional + attention) smoothing module, outperforms RFmix, FLARE and Gnomix in precision and recall across simulated admixture generations. method
- ★ Orchestra maintains high per-population accuracy across 35 worldwide populations, including closely related ancestries, unlike competing methods which drop below 50% accuracy for a third of populations. finding
- ★ Orchestra accurately reconstructs Latin American admixture patterns and detects trace ancestries (e.g., Central/South/Southeast African in Brazil, Indian in the Guianas, Ashkenazi Jewish in Argentina, Japanese/Korean in Brazil and Peru) consistent with historical records. finding
- ★ Orchestra can be used to estimate genetic closeness between 35 populations by omitting each target population from its own reference panel and projecting resulting admixture distances via SMACOF, replicating known geographic/genetic relationships. method
- ★ Orchestra sheds light on debated Ashkenazi Jewish origins, highlighting South European heritage. finding
- ★ A curated high-quality reference panel of 10,169 non-admixed individuals from 35 world regions was assembled by combining >30 studies, filtered by MAF, batch-effect GWAS, and PCA/UMAP and t-SNE/GNN outlier removal. resource
- Orchestra enables mapping of selection signatures, including trace Scandinavian ancestry in British samples linked to an immune-rich region from Viking-era admixture. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Simulated admixture benchmarking (precision/recall) | 1KGP-16pops (N=3276) and custom-35pops (N=24,408) reference panels | 6 generations of simulated random admixture via SLiM | precision and recall per generation and per population | SLiM |
| Population-level and region/continental accuracy comparison | 1KGP-16pops and custom-35pops panels | simulated admixture | per-population accuracy (%), r2 metric | — |
| Local ancestry inference on real-world samples | Independent set of >10,000 UK Biobank samples not in training panel (103 countries) | none (real admixed individuals) | accuracy compared across LAI models | — |
| Simulated Latin American admixture benchmarking | Simulated Latin Americans from Southern/Northern European, Western/Central-Southern African and reconstructed Native American (1KGP East Asian segments as proxy) samples (N=4068) | 12 generations of simulated admixture via SLiM, region-specific ancestry proportions (Antilles, Mexico/Central America, South America) | precision and recall | SLiM |
| Real-world local ancestry deconvolution | 1KGP admixed American populations and UKBB participants born in Latin America (N=16,224) | none (real admixed individuals) | ancestral composition proportions and chromosome tract length distributions per country | — |
| Ancestral mapping (leave-one-population-out) | 35 custom reference populations, each iteratively omitted from its own reference panel | none / population omission | admixture proportions converted to distance matrix, projected via SMACOF algorithm | SMACOF algorithm |
| Reference panel curation (PCA/UMAP, t-SNE/GNN outlier detection, GWAS batch-effect filtering) | 10,169 non-admixed individuals from 35 world regions, incl. UK Biobank non-UK ancestries | none | sample inclusion/exclusion based on agreement between reported and inferred ancestry | PCA, UMAP, t-SNE, tsinfer (GNN statistics) |
- ▲ Orchestra average recall and precision on 1KGP-16pops across generations 90.17% recall, 90.22% precision (+15.89% and +14.03% vs. Gnomix)
- ▲ Orchestra average recall and precision on custom-35pops across generations 79.54% recall, 80.54% precision (+15.04% and +13.99% vs. RFmix)
- ▲ Orchestra accuracy exceeded 75% for all populations in 1KGP-16pops panel >75%
- ▲ Orchestra accuracy in custom-35pops panel >50% for all 35 populations, >75% for 26/35 populations
- ▲ Improvement in r2 metric at population level +24% (custom-35pops), +23.9% (1KGP-16pops)
- ▲ Orchestra outperformed other LAI methods on independent UKBB samples >90% of 103 evaluated countries
- ▲ Orchestra performance on simulated Latin American admixture 77.17% precision, 76.73% recall, outperforming other models in all three regions
- – Detection of trace ancestries matching historical migration events (SAF in Brazil, Indian in Guianas, Ashkenazi in Argentina, Japanese/Korean in Brazil/Peru)
- fold_change +15.89% recall / +14.03% precision improvement over Gnomix (1KGP-16pops benchmarking)
- fold_change +15.04% recall / +13.99% precision improvement over RFmix (custom-35pops benchmarking)
- mean 90.17% recall, 90.22% precision (Orchestra average across 6 generations, 1KGP-16pops)
- mean 79.54% recall, 80.54% precision (Orchestra average across 6 generations, custom-35pops)
- fold_change +24% improvement in r2 (custom-35pops population-level r2 metric)
- fold_change +23.9% improvement in r2 (1KGP-16pops population-level r2 metric)
- mean 77.17% precision, 76.73% recall (simulated Latin American admixture (12 generations))
- count 10,169 non-admixed individuals from 35 world regions (curated reference panel size)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper is a methods/benchmarking study for a local ancestry inference (LAI) algorithm (Orchestra), evaluated primarily via simulation. Performance is reported as precision, recall, and accuracy percentages across simulated generations of admixture and across reference populations, comparing Orchestra to three existing LAI tools (RFmix, FLARE, Gnomix). Reference panel construction used two GWAS-based filters plus two dimensionality-reduction pipelines (PCA+UMAP and t-SNE on GNN statistics) to screen samples, and ancestry tract-length distributions are summarized with boxplots/violin plots and 95% CI bands. No formal significance-testing framework (e.g., p-values from group comparisons) is described in the visible text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| GWAS-based association filtering (specific statistical test not named) | Filtering SNPs associated with genotyping platform or ancestry during reference panel curation | — | not stated |
| Precision/recall/accuracy comparison across LAI methods (descriptive metric comparison, no named inferential test) | Fig. 1b,c (1KGP-16pops N=3276; custom-35pops N=24,408) and Fig. 2a (Latin American simulations N=4068) | N=3276 (1KGP-16pops); N=24,408 (custom-35pops); N=4068 (Latin American simulations) | not stated |
-
Performance differences between Orchestra and other LAI methods are reported as point-estimate precision/recall/accuracy percentages without an accompanying measure of uncertainty or a stated significance test.↳ Could also: Bootstrap resampling or a paired test (e.g., paired t-test or Wilcoxon signed-rank test across generations/populations) could also be used — This would quantify whether the observed accuracy gains exceed the variability expected from sampling or simulation noise.
-
Accuracy/precision/recall are compared across many populations and generations without a stated correction for multiple comparisons.↳ Could also: A multiplicity correction such as Benjamini-Hochberg FDR or Bonferroni could also be applied across the population- and generation-level comparisons — This would help control the family-wise error rate or false discovery rate when many comparisons are summarized together.
-
Two separate dimensionality-reduction pipelines (PCA+UMAP and t-SNE on GNN statistics) were used to flag samples with a mismatch between reported and inferred ancestry.↳ Could also: A formal outlier-detection method with a quantitative threshold (e.g., Mahalanobis distance or a Gaussian mixture model) could also be used — This would provide an explicit, reproducible numeric cutoff for sample exclusion alongside the visual clustering approach.
-
SNPs were filtered for platform/ancestry association by ranking GWAS results and removing those in the 'top and low end', without a stated significance threshold.↳ Could also: A defined p-value or FDR threshold for the GWAS filtering step could also be used — This would make the batch-effect filtering criterion explicit and directly reproducible.
-
Ancestry tract-length distributions are shown with 95% CI bands in some figures and boxplot/violin summaries (median, quartiles, min–max) in others.↳ Could also: A single consistent dispersion measure (e.g., IQR or SD alongside the median) across all tract-length figures could also be used — This would standardize how spread is conveyed across the different ancestry-tract comparisons.
-
Relationships among the 35 reference populations were visualized by projecting a pairwise distance matrix with the SMACOF algorithm.↳ Could also: Alternative ordination approaches such as classical MDS, PCoA, or a neighbor-joining tree could also be used — These provide complementary ways to represent between-population genetic distance and may support hierarchical or branching relationships in addition to a 2D projection.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 40379651
Paper: Lerga-Jaso, J. et al. Tracing human genetic histories and natural selection with precise local ancestry inference. Nat Commun 16, 4422 (2025). DOI 10.1038/s41467-025-59936-3 · PMCID PMC12084304.
Tool introduced: Orchestra — a deep-learning local ancestry inference (LAI) method (two-stage: a base layer computing recombination-distance over genomic windows + a smoothing layer with conv/attention). Authors: Omicsedge (Lerga-Jaso, Novković, Unnikrishnan, …, Yazdi).
CORRECTION to BRIEF defaults
The BRIEF template listed Code: github.com/odelaneau/shapeit4 and
Data: geo:GSE80534. Both are wrong template defaults for this paper:
- Real code repo: https://github.com/omicsedge/orchestra-paper
(HEAD
0c70c2f1d4dfda8dad81899e32922ba646e25061, default branchmaster, Non-Commercial academic license). shapeit4 is unrelated (a phasing tool). - Data: the paper uses public reference panels (1000 Genomes 30x, HGDP, SGDP)
- GWAS/selection data; GSE80534 (a 506-sample SNP-array GEO series) is unrelated and is NOT cited by this paper.
In scope (pipeline-derived, attempted)
The repo ships a self-contained, runnable toy reproduction with bundled data and canonical reference results — the ideal reproduction target.
| ID | Result | Pipeline | Comparison target |
|---|---|---|---|
| R1 | Build the exact shipped environment (Singularity.def → orchestra.sif) | apptainer build | container builds & runs |
| R2 | Orchestra LAI pipeline end-to-end on toy data: SLiM simulation → base+smooth training (example-0.01, 100 epochs/chr-pair) → inference on 20 Admixed Mexicans, 3-way reference (Iberian/EUR, NativeAmerican/NAM, SubSaharanAfrican/AFR) |
simulation→training→inference |
produces summary_results/ancestry.tsv; MXL global ancestry should be ≈ Native American + European + small African (literature ≈47% NAM / 48% EUR / 5% AFR) |
| R3 | Downstream selection scan on admixed Mexicans: LAD score, Fadm score, genome-wide Manhattan plot of adaptive-admixture signals (prepare_gt_data/ Step1–5; replicates Cuadros-Espinoza 2022 selection signal using Orchestra LAI) |
R + PLINK2 + Beagle + Picard | repo prepare_gt_data/canonical_results/{LAD_score.png, Fadm_score.png, ManhattanPlot.GWsignals_selection.Mexicans_1kgp.png} |
R1+R2 = the ~80% floor (fully self-contained, all data shipped in repo). R3 = harder stretch (needs 1000G download + liftover/imputation + extra tools; canonical PNGs shipped for comparison).
Out of scope (not pipeline-reproducible from shipped artifacts)
- Headline accuracy benchmark of Orchestra vs RFMix/FLARE/Gnomix on 35 populations / >10,000 single-origin individuals — the full training panel (HGDP+SGDP+1000G+proprietary) and trained production weights are not shipped.
- Ashkenazi Jewish origin analysis; Latin-American demographic-history inference; trace Scandinavian ancestry / Viking immune-region selection in British samples — depend on large/custom (partly access-restricted) cohorts and the production model, not the toy artifacts.
Compute plan
All compute on «our HPC» («infra»). Build container on front1 (internet); run
sim→train→infer as a SLURM job on std/big (CPU; paper quotes ~5 h training +
20 min inference on 48 cores/192 GB). All data/intermediates on «infra» work dir
«path».
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.