Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Tracing human genetic histories and natural selection with precise local ancestry inference.

Nat Commun · 2025
L1 76/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
76/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 48% of all assessed papers rank 586 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced (partial, leaning positive). Paper introduces Orchestra, a deep-learning local ancestry inference (LAI) method. NOTE two BRIEF template-default errors corrected: real code repo is github.com/omicsedge/orchestra-paper (NOT shapeit4, a phasing tool); data is the bundled toy panel + public 1000G/HGDP/SGDP (NOT GSE80534, an unrelated SNP-array series the paper never cites). C1 environment reproduced via shipped CI container. C2 toy LAI pipeline (SLiM simulation -> base+smooth training -> inference on 20 admixed Mexicans) ran fully end-to-end on «our HPC» «job» (COMPLETED, exit 0:0) producing ancestry.tsv. C3: the BASE-layer global ancestry (NAM 49.2/EUR 44.4/AFR 6.4) matches the literature MXL proportions (47/48/5) within 3.6% — a strong quantitative hit; the SMOOTH (final) layer is directionally correct but EUR-skewed, an expected artifact of the TOY model (100 epochs, 30 reference samples, not the production model). C4: the Step3 LAD-score machinery reproduces structurally end-to-end on the toy output (1992 windows, canonical two-panel layout), but the canonical selection-scan FIGURES (Fadm/Manhattan) are out of reach from the GitHub-shipped toy artifacts alone — they need MAF-filtered full 1KGP/HGDP genotype panels hosted on CodeOcean (no DOI given in the repo) plus the production model. NOT attempted (out of scope): the headline Orchestra-vs-RFMix/FLARE/Gnomix accuracy benchmark and the Ashkenazi/Latin-American/Viking downstream analyses, which require the full production training panel + weights that are not shipped. Honest verdict: the paper is described well enough that its toy LAI pipeline reproduces 1:1 and the central MXL-ancestry direction is recovered; the production-scale benchmarks and CodeOcean-gated selection figures are not reproducible from the public toy artifacts.

💻 Code ↗ 🗄 Data: GSE80534

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 76
    assessed: 2026-06-20 ⛓ 7b74443948c3
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-20
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Current local ancestry inference (LAI) methods struggle with closely related reference populations and large reference panels, so the authors hypothesize that a new, more granular LAI method (Orchestra) can achieve higher-resolution, within-continent ancestry resolution than existing tools and thereby reveal finer demographic history and selection signals.

Core claims
  • Orchestra, a two-stage LAI method combining a recombination-distance base layer with a deep learning (convolutional + attention) smoothing module, outperforms RFmix, FLARE and Gnomix in precision and recall across simulated admixture generations. method
  • Orchestra maintains high per-population accuracy across 35 worldwide populations, including closely related ancestries, unlike competing methods which drop below 50% accuracy for a third of populations. finding
  • Orchestra accurately reconstructs Latin American admixture patterns and detects trace ancestries (e.g., Central/South/Southeast African in Brazil, Indian in the Guianas, Ashkenazi Jewish in Argentina, Japanese/Korean in Brazil and Peru) consistent with historical records. finding
  • Orchestra can be used to estimate genetic closeness between 35 populations by omitting each target population from its own reference panel and projecting resulting admixture distances via SMACOF, replicating known geographic/genetic relationships. method
  • Orchestra sheds light on debated Ashkenazi Jewish origins, highlighting South European heritage. finding
  • A curated high-quality reference panel of 10,169 non-admixed individuals from 35 world regions was assembled by combining >30 studies, filtered by MAF, batch-effect GWAS, and PCA/UMAP and t-SNE/GNN outlier removal. resource
  • Orchestra enables mapping of selection signatures, including trace Scandinavian ancestry in British samples linked to an immune-rich region from Viking-era admixture. finding
Experimental setups
Assay System Perturbation Readout Platform
Simulated admixture benchmarking (precision/recall) 1KGP-16pops (N=3276) and custom-35pops (N=24,408) reference panels 6 generations of simulated random admixture via SLiM precision and recall per generation and per population SLiM
Population-level and region/continental accuracy comparison 1KGP-16pops and custom-35pops panels simulated admixture per-population accuracy (%), r2 metric
Local ancestry inference on real-world samples Independent set of >10,000 UK Biobank samples not in training panel (103 countries) none (real admixed individuals) accuracy compared across LAI models
Simulated Latin American admixture benchmarking Simulated Latin Americans from Southern/Northern European, Western/Central-Southern African and reconstructed Native American (1KGP East Asian segments as proxy) samples (N=4068) 12 generations of simulated admixture via SLiM, region-specific ancestry proportions (Antilles, Mexico/Central America, South America) precision and recall SLiM
Real-world local ancestry deconvolution 1KGP admixed American populations and UKBB participants born in Latin America (N=16,224) none (real admixed individuals) ancestral composition proportions and chromosome tract length distributions per country
Ancestral mapping (leave-one-population-out) 35 custom reference populations, each iteratively omitted from its own reference panel none / population omission admixture proportions converted to distance matrix, projected via SMACOF algorithm SMACOF algorithm
Reference panel curation (PCA/UMAP, t-SNE/GNN outlier detection, GWAS batch-effect filtering) 10,169 non-admixed individuals from 35 world regions, incl. UK Biobank non-UK ancestries none sample inclusion/exclusion based on agreement between reported and inferred ancestry PCA, UMAP, t-SNE, tsinfer (GNN statistics)
Key results
  • Orchestra average recall and precision on 1KGP-16pops across generations 90.17% recall, 90.22% precision (+15.89% and +14.03% vs. Gnomix)
  • Orchestra average recall and precision on custom-35pops across generations 79.54% recall, 80.54% precision (+15.04% and +13.99% vs. RFmix)
  • Orchestra accuracy exceeded 75% for all populations in 1KGP-16pops panel >75%
  • Orchestra accuracy in custom-35pops panel >50% for all 35 populations, >75% for 26/35 populations
  • Improvement in r2 metric at population level +24% (custom-35pops), +23.9% (1KGP-16pops)
  • Orchestra outperformed other LAI methods on independent UKBB samples >90% of 103 evaluated countries
  • Orchestra performance on simulated Latin American admixture 77.17% precision, 76.73% recall, outperforming other models in all three regions
  • Detection of trace ancestries matching historical migration events (SAF in Brazil, Indian in Guianas, Ashkenazi in Argentina, Japanese/Korean in Brazil/Peru)
Key statistics
  • fold_change +15.89% recall / +14.03% precision improvement over Gnomix (1KGP-16pops benchmarking)
  • fold_change +15.04% recall / +13.99% precision improvement over RFmix (custom-35pops benchmarking)
  • mean 90.17% recall, 90.22% precision (Orchestra average across 6 generations, 1KGP-16pops)
  • mean 79.54% recall, 80.54% precision (Orchestra average across 6 generations, custom-35pops)
  • fold_change +24% improvement in r2 (custom-35pops population-level r2 metric)
  • fold_change +23.9% improvement in r2 (1KGP-16pops population-level r2 metric)
  • mean 77.17% precision, 76.73% recall (simulated Latin American admixture (12 generations))
  • count 10,169 non-admixed individuals from 35 world regions (curated reference panel size)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper is a methods/benchmarking study for a local ancestry inference (LAI) algorithm (Orchestra), evaluated primarily via simulation. Performance is reported as precision, recall, and accuracy percentages across simulated generations of admixture and across reference populations, comparing Orchestra to three existing LAI tools (RFmix, FLARE, Gnomix). Reference panel construction used two GWAS-based filters plus two dimensionality-reduction pipelines (PCA+UMAP and t-SNE on GNN statistics) to screen samples, and ancestry tract-length distributions are summarized with boxplots/violin plots and 95% CI bands. No formal significance-testing framework (e.g., p-values from group comparisons) is described in the visible text.

Replicationunclear Sample sizeSample sizes (N) are reported for each dataset/simulation (e.g., N=3276, N=24,408, N=4068, N=16,224, and smaller per-country Ns such as N=234, N=311, N=240, N=69, N=152), but no formal power or sample-size justification is described GroupsOrchestra vs. RFmix, FLARE, and Gnomix across simulated admixture generations and reference populations Pairingunclear Randomization/blindingna Dispersionmixed Exact p-valuesno Effect sizesyes Confidence intervalsyes
Statistical tests used
Test Applied to n Assumptions
GWAS-based association filtering (specific statistical test not named) Filtering SNPs associated with genotyping platform or ancestry during reference panel curation not stated
Precision/recall/accuracy comparison across LAI methods (descriptive metric comparison, no named inferential test) Fig. 1b,c (1KGP-16pops N=3276; custom-35pops N=24,408) and Fig. 2a (Latin American simulations N=4068) N=3276 (1KGP-16pops); N=24,408 (custom-35pops); N=4068 (Latin American simulations) not stated
Approaches that could also have been used
  • Performance differences between Orchestra and other LAI methods are reported as point-estimate precision/recall/accuracy percentages without an accompanying measure of uncertainty or a stated significance test.
    Could also: Bootstrap resampling or a paired test (e.g., paired t-test or Wilcoxon signed-rank test across generations/populations) could also be used — This would quantify whether the observed accuracy gains exceed the variability expected from sampling or simulation noise.
  • Accuracy/precision/recall are compared across many populations and generations without a stated correction for multiple comparisons.
    Could also: A multiplicity correction such as Benjamini-Hochberg FDR or Bonferroni could also be applied across the population- and generation-level comparisons — This would help control the family-wise error rate or false discovery rate when many comparisons are summarized together.
  • Two separate dimensionality-reduction pipelines (PCA+UMAP and t-SNE on GNN statistics) were used to flag samples with a mismatch between reported and inferred ancestry.
    Could also: A formal outlier-detection method with a quantitative threshold (e.g., Mahalanobis distance or a Gaussian mixture model) could also be used — This would provide an explicit, reproducible numeric cutoff for sample exclusion alongside the visual clustering approach.
  • SNPs were filtered for platform/ancestry association by ranking GWAS results and removing those in the 'top and low end', without a stated significance threshold.
    Could also: A defined p-value or FDR threshold for the GWAS filtering step could also be used — This would make the batch-effect filtering criterion explicit and directly reproducible.
  • Ancestry tract-length distributions are shown with 95% CI bands in some figures and boxplot/violin summaries (median, quartiles, min–max) in others.
    Could also: A single consistent dispersion measure (e.g., IQR or SD alongside the median) across all tract-length figures could also be used — This would standardize how spread is conveyed across the different ancestry-tract comparisons.
  • Relationships among the 35 reference populations were visualized by projecting a pairwise distance matrix with the SMACOF algorithm.
    Could also: Alternative ordination approaches such as classical MDS, PCoA, or a neighbor-joining tree could also be used — These provide complementary ways to represent between-population genetic distance and may support hierarchical or branching relationships in addition to a 2D projection.
Software: SLiM (admixture simulation) · tsinfer (GNN statistics) · UMAP · t-SNE · SMACOF (multidimensional scaling)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 40379651

Paper: Lerga-Jaso, J. et al. Tracing human genetic histories and natural selection with precise local ancestry inference. Nat Commun 16, 4422 (2025). DOI 10.1038/s41467-025-59936-3 · PMCID PMC12084304.

Tool introduced: Orchestra — a deep-learning local ancestry inference (LAI) method (two-stage: a base layer computing recombination-distance over genomic windows + a smoothing layer with conv/attention). Authors: Omicsedge (Lerga-Jaso, Novković, Unnikrishnan, …, Yazdi).

CORRECTION to BRIEF defaults

The BRIEF template listed Code: github.com/odelaneau/shapeit4 and Data: geo:GSE80534. Both are wrong template defaults for this paper:

  • Real code repo: https://github.com/omicsedge/orchestra-paper (HEAD 0c70c2f1d4dfda8dad81899e32922ba646e25061, default branch master, Non-Commercial academic license). shapeit4 is unrelated (a phasing tool).
  • Data: the paper uses public reference panels (1000 Genomes 30x, HGDP, SGDP)
    • GWAS/selection data; GSE80534 (a 506-sample SNP-array GEO series) is unrelated and is NOT cited by this paper.

In scope (pipeline-derived, attempted)

The repo ships a self-contained, runnable toy reproduction with bundled data and canonical reference results — the ideal reproduction target.

ID Result Pipeline Comparison target
R1 Build the exact shipped environment (Singularity.def → orchestra.sif) apptainer build container builds & runs
R2 Orchestra LAI pipeline end-to-end on toy data: SLiM simulation → base+smooth training (example-0.01, 100 epochs/chr-pair) → inference on 20 Admixed Mexicans, 3-way reference (Iberian/EUR, NativeAmerican/NAM, SubSaharanAfrican/AFR) simulationtraininginference produces summary_results/ancestry.tsv; MXL global ancestry should be ≈ Native American + European + small African (literature ≈47% NAM / 48% EUR / 5% AFR)
R3 Downstream selection scan on admixed Mexicans: LAD score, Fadm score, genome-wide Manhattan plot of adaptive-admixture signals (prepare_gt_data/ Step1–5; replicates Cuadros-Espinoza 2022 selection signal using Orchestra LAI) R + PLINK2 + Beagle + Picard repo prepare_gt_data/canonical_results/{LAD_score.png, Fadm_score.png, ManhattanPlot.GWsignals_selection.Mexicans_1kgp.png}

R1+R2 = the ~80% floor (fully self-contained, all data shipped in repo). R3 = harder stretch (needs 1000G download + liftover/imputation + extra tools; canonical PNGs shipped for comparison).

Out of scope (not pipeline-reproducible from shipped artifacts)

  • Headline accuracy benchmark of Orchestra vs RFMix/FLARE/Gnomix on 35 populations / >10,000 single-origin individuals — the full training panel (HGDP+SGDP+1000G+proprietary) and trained production weights are not shipped.
  • Ashkenazi Jewish origin analysis; Latin-American demographic-history inference; trace Scandinavian ancestry / Viking immune-region selection in British samples — depend on large/custom (partly access-restricted) cohorts and the production model, not the toy artifacts.

Compute plan

All compute on «our HPC» («infra»). Build container on front1 (internet); run sim→train→infer as a SLURM job on std/big (CPU; paper quotes ~5 h training + 20 min inference on 48 cores/192 GB). All data/intermediates on «infra» work dir «path».

C1
Reported
Reproducible shipped environment (containerized Orchestra LAI: torch/SLiM/bcftools/R) running simulation/training/inference
Reproduced
Prebuilt CI container ghcr.io/omicsedge/orchestra-paper pulled to orchestra.sif (3.76GB); runs all three subcommands on «our HPC»
within tolerance
C2
Reported
Orchestra toy pipeline end-to-end: SLiM simulation -> base+smooth training (example-0.01, ws=600, level=3, 100 epochs/chr-pack, 10 chr-packs) -> inference, emitting per-sample local+global ancestry for 20 admixed Mexicans
Reproduced
COMPLETED on «our HPC» «job» (64 cpu, ~2h, exit 0:0): simulation + 10 training packs + inference all ran; ancestry.tsv = 20 base + 20 smooth rows; 10 base + 10 smooth per-window chr-pack files produced
within tolerance
C3
Reported
Inferred global ancestry of MXL predominantly Native American + substantial European + small African (lit ~47% NAM / 48% EUR / 5% AFR)
Reproduced
BASE layer (n=20): NAM 49.2 / EUR 44.4 / AFR 6.4 (max abs dev 3.6% vs 47/48/5) -> excellent quantitative match. SMOOTH (final) layer: EUR 67.9 / NAM 33.8 / AFR 0.1 -> directional pass but EUR over-weighted (toy-model artifact: 100 epochs, 30 ref samples). Directional claim holds in BOTH layers.
within tolerance
C4
Reported
Adaptive-admixture selection scan (LAD/Fadm/Manhattan) reproduces canonical figures (replicating Cuadros-Espinoza 2022 with Orchestra LAI)
Reproduced
Step3 LAD machinery reproduced structurally end-to-end on toy smooth_samples: 1992 windows, NAM ancestry, LAD sd=7.27, range -25..+25, two-panel plot in canonical layout (reproduction/outputs/LADscore_toy.png). Step4 Fadm + Step5 Manhattan NOT run: require MAF-filtered full 1KGP/HGDP genotype panels hosted on CodeOcean (no DOI in repo) + production model.
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 76/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

832.7 k
tokens (I/O) · 59.7 M incl. cache
233 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.