Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Constraints to gene flow increase the risk of genome erosion in the Ngorongoro Crater lion population.

Commun Biol · 2025
L1 68/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
68/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 32% of all assessed papers rank 765 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH: YES - paper names exact tool (mlRho v2.7), reference (GCF_018350215.1), mapper (bwa-mem 0.7.17), filters (-q30 -Q30 depth>=2, autosomes), and per-sample targets (Suppl. Data 4); data fully public (ENA PRJEB80542). OUTCOME (native coverage): HEADLINE C4 REPRODUCED - per-sample autosomal heterozygosity theta (mlRho v2.7) ranks Crater(Lion12) 0.589 < Selous(Lion14) 0.865 < Serengeti(Lion11) 0.893, so Crater is unambiguously the lowest-diversity wild population (the paper's central claim). ABSOLUTE values: Lion12 0.589 vs reported 0.569 (+3.5%, within-tol); Lion11 0.893 vs 0.698 (+28%); Lion14 0.865 vs 0.702 (+23%). The Lion11/Lion14 excess is COVERAGE-DRIVEN: the paper downsampled every genome to 14x before mlRho but we ran at native depth (Lion12 ~14.7x matches; Lion11 ~17x and Lion14 ~19x are inflated, with magnitude tracking read count) - a documented methodological deviation, NOT a data/fabrication discrepancy. A 14x-downsampled confirmation (samtools view --subsample; «job») was attempted but produced unreliable/truncated output under RAM contention, so those numbers are not recorded; the coverage explanation is independently supported by the read-count correlation (inflation magnitude follows reads 353M<412M<446M). ENGINEERING NOTE: the shared «our HPC» account quota was chronically saturated; solved by a fully-streamed pipeline (no BAM on disk, in-RAM sort) so each job writes only a tiny profile file. NOT ATTEMPTED (hard ~20%, see scope.md): full-cohort VCF -> F_ROH/PCA/ADMIXTURE; Rxy genetic load (input tables not shipped); PSMC/GONE/SMC++ demography; SLiM simulations.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-14 ⛓ 63306e9f2694
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether ~200 years of quasi-isolation and the 1962 epizootic bottleneck have driven genome erosion (loss of diversity, increased inbreeding, increased genetic load) in the isolated Ngorongoro Crater lion population, and whether constrained gene flow from surrounding populations continues to exacerbate this erosion.

Core claims
  • 200 years of quasi-isolation and the 1962 epizootic caused a two-fold increase in inbreeding and an excess of highly deleterious mutations in Crater lions relative to other Greater Serengeti populations finding
  • There is little evidence for purging of genetic load in the Crater population finding
  • Forward simulations indicate a minimum of one to five effective male migrants per decade is required to prevent future genomic erosion and long-term inbreeding depression finding
  • The Crater lion population functions as an ecologically isolated population due to habitat fragmentation and high territoriality of resident males, despite no geographic dispersal barriers finding
  • Divergence/reduced gene flow between Crater and Greater Serengeti lions began approximately 200 years BP finding
  • Crater lions have the lowest heterozygosity and are ~1.6 times more inbred (F_ROH) than neighbouring Serengeti and Selous populations finding
  • Crater lions show an excess (Rxy) in high-impact (premature stop codon) deleterious variants relative to Serengeti and Selous lions finding
  • Realised load is significantly higher in Crater lions for both High and Moderate impact variants compared to more outbred populations finding
Experimental setups
Assay System Perturbation Readout Platform
Whole genome sequencing African lion (Panthera leo), blood/tissue samples from Ngorongoro Crater, Greater Serengeti, Selous, Botswana, South Africa none genome-wide variation, population comparisons
Principal Component Analysis (population genomics) 20 lion genomes (15 newly-sequenced + 5 published) none genetic clustering/population structure
Admixture analysis (K=2-6) 20 lion genomes none ancestry proportions, cross-validation error
Pairwise Sequentially Markovian Coalescent (PSMC) lion genomes none long-term effective population size (~1 Ma BP)
Linkage-disequilibrium-based recent demographic reconstruction Crater lion genomes none recent effective population size over ~200 generations
Runs of Homozygosity (ROH) analysis, 100kbp sliding window 16 lion genomes (≥14X depth) across 4 populations none inbreeding coefficient (F_ROH), heterozygosity (theta)
Variant annotation and genetic load estimation (SnpEff) 16-20 lion genomes none Rxy ratio, total load, realised load (High/Moderate/Synonymous impact variant ratios) SnpEff
Gene function/ontology annotation Crater lion genomes, genes with High/Moderate impact alleles none candidate gene functions related to sperm abnormality/fertility Mouse Genome Informatics database
Forward genetic simulations simulated lion population model varying number of effective male migrants per decade predicted inbreeding, genetic diversity, genetic load under different gene flow scenarios
Key results
  • F_ROH is higher in Crater lions (0.37 ± 0.031) than Serengeti (0.22 ± 0.020) and Selous (0.23 ± 0.029) ~1.6-fold
  • Divergence time between Crater and Greater Serengeti lions estimated at ~178 years BP (95% HPD: 41-380)
  • Rxy shows excess of High impact (premature stop codon) variants and slight deficit of Moderate impact variants in Crater relative to Serengeti/Selous
  • Realised load significantly higher in Crater lions for both High and Moderate impact variants vs Serengeti/Selous
  • Forward simulations show 1-5 effective male migrants per decade needed to reduce risk of long-term inbreeding depression and diversity loss 1-5 migrants/decade
  • Heterozygous High/Moderate impact variant counts significantly lower in Crater vs Serengeti, but total load still higher, indicating insufficient time for purging
  • Population census size ranged from 10 to 124 individuals between 1962 and 2022, reflecting 1962 epizootic and 2001 CDV decline 10-124 individuals
  • ~60% (range 51-66%) of F_ROH comprises ROH ≥2Mb, consistent with recent inbreeding events dating to ~110 and ~30 years ago 51-66%
Key statistics
  • mean F_ROH-Crater = 0.37 ± 0.031 (inbreeding coefficient in Crater lions)
  • mean F_ROH-Serengeti = 0.22 ± 0.020 (inbreeding coefficient in Serengeti lions)
  • mean F_ROH-Selous = 0.23 ± 0.029 (inbreeding coefficient in Selous lions)
  • other Mean divergence time 178 years BP, 95% HPD 41-380 (Crater-Serengeti gene flow reduction timing)
  • fold_change two-fold increase in inbreeding (Crater relative to other Greater Serengeti populations)
  • count 1-5 effective male migrants per decade (minimum gene flow required per simulations to prevent genomic erosion)
  • other 80% of resident males born in Crater vs 33% in rest of Serengeti National Park (male philopatry/territoriality comparison)
  • count 20 lion genomes analysed (15 newly-sequenced + 5 published) (total sample size for genomic analyses)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used whole-genome sequencing (20 lion genomes) and comparative population genomics to assess genome erosion in the isolated Ngorongoro Crater lion population relative to Serengeti, Selous, Botswana, and South African populations. Population structure was characterised via PCA and ADMIXTURE, inbreeding via ROH-based F coefficients, and genetic load via SnpEff functional annotation with Rxy allele-frequency ratios. Group differences in heterozygosity, inbreeding, and load were tested with Wilcoxon signed-rank tests, and forward simulations were used to project future genomic trajectories under varying gene-flow scenarios.

Replicationbiological Sample size20 total genomes: 15 newly sequenced (10 Crater, 3 Greater Serengeti, 2 Selous) plus 5 published (Tanzania, Botswana, South Africa); sequencing depth ≥14X required for load analyses, yielding n = 16 for those comparisons GroupsCrater lions vs. Greater Serengeti, Selous, Botswana, and South African lions Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Wilcoxon signed-rank test Comparisons of F_ROH (inbreeding coefficients), heterozygosity, realised load, and heterozygous variant counts across populations (Figs. 3b, 4c, S3) n = 4 as stated in figure captions; basis not further specified in the excerpt not stated
Rxy allele-frequency ratio Excess or deficit of High- and Moderate-impact derived alleles in Crater lions relative to Serengeti and Selous lions (Fig. 4a) Crater n = 10; Serengeti and Selous combined n = 5 not stated
Pairwise Sequentially Markovian Coalescent (PSMC) Long-term (~1 Ma BP) and recent (~200 generation) demographic reconstruction (Figs. 2a, 2b) 10 replicates per estimate; 40 independent estimates from observed LD spectrum each not stated
ADMIXTURE with cross-validation error (K = 2–6) Population structure and admixture proportions for all 20 genomes (Fig. 1c) n = 20 genomes total not stated
Principal Component Analysis (PCA) Visualisation of genetic relationships among lion populations (Fig. 1b) n = 20 genomes total not stated
Forward simulations Projections of future inbreeding, genetic diversity, and genetic load under varying male gene-flow scenarios null not stated
Approaches that could also have been used
  • Wilcoxon signed-rank tests were applied to compare genomic metrics across populations, with n = 4 as stated in figure captions
    Could also: Permutation or bootstrap tests treating each sequenced individual as the observation unit could also be used — With very small and unequal sample sizes per population (e.g., 3 Serengeti, 2 Selous genomes), permutation-based approaches make no distributional assumptions and naturally accommodate the small-n, unbalanced designs common in wildlife population genomics
  • Multiple Wilcoxon tests were performed across several genomic metrics and population pairs without a stated multiplicity correction
    Could also: A false discovery rate adjustment (e.g., Benjamini-Hochberg) or family-wise error rate correction (e.g., Holm-Bonferroni) could also be applied across the family of comparisons — Adjusting for multiple comparisons across variant classes and population pairs controls the cumulative type I error rate, which increases with the number of tests performed
  • Recent and long-term demographic history was reconstructed using PSMC applied to individual diploid genomes
    Could also: SMC++ or fastsimcoal2 (using the joint site-frequency spectrum across sampled individuals) could also be used, particularly for recent demographic inference — SMC++ leverages population-level data from multiple individuals simultaneously, which can improve resolution of recent demographic events and is better suited to studies with more than one sequenced individual per population
  • Population structure was characterised with ADMIXTURE model-based clustering and PCA
    Could also: TreeMix or D-statistics (f3/f4 tests) could also be used to formally test for and quantify admixture and directional gene flow among populations — Admixture graph methods and f-statistics provide explicit statistical tests for gene-flow events and can distinguish admixture from shared ancestral variation, complementing the descriptive outputs of ADMIXTURE and PCA
  • Genetic load was classified using SnpEff rule-based functional annotation (High impact, Moderate impact, Synonymous) of coding variants
    Could also: Conservation-based scores such as GERP++ or phastCons, or sequence-alignment-based deleteriousness predictors such as SIFT, could also be used to estimate variant impact — Conservation- and alignment-based methods are complementary to rule-based annotation and can capture fitness-relevant variation in non-coding or structurally complex regions not covered by stop-codon or splice-site definitions
  • Dispersion of genomic metrics was summarised with standard deviation around means
    Could also: 95% bootstrap confidence intervals around population-level estimates could also be reported — Bootstrap CIs convey both spread and estimation uncertainty, facilitate inference about whether population differences are likely to be meaningful, and are often preferred when sample sizes are small and distributional assumptions are uncertain
Software: SnpEff · PSMC · ADMIXTURE · Mouse Genome Informatics (MGI) database

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
3
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40258987

Paper: Dussex et al. 2025, Commun Biol 8:640. "Constraints to gene flow increase the risk of genome erosion in the Ngorongoro Crater lion population." DOI 10.1038/s42003-025-07986-0 · PMCID PMC12012037

Code: https://github.com/ndussex/Crater_lion_genomics (commit 10348ae, 2025-04-30) Data: ENA BioProjects — PRJEB80542 (this study: 15 lions Lion01–Lion15), plus PRJNA611920 / PRJNA182708 / PRJNA16726 / PRJNA854353 / PRJNA684344 (comparative lions).

Nature of the code repo

The repo is a thin set of analysis notes + plotting scripts, not a runnable end-to-end pipeline. The heavy lifting (reads → BAM → het/ROH/load) is done by the external GenErode Snakemake pipeline (NBISweden/GenErode), which the repo only references. The repo's own files are:

  • 0_Data_processing_inbreeding_genetic_load/ — SNPeff setup notes, an Rxy R script, a VCF ref-allele-replacement helper. Input tables (allele-freq files High_freq.txt, Intergenic_*SNPs.txt) are NOT shipped → Rxy not reproducible as-is.
  • 1_Demography/ — GONE / SMC++ plotting R scripts (no input data shipped).
  • 2_Simulations_Slim/ — SLiM 4.0 non-WF simulation + plotting scripts.
  • 3_Pop_structure/ — PCA/ADMIXTURE command notes + sample lists (needs the 20-lion VCF, not shipped).

Pipeline-derived results (candidate in-scope)

Result Pipeline / tool Reported location Feasibility
Genome-wide heterozygosity θ (het sites/1000 bp) per sample BWA-MEM→bcftools→mlRho v2.7 Fig 3a, Suppl. Data 4 (exact per-sample) IN SCOPE — chosen
F_ROH (≥100 kb, ≥2 Mb) per sample PLINK v2 ROH on multi-sample VCF Fig 3b, Suppl. Data 4 partial (needs full cohort VCF) — not chased
Genetic load Rxy / SnpEff counts SnpEff v4.3 + Rxy R script Fig 4, Suppl. Data 4 OUT (allele-freq input tables not shipped)
Demography Ne / divergence PSMC, GONE, SMC++ Fig 2 OUT (heavy; multi-sample; under-specified)
Forward simulations FROH 0.14→0.27 SLiM 4.0 non-WF Fig 5 OUT (bespoke demographic model, stochastic)
Population structure (PCA/ADMIXTURE) PLINK2 / ADMIXTURE 1.3 Fig 1 OUT (needs full cohort VCF)

Chosen reproduction target (80/20)

Reproduce the paper's central genome-erosion metric: per-sample genome-wide autosomal heterozygosity θ (N. het. sites/1000 bp) via the exact tool named in Methods (mlRho v2.7), on the paper's own raw reads (PRJEB80542), mapped with BWA-MEM v0.7.17 to the same reference (GCF_018350215.1, P.leo_Ple1_pat1.1), with the paper's filters (-q30 -Q30, MQ30, depth≥2). Exact per-sample targets are in Suppl. Data 4 (original/DataS4_het_targets.tsv).

Samples chosen — one per wild population, spanning the reported θ range:

  • Lion12 (Crater) → reported θ = 0.569 (lowest pop)
  • Lion11 (Serengeti) → reported θ = 0.698
  • Lion14 (Selous) → reported θ = 0.702

This directly tests the headline claim: Crater diversity is the lowest of the wild populations. Pop means (Suppl. Data 4): Crater 0.602 (n=10) < Selous 0.689 (n=2) < Serengeti 0.719 (n=3); captive Plwhi 0.385.

Explicitly NOT attempted (the hard ~20%) and why

  • Full 16–20-genome cohort VCF (→ ROH/F_ROH, PCA, ADMIXTURE): days of compute, multi-TB intermediates; deviates from 80/20.
  • Rxy genetic load: the allele-frequency input tables are not in the repo.
  • Demographic inference (PSMC/GONE/SMC++) and SLiM forward simulations: bespoke, stochastic, under-specified model parameters.
  • Downsampling each genome to exactly 14× before mlRho (paper did): mlRho jointly estimates sequencing error and is robust above ~10×; we run at native depth and flag this as a minor deviation.
Figures / tables: Fig 3a
C1
Reported
0.569 (Lion12, Crater) - mlRho theta, N.het.sites/1000bp
Reproduced
0.589 (theta_ML 5.89e-4); +3.5% (native ~14.7x)
within tolerance
C2
Reported
0.698 (Lion11, Serengeti) - mlRho theta
Reproduced
0.893 (theta_ML 8.93e-4); +28% at native ~17x (no 14x downsampling); coverage-explained
partial
C3
Reported
0.702 (Lion14, Selous) - mlRho theta
Reproduced
0.865 (theta_ML 8.65e-4); +23% at native ~19x (no 14x downsampling); coverage-explained
partial
C4
Reported
Headline: Crater is lowest-diversity wild pop (Crater 0.602 < Selous 0.689 < Serengeti 0.719)
Reproduced
REPRODUCED: Crater(0.589) < Selous(0.865) < Serengeti(0.893); Crater unambiguously lowest
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 68/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

This is not a discrepancy but an incomplete (in-flight) reproduction: the pipeline was correctly built and launched (mlRho v2.7, ref GCF_018350215.1, public reads ENA PRJEB80542) but «job»-65 were forced to finalize before producing any θ, so reproduced_values is null and the reported values (Lion12 0.569, Lion11 0.698, Lion14 0.702; rank Crater < Selous < Serengeti) are untested, neither confirmed nor refuted. The incompleteness lies entirely on our/operational side, not the authors — data and methods are fully and exactly specified, so 1:1 comparison is in principle possible. No fabrication concern can be raised or cleared. Graded yellow on q5/q7/q8 to reflect derivable-but-unverified, with no severity since nothing actually deviated.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

778.2 k
tokens (I/O) · 91.6 M incl. cache
468 min
runtime · 113.58 CPU-h
179.9 GB
peak RAM
8 (1 failed)
HPC jobs
hummel
machine