Candidates for Balancing Selection in Leishmania donovani Complex Parasites.
The main results reproduced, with only marginal, non-material deviations.
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL reproduction, run completed as intended. EXACT: dataset integrity (93/93 PE WGS L. infantum on ENA, C2), reference assembly (BPK282A1 TriTrypDB Nov-2019, 36 contigs/32.44 Mbp, C3), and the full tool stack at the paper's exact versions (bwa0.7.17/samtools1.9/bcftools1.9/freebayes1.3.2/vcftools0.1.15, CENV). The named bwa->Freebayes1.3.2 (--min-alt-count 5) pipeline was executed on a 6-isolate subset of the paper's own new data and runs cleanly with consistent, biologically sensible output (mean 128,490 SNPs/isolate, sd 1,719; 95.9% genome breadth; 95% reads mapped; bcftools stats concurs) -> C10 reproduced as internal-consistency. NOT ATTEMPTED (honest, out-of-scope): the cohort-level headline numbers (339,367 SNPs / 14,383 indels / 283,378 variable sites / 194,351 LD-pruned / 5 ADMIXTURE populations / 24 balancing-selection genes) require full 477-genome multi-study joint calling + a manual 38->24 gene curation step; reproducing those exactly is many CPU-days across hundreds of GB from 4 external studies and the final gene count is non-deterministic. Per brief P16, running the named third-party caller on the paper's own data is a valid reproduction. All grades provisional; human reviewer decides.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 58assessed: 2026-06-18 ⛓ 20a6bf0542b1
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether balancing selection (BS) operates in Leishmania donovani species complex parasites, and if genomic signatures of BS can identify genes important for host-parasite interaction, immune evasion, or ecological adaptation, given BS has been documented in the related parasite genus Plasmodium.
- ★ Five genetically distinct, well-sampled populations exist within the L. donovani complex: EA1, EA2 (East Africa), ISC1, ISC2 (Indian subcontinent), and BM (Brazil-Mediterranean) finding
- ★ Signals of balancing selection are generally not shared between populations, consistent with transient adaptive events rather than long-term balancing selection finding
- ★ A curated set of 24 candidate genes show robust signatures of balancing selection, including zeta toxin, nodulin-like, and flagellum attachment proteins finding
- ★ NCD2 and Betascan1* tests, together with nucleotide diversity and Tajima's D, are largely complementary metrics for detecting balancing selection signals method
- ★ Nucleotide diversity is significantly elevated in genomic regions surrounding BS target genes, extending up to 250 kb from targets finding
- BM and EA2 populations show an excess of indel polymorphisms relative to SNPs, consistent with a population bottleneck/founder effect finding
- ★ A dataset of 477 sequenced L. donovani complex isolates (including 93 newly sequenced Brazilian isolates) expands prior global genomic analyses resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-genome sequencing | 93 L. infantum isolates from Piauí state, Brazil | none | SNPs and indels | — |
| Read mapping and variant calling | 477 L. donovani/L. infantum clinical isolates mapped to BPK282A1 reference genome | none | SNPs and indel polymorphisms | — |
| ADMIXTURE clustering analysis | 477 L. donovani complex genomes | none | Population assignment/ancestry proportions (K=8-11) | ADMIXTURE |
| Principal component analysis (PCA) | 477 L. donovani complex genomes (SNP data) | none | Population structure/clustering | — |
| Maximum-likelihood phylogenetic analysis | SNP alignments (477 sequences, 283,378 variable sites; BM subset 158 sequences) | none | Unrooted/midpoint-rooted phylogenetic trees | — |
| NCD2 test | 5 populations (EA1, EA2, ISC1, ISC2, BM), 1/5/10 kb genome windows | none | Score reflecting allele frequency similarity to 0.5 target (balancing selection signal) | — |
| Betascan1* test | 5 populations (EA1, EA2, ISC1, ISC2, BM), 1/5/10 kb genome windows | none | Correlated allele frequency score (balancing selection signal) | — |
| Population genetics statistics (nucleotide diversity π, Tajima's D, minor allele frequency, FST) | 5 populations, per-gene and 10 kb windows | none | π, Tajima's D, MAF, FST values | — |
- – 477 sequenced isolates yielded 339,367 SNPs and 14,383 indels genome-wide
- – Five major populations (EA1, EA2, ISC1, ISC2, BM) identified with ≥99% confidence assignment for most isolates
- – FST between populations ranges from 0.27 to 0.90; only 7% of polymorphic sites shared between two or more populations 0.27-0.90
- – Betascan1* outlier genes largely population-specific: 701 genes were outliers in ≥1 population, only 42 (6%) in ≥2 populations, only 9 (1%) in ≥3 populations 6%/1%
- – NCD2 outlier genes: 1,627 genes were outliers in ≥1 population, only 195 (11%) in more than one population 11%
- ▲ ISC1 Betascan1* outliers significantly predict higher scores in ISC2, the only cross-population enrichment found P = 2 x 10^-4
- – Screen identified 24 vetted candidate genes with robust BS signatures after filtering 38 initial candidates for hitchhikers and coverage artifacts
- ▲ EA1 population candidate genes show nucleotide diversity 34-fold higher than genome-wide median, with elevated diversity extending up to 250 kb from targets 34-fold
- count 477 sequenced isolates; 339,367 SNPs; 14,383 indels (Full dataset used for the study)
- fold_change 34-fold higher nucleotide diversity than genome-wide median (EA1 population BS candidate genes vs genome-wide)
- pvalue P = 2 x 10^-4 (ISC1 outlier enrichment predicting higher Betascan1* scores in ISC2)
- other FST 0.27-0.90 (Range of fixation index values between the five populations)
- count 24 candidate genes (Final vetted set of genes with robust balancing selection signatures)
- count 701 Betascan1* outlier genes, 42 (6%) shared in >=2 populations, 9 (1%) shared in >=3 (Cross-population sharing of Betascan1* outliers)
- count 1,627 NCD2 outlier genes, 195 (11%) shared in more than one population (Cross-population sharing of NCD2 outliers)
- pvalue Wilcoxon signed rank tests < 1.5 x 10^-4 (Significance of elevated diversity up- and downstream of BS target genes in EA1)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This population genomics study characterized 477 Leishmania donovani complex clinical isolates across five geographically defined populations using ADMIXTURE, PCA, and maximum-likelihood phylogenetics to establish population structure. Balancing selection was assessed genome-wide using the NCD2 and Betascan1* statistics in sliding windows, supplemented by Tajima's D and nucleotide diversity (π) per gene, with candidates identified by intersection of percentile-based outlier thresholds across metrics. Inter-population sharing of balancing selection signals was assessed by comparing outlier gene-set overlaps and using a rank-based enrichment test, and spatial elevation of diversity around candidate loci was tested with Wilcoxon signed-rank tests.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| ADMIXTURE clustering (likelihood-based population structure) | Assignment of all 477 isolates to K populations; cross-validation error used to select K (fig. 1A, supplementary fig. 1) | 477 isolates total; 433 non-admixed retained for downstream analyses | not stated |
| Principal component analysis (PCA) | Visualization of global population structure (fig. 1B) | 477 isolates | na |
| Maximum likelihood phylogeny | Unrooted global ML phylogeny from SNP alignment (fig. 1C) and BM ML tree (fig. 1D); bootstrap support 100% mlBP on all visible branches unless noted | 477 isolates global tree (283,378 variable sites); 158 isolates BM tree (81,018 variable sites) | not stated |
| NCD2 (normalised centred deviation) statistic (Bitarello et al. 2018) | Genome-wide balancing selection scan in each of five populations; applied in 1, 5, and 10 kb windows; outliers defined as 1st percentile (minimum NCD2) per population | Per-population n: EA1=41, EA2=18, ISC1=211, ISC2=15, BM=127 non-admixed isolates | not stated |
| Betascan1* statistic (Siewert and Voight 2017) | Genome-wide balancing selection scan in each of five populations; applied in 1, 5, and 10 kb windows; outliers defined as 99th percentile (maximum score) and upper 5% per population | Per-population n: EA1=41, EA2=18, ISC1=211, ISC2=15, BM=127 non-admixed isolates | not stated |
| Tajima's D (Tajima 1989) | Genome-wide in 10 kb windows (mean values reported in table 1) and per-gene; genes in 90th percentile used as supporting evidence for candidate identification | Per-population isolate sets across all genome windows | not stated |
| Nucleotide diversity (π) | Per 10 kb window (table 1) and per gene; 90th percentile threshold used in candidate selection; also measured in 500 kb flanking regions around BS candidate loci (fig. 3) | Per-population isolate sets | not stated |
| Wilcoxon signed-rank test | Comparison of π at successive 10 kb distances (up to 500 kb) from BS target genes vs. genome-wide π distribution in EA1 (fig. 3); significance threshold stated as p < 1.5 × 10^-4 | 20 vetted BS targets in EA1; genome-wide 10 kb window distribution | not stated |
| Rank-based enrichment test (specific test not named in excerpt) | Testing whether 5% Betascan1* outliers from one population show elevated Betascan1* scores in other populations (supplementary fig. 18); significant result for ISC1 predicting ISC2 scores (P = 2 × 10^-4) | Not explicitly stated | not stated |
| Fixation index (FST) | Between-population genetic differentiation for all population pairs (supplementary table 2); range 0.27–0.90 reported | All five populations | not stated |
-
Candidate balancing selection loci were identified by intersecting percentile-based outlier thresholds across multiple statistics (NCD2, Betascan1*, π, Tajima's D) without a formal genome-wide multiple-testing correction↳ Could also: A permutation-based false discovery rate (FDR) or Benjamini-Hochberg correction applied across all genomic windows tested could also be used — With tens of thousands of genomic windows tested across five populations, a genome-wide FDR would provide a formal estimate of the expected false-positive rate among candidate windows, making the confidence in the candidate set more precisely quantifiable alongside the percentile approach
-
Inter-population sharing of balancing selection signals was assessed by counting overlap in the top-5% Betascan1* outlier gene sets between populations↳ Could also: A hypergeometric test or Fisher's exact test with a permutation-derived null (shuffling genomic coordinates to preserve linkage structure) could also quantify whether overlap exceeds chance — A permutation-based null accounts for non-independence of adjacent genomic windows and uneven gene density across chromosomes, providing a genome-structure-aware significance estimate for overlap
-
Genetic differentiation between populations was summarized with FST↳ Could also: Dxy (absolute nucleotide divergence) or the population branch statistic (PBS) could also be reported alongside FST — FST is sensitive to within-population diversity levels; given the large range of π across these populations (~6 to ~424 × 10^-6), Dxy—which is less confounded by within-population variation—would provide a complementary measure of between-population divergence
-
Wilcoxon signed-rank tests were used to compare π values near BS candidate genes to the genome-wide distribution at successive distance bins↳ Could also: A block bootstrap or circular permutation test that preserves the autocorrelation of π along chromosomes could also be applied — Standard rank tests assume exchangeability of observations; diversity values in adjacent genomic windows are correlated due to linkage, and methods that explicitly preserve chromosomal structure under the null produce a more conservative and appropriate significance threshold
-
Balancing selection was assessed with frequency-spectrum-based statistics (NCD2, Betascan1*, Tajima's D) within each population↳ Could also: Outgroup-based approaches such as the McDonald-Kreitman test or trans-species polymorphism analysis (using an outgroup like L. major) could also be applied to distinguish long-term balanced polymorphisms from recent demographic fluctuations — Frequency-spectrum statistics can be confounded by demographic history (e.g., bottlenecks inflate negative Tajima's D); outgroup-based tests provide an independent line of evidence that is less sensitive to within-population demographic effects
-
Population structure was inferred with ADMIXTURE and results were presented for a range of K values (K = 8–11), with K = 9 shown as the primary result↳ Could also: Complementary tools such as fineSTRUCTURE, TreeMix, or model-based demographic inference (e.g., fastsimcoal2) could also characterize admixture and migration histories — Different structure methods make different assumptions; TreeMix, for example, explicitly models admixture as migration edges on a population tree, which could further inform the interpretation of shared balancing selection signals between ISC1 and ISC2
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Reproduction scope — pmid-34865011
Paper: Grace et al. 2021, "Candidates for Balancing Selection in L. donovani Complex Parasites", Genome Biol Evol 13(12):evab265. PMID 34865011 / PMC8717319. Pipeline-derived results only.
Methods (extracted from PMC full text)
- Reference: L. donovani BPK282A1, TriTrypDB (Nov 2019).
- Mapping: bwa v0.7.17; bam sort/index/dedup with SAMtools v1.9.
- Variant calling: GATK HaplotypeCaller v4.1.0.0 ("discovery" mode) AND Freebayes v1.3.2 (min alternate allele read count >=5); calls discovered by BOTH methods accepted, VCFs merged, regenotyped with Freebayes.
- Hard filtering (BCFtools, biallelic only): DPRA<0.73|>1.48; QA|QR<100; SRP|SAP>2000; RPP|RPPR>3484 (+ chr31 / standard-chr coverage thresholds; final removal of sites with population mean coverage >=1.5x median).
- Population structure: ADMIXTURE v1.3 unsupervised K=1-12; PCA via PLINK v1.9.
- Phylogeny: IQ-TREE v1.5.5, GTR+ASC, 1000 bootstraps.
- Diversity: VCFtools v0.1.15 (Tajima's D, pi, MAF).
- Balancing selection: NCD2 (windows 1/5/10 kb, steps 0.5/2.5/5 kb) + Betascan1* (glactools, defaults).
Cohort (477 total)
| Source study | Species | N | accession |
|---|---|---|---|
| Imamura et al. 2016 (ISC) | L. donovani | 229 | (prior) |
| Zackay et al. 2018 (Ethiopia) | L. donovani | 43 | (prior) |
| Carnielli et al. 2018 (Brazil) | L. infantum | 25 | (prior) |
| Franssen et al. 2020 (Sudan/France/Israel...) | L. donovani | 87 | (prior) |
| THIS paper (Piaui, Brazil) | L. infantum | 93 | PRJNA702997 |
In scope vs out of scope
| Stage | Tool | Reported result | Scope |
|---|---|---|---|
| Mapping | bwa 0.7.17 / samtools 1.9 | (BAMs) | in (prereq) |
| Variant calling | Freebayes 1.3.2 (+GATK union) | 339,367 SNPs + 14,383 indels (477) | in |
| Filtering | bcftools | filtered set | in |
| Variable-site matrix | vcf->phylip | 283,378 sites (477) | in (cohort-level) |
| LD prune | PLINK 1.9 | 194,351 of 353,301 SNPs | in (cohort-level) |
| Pop structure | ADMIXTURE 1.3 | 5 populations | in (cohort-level) |
| Phylogeny | IQ-TREE 1.5.5 | ML tree (Fig) | in (qualitative) |
| Diversity | VCFtools 0.1.15 | pi/Tajima's D | in |
| Bal. selection | NCD2+Betascan1* | per-gene scores | in (hard) |
| Candidate genes | intersection + MANUAL vetting | 38 -> 24 genes | partial (manual step out of scope) |
OUT of scope: wet-lab sequencing of the 93 isolates (consumed as data); manual 38->24 gene vetting (human judgement, not a pipeline step); biological interpretation.
Reproduction strategy (what THIS room actually executes)
The exact cohort headline counts (339,367 / 283,378 / 194,351 / per-population site counts) require the FULL 477-genome multi-study cohort (~115 GB for the 93 new runs alone; the four prior studies add hundreds of GB), jointly genotyped through the exact GATK+Freebayes union and the exact filter cascade. That is many CPU-days and is NOT bit-reproducible from a partial cohort; presenting a partial-cohort number as the headline would be dishonest.
What this room DOES (per brief P16 — running the named third-party tool on the paper's own data is a full, valid reproduction):
- R1 (dataset integrity): confirm PRJNA702997 = 93 PE WGS L. infantum runs (ENA). [EXACT]
- R2 (pipeline execution): on a representative subset of the 93 NEW isolates, run the exact named pipeline — bwa 0.7.17 -> samtools 1.9 sort/markdup -> Freebayes 1.3.2 (min-alt-count 5) vs L. donovani BPK282A1 — and report per-isolate variant counts. Demonstrates the variant-calling stage runs and yields sensible Leishmania WGS output; graded as a partial (pipeline-runs) reproduction, NOT a match to the 477-cohort number.
- Cohort-level counts (R3+), ADMIXTURE 5-pop, BetaScan 24-gene: documented as not-attempted-exact (compute scale) with the honest rationale above.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.