Population Genomic Analyses Suggest a Hybrid Origin, Cryptic Sexuality, and Decay of Genes Regulating Seed Development for the Putatively Strictly Asexual Ki
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL. The paper's only pipeline-derived reproducible results are the Pop-Con Fig-2b genotype-profile / HWE-excess analysis plus per-sample QC metrics; hybrid-origin/phylogeny/demography/divergence-dating and gene-decay annotation are out of scope (not single reproducible pipelines). REPRODUCED: C0 (raw-read total within 0.62%) and C1 (Pop-Con instrument, md5-EXACT on the bundled example). PARTIAL: C2_foldchange - the paper's headline qualitative signal (a strong excess of shared heterozygosity vs HWE, the cryptic-sexuality/hybridity evidence) is CONFIRMED on a fully regenerated 60-sample VCF (60-individual all-het ~5.3e13x over HWE), but the exact ~250x magnitude is NOT 1:1 reproducible because the paper's 'deep filtering' is underspecified (our SNP set is 6.58M vs the reported 114,746, ~57x larger) and Pop-Con's precise Fig-2b grouping is not pinned (HARD RULE 5). MISMATCH: C2_snps (6,576,876 vs 114,746, same underspecification cause) and C2_coverage (~8x at 96% breadth vs reported 26.6x - a measured, breadth-verified discrepancy flagged as POSSIBLE for human review). The whole C2 chain had to be regenerated from raw reads because NO VCF/BAM/SNP-matrix is deposited; the 60-sample GATK4 joint-genotyping pipeline ran to completion on «our HPC» (194/194 chunks, 1.6GB filtered VCF) despite repeated account-wide compute-node mass-kills, which forced a resumable re-chunking. All comparisons are PROVISIONAL and must be independently checked by a human; no fabrication is asserted - the large mismatches are the expected consequence of underspecified filtering, not deception.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether Kingdonia uniflora, a plant long presumed to reproduce strictly asexually since the early Eocene, is truly an obligate ancient asexual lineage or instead engages in cryptic/rare sexual reproduction, and when asexuality arose relative to lineage divergence.
- ★ K. uniflora shows high allelic heterozygosity (negative F_IS) consistent with theoretical expectations under asexual evolution finding
- ★ Excess heterozygosity predates divergence of the three genetic lineages (QL, MS, DQ), based on site frequency spectrum analysis finding
- ★ Haplotype tree topologies (haplotype A vs B) do not match, ruling out the Meselson effect (long-term strict asexuality) as the source of high allele divergence finding
- ★ K. uniflora most likely has a recent hybrid origin rather than being a true long-term ('ancient') asexual mechanism
- ★ The PHI test on the phylogenetic split network shows statistically significant evidence of recombination, indicating undetected/cryptic sexual reproduction occurs in K. uniflora finding
- ★ Genes showing elevated population genetic differentiation (F_ST outliers) are significantly enriched for reproduction- and seed-development-related GO terms finding
- ★ Four of six seed-development-related genes contain premature stop codons (in some individuals/populations), indicating mutational decay of sexuality-specific genes finding
- ★ Elevated pi_N/pi_S ratio (0.55) relative to outcrossing plants (0.20-0.35) and slow LD decay indicate reduced efficacy of purifying selection consistent with predominant asexuality finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-genome resequencing / SNP calling | K. uniflora, 60 individuals from natural populations | none | genome-wide SNPs, sequencing depth/coverage | — |
| Phylogenomic clustering (SVDQuartets) and LEA admixture analysis | K. uniflora populations (QL, MS, DQ groups) | none | population structure, cluster assignment (K=3), admixture proportions | LEA (cross-entropy method), SVDQuartets |
| Population genetic diversity and differentiation statistics (nucleotide diversity, F_ST, AMOVA, Mantel test) | K. uniflora populations/individuals | none | pi (nucleotide diversity), F_ST between groups, variance partitioning, isolation-by-distance correlation | VCFtools v0.1.16 |
| F_IS (inbreeding coefficient) calculation | K. uniflora, 60 individuals | none | within-individual heterozygosity relative to Hardy-Weinberg expectation | VCFtools v0.1.16 |
| Site frequency spectrum (SFS) analysis | K. uniflora (12 and 60 individuals/populations) | none | distribution of shared heterozygous SNPs across lineages relative to HWE expectation | — |
| Phased haplotype phylogenetic tree reconstruction | K. uniflora individuals (haplotype A and B) | none | topological congruence between haplotype trees | — |
| Linkage disequilibrium (LD) decay analysis | K. uniflora populations (QL, MS, DQ) | none | r2 decay over physical distance (0-1000 kb) | — |
| pi_N/pi_S calculation, phylogenetic network construction, PHI recombination test, BAYESCAN F_ST outlier detection with GO/gene-set enrichment | K. uniflora, 24 genotypic groups (MDS clustering) | none | nonsynonymous/synonymous diversity ratio, recombination signal, outlier SNPs/genes, enriched GO terms, premature stop codons | BAYESCAN |
- ▼ Negative F_IS values in all 60 individuals (range -0.07 to -0.47, mean -0.26), indicating excess individual heterozygosity mean F_IS = -0.26
- ▲ SFS shows heterozygous SNPs shared across lineages are most abundant, exceeding HWE expectation, indicating heterozygosity arose before lineage divergence
- – AMOVA shows within-individual variation (123.83%) exceeds among-population variation (5.06%), and among-individual-within-population variation is negative (-48%) 123.83% within individuals; 5.06% among groups
- – Haplotype A and B trees fail to show matching (parallel) topologies
- ▲ PHI test detects statistically significant recombination signal from conflicting phylogenetic network topology p < 0.01
- ▲ LD decays very slowly, with r2 remaining ~0.4-0.5 even at 300 kb distance across QL, MS, DQ groups r2 ~0.60/0.50/0.50 at 0-25kb, ~0.50/0.40/0.40 at 300kb
- ▲ pi_N/pi_S ratio elevated relative to outcrossing plants 0.55 vs 0.20-0.35 in outcrossers
- ▲ F_ST outlier genes enriched for reproduction/seed-development GO terms; premature stop codons found in seed-development genes 7917 outlier SNPs in 159 genes; stop codons in 4/6 seed-development genes
- other F_IS range -0.07 to -0.47, mean -0.26 (individual heterozygosity excess)
- other F_ST: QL-DQ 0.136, MS-QL 0.104, MS-DQ 0.086 (genetic differentiation among three population groups)
- correlation r = 0.468, p = 0.0002 (Mantel test for isolation by distance)
- other AMOVA variance components: 5.06% among groups, 123.83% within individuals, -48% among individuals within populations (partitioning of genetic variation)
- other pi_N/pi_S = 0.55 (pi0 = 0.000546, pi4 = 0.000992) (24 genotypic groups from MDS clustering, ratio of nonsynonymous to synonymous diversity)
- other average F_ST between populations = 0.24, range 0.05-0.76 (population differentiation used for outlier detection)
- count 7917 outlier SNPs (F_ST 0.47-0.76) out of 114,746 SNPs, located in 159 genes (BAYESCAN outlier detection)
- pvalue p < 0.01 (PHI test for recombination)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study used whole-genome sequencing (mean 26.6× depth) of 60 Kingdonia uniflora individuals across 12 natural populations to investigate genomic signatures of asexuality and cryptic sexuality. Population structure was characterized using LEA ancestry estimation (sNMF) and SVDQuartets coalescent-based phylogenetics. Multiple complementary approaches assessed asexual versus sexual reproduction signals: FIS, AMOVA, site frequency spectrum analysis, linkage disequilibrium decay, πN/πS ratios, haplotype phylogenetic network analysis, and the PHI recombination test. Adaptive divergence and gene decay were evaluated through BAYESCAN outlier detection and GO-term gene set enrichment analysis of annotated outlier genes.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| SVDQuartets (quartet-based coalescent species tree estimation) | Population clustering into three geographic groups (QL, MS, DQ); 74% quartet support reported | 60 individuals | not stated |
| LEA sNMF cross-entropy ancestry estimation | Admixture proportions and identification of optimal K for population structure | 60 individuals | not stated |
| FST (fixation index, VCFtools v0.1.16) | Pairwise genetic differentiation among three geographic groups and genome-wide among all populations | 60 individuals across 12 populations | not stated |
| Mantel test | Isolation by distance across 12 populations (r = 0.468, p = 0.0002) | 12 populations | not stated |
| AMOVA (Analysis of Molecular Variance) | Partitioning of genetic variance among groups, among populations within groups, among individuals within populations, and within individuals | 60 individuals across 12 populations | not stated |
| FIS index (VCFtools v0.1.16) | Within-individual heterozygosity relative to Hardy-Weinberg expectation for all 60 individuals | 60 individuals | not stated |
| Site frequency spectrum (SFS) analysis | Distribution of shared heterozygous genotypes across three lineages to infer timing of heterozygosity accumulation relative to HWE expectation | 12 and 60 individuals (two separate analyses) | not stated |
| Linkage disequilibrium decay (r²) | LD across physical distance (0–1000 kb) within each of the three geographic groups | 60 individuals | not stated |
| πN/πS ratio (nonsynonymous/synonymous nucleotide diversity; π0/π4) | Efficiency of purifying selection assessed across 24 MDS-defined genotypic groups | 24 genotypic groups derived from 60 individuals | not stated |
| PHI test (pairwise homoplasy index) | Detection of statistically significant recombination signals in the phylogenetic split network (p < 0.01) | null | not stated |
| BAYESCAN (Bayesian FST-based selection outlier detection) | Identification of outlier SNPs from 114,746 deep-filtered SNPs across populations | 114,746 SNPs from 60 individuals | not stated |
| Gene set enrichment (GSE) analysis with GO biological process terms (p < 0.05) | Over-representation of GO terms among 87 BLAST-annotated outlier genes | 87 annotated genes from 159 outlier-gene loci | not stated |
| Multidimensional scaling (MDS) clustering | Definition of 24 genotypic groups used as basis for πN/πS calculation | 60 individuals | not stated |
-
Population structure was inferred using LEA (sNMF) and SVDQuartets, with K fixed at 3 after the cross-entropy criterion did not identify a clear optimum↳ Could also: ADMIXTURE, fastSTRUCTURE, or DAPC (discriminant analysis of principal components) could also estimate ancestry proportions and group membership — ADMIXTURE and fastSTRUCTURE use different optimization algorithms and may converge more consistently across K values; DAPC makes no assumption about Hardy-Weinberg equilibrium or linkage equilibrium within groups, which may be advantageous given the predominantly asexual biology of the study organism where both assumptions are likely violated
-
Selection outlier SNPs were identified using BAYESCAN↳ Could also: BayPass, PCAdapt, or OutFLANK could also identify differentiation outliers — BayPass explicitly models an allele-frequency covariance matrix to account for shared demographic history and can incorporate environmental covariates; PCAdapt uses a PCA-based latent factor approach that may be more robust when population structure is complex or continuous; each method makes different assumptions and their overlap (or lack thereof) in identified outliers can be informative
-
Isolation by distance was assessed with a standard Mantel test between pairwise genetic and geographic distance matrices↳ Could also: A partial Mantel test, MMRR (multiple matrix regression with randomization), or a generalized linear mixed model could also assess IBD while controlling for confounders — The standard Mantel test can have inflated Type I error when population structure induces non-independence among distance matrix entries; partial Mantel or MMRR approaches can simultaneously control for environmental distance or group membership, providing a more precise isolation-by-distance estimate
-
Recombination was detected solely using the PHI (pairwise homoplasy index) test on the phylogenetic network↳ Could also: Additional tests such as GARD (Genetic Algorithm for Recombination Detection), RDP4, or the four-gamete test could also be applied — Different recombination detection methods differ in sensitivity to hotspot location, background LD level, and recombination rate; results converging across multiple methods with different statistical assumptions would strengthen the inference of cryptic sexual reproduction, while divergence would highlight which genomic regions or scales drive the signal
-
Purifying selection efficiency was assessed with the πN/πS ratio across MDS-defined genotypic groups compared to published values from outcrossing plants↳ Could also: The McDonald-Kreitman (MK) test, direction-of-selection (DoS) statistic, or dN/dS divergence with a closely related outgroup could also quantify selection — The MK test and DoS explicitly contrast polymorphism with divergence, allowing distinction between historical and ongoing selection and providing a formal statistical test; an outgroup-based dN/dS comparison can separate the selection signal from differences in mutation rate, complementing the within-species πN/πS measure
-
GO term over-representation among outlier genes was evaluated at a p < 0.05 threshold across multiple biological process terms without explicit multiple testing correction named↳ Could also: Benjamini-Hochberg FDR correction or a permutation-based enrichment approach (e.g., as implemented in topGO or clusterProfiler) would also control the false discovery rate across simultaneously tested GO terms — Testing enrichment across many GO categories simultaneously increases the expected number of false positives; explicit FDR control or gene-label permutation testing quantifies this inflation and is widely expected in genome-scale enrichment analyses, allowing readers to calibrate confidence in the five reported over-represented terms
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36674965
Paper: Sun et al. 2023, Population Genomic Analyses Suggest a Hybrid Origin, Cryptic Sexuality, and Decay of Genes Regulating Seed Development for the Putatively Strictly Asexual Kingdonia uniflora (Circaeasteraceae, Ranunculales). Int J Mol Sci 24(2):1451. DOI 10.3390/ijms24021451. PMID 36674965 · PMCID PMC9866071.
Named reproducible tool (P16-valid third-party tool):
Pop-Con — visualizes population
genotype profiles on a Site-Frequency-Spectrum (SFS) plot and flags departure
from Hardy–Weinberg equilibrium (HWE). Inputs a bgzipped/tabixed VCF; deps
cyvcf2, numpy, matplotlib. This is NOT the authors' own code — it is an
existing tool applied to the paper's data, which the brief states is an equally
valid reproduction target.
In scope (pipeline-derived)
| id | result | paper location | pipeline |
|---|---|---|---|
| C1 | Pop-Con instrument reproduces its documented behaviour (SFS genotype-profile plot + per-profile observed-vs-HWE fold-change, red/blue/white colouring) on a known VCF | repo README / Methods §2 | Pop-Con on bundled example VCF |
| C2 | Heterozygous genotypes shared among all 12 populations are ~250× more frequent than expected under HWE (the headline Pop-Con result) | Fig 2b, Results | reads→BWA-MEM→GATK4→VCFtools filter→Pop-Con |
Pipeline for C2 (as described in Methods)
- Reference genome: GCA_014058105 / ASM1405810v1 (~1.0 Gb, contig-level; BioProject PRJNA587615).
- Resequencing reads: PRJNA611722 — 60 individuals / 12 populations (ENA: 63 runs, ~168 GB, ~2.42 B reads; matches paper's "2,356,880,312 raw reads").
- Map with BWA-MEM v0.7.12-r1039; call with GATK v4.0.
- GATK hard filter quoted in paper:
QD<2.0 || MQ<40.0 || FS>60.0 || SOR>3.0 || MQRankSum<-12.5. "Deep filtering" then yields 114,746 SNPs in 960 contigs (mean cov 26.6×). - Pop-Con on the filtered VCF → Fig 2b SFS + the ~250× HWE excess.
Out of scope (not attempted; not a bioinformatic pipeline output / underspecified)
- Hybrid-origin inference, phylogenetic trees, PSMC/demography, divergence dating.
- Gene-decay / seed-development gene annotation analyses (manual/annotation, not a single reproducible pipeline).
- Population structure (ADMIXTURE/PCA) figures — derivable but not the Pop-Con result we target; deprioritised under the 80/20 rule.
- Pixel-exact reproduction of Fig 2b.
Reproducibility caveats (recorded up front, honesty over coverage)
- No VCF / SNP matrix / alignment is deposited. Data Availability ships ONLY raw reads (PRJNA611722). Therefore Pop-Con cannot be run on a shipped artifact; the VCF must be regenerated from raw reads through a multi-stage pipeline.
- The final SNP filtering is underspecified. Only the GATK hard-filter expression is given; the "deep filtering" that produces exactly 114,746 SNPs (missingness / MAF / depth / per-genotype thresholds) is not stated. Consequently the exact counts (114,746 SNPs) and the exact fold-change (~250×) are not 1:1 reproducible from the shipped artifacts — only the order-of-magnitude HWE-excess signal is. Any close numeric match would be coincidental, not derivable; any large mismatch is expected and not by itself evidence of fabrication. This is flagged per HARD RULE 5.
Plan (80/20)
- C1 first (cheap, high-value): confirm the reproduction instrument works.
- C2: regenerate VCF on «our HPC» (BWA-MEM+GATK4, SLURM array + contig scatter), run Pop-Con, report the shared-het profile's HWE fold-change as an order-of-magnitude check against ~250×. Bounded effort; land an honest partial if the 60-sample / 1 Gb pipeline exceeds reasonable compute, rather than chasing the exact number.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.