Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Natural clines and human management impact the genetic structure of Algerian honey bee populations.

Genet Sel Evol · 2023
L1 55/100 PQI 87
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🔴The deviation was non-trivial in magnitude
  • 🔴The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
How its reproducibility compares
55/100
Reproducibility score
1.1 SD below mean
vs. all fields · 1187 studies
🎯 Scores higher than 15% of all assessed papers rank 994 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL. The paper (Algerian honey bee population genomics) is described well at the methods level, but its computational results are NOT reproducible from deposited artifacts: the authors' own analysis pipeline (BWA-MEM->GATK->ADMIXTURE/SMARTPCA/hmmIBD/Random-Forest/cline scripts) is not deposited anywhere (no authors' GitHub/Zenodo found), and only raw FASTQ is in SRA PRJNA1044268 (no genotype matrix/VCF). The ONLY deposited runnable code is the cited third-party R package gscramble (eriqande); per P16 we reproduced it faithfully: pinned to v1.0.1 (commit 7018bf7) on «our HPC», it builds, installs and runs a deterministic gene-drop (set.seed(15) -> 78x200 simulated genotype matrix, digest d31fd8b7). NOT attempted: the full raw-reads->SNP pipeline (out of 80/20, TB-scale, needs an external 228-sample reference panel not in this accession) and all cohort-level numbers (C1-C9) which depend on the undeposited VCF + undeposited scripts. The paper's gscramble-based hybrid-misclassification numbers (C10) are not regenerable because the gscramble inputs (real founder genotypes, Amel_HAv3.1 recomb map, 96-SNP panel, trained RF) are not deposited. Not a clean drop (cited code runs); not a full reproduction (no paper number regenerable from shipped artifacts).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 55
    assessed: 2026-06-15 ⛓ b6c86286075c
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether the two Algerian honey bee subspecies (A. m. intermissa and A. m. sahariensis) show distinct population genetic structure, whether admixture with European honey bee subspecies has occurred, and whether a reduced SNP panel can reliably distinguish Algerian from European honey bees for conservation purposes.

Core claims
  • Algerian honey bees show no significant admixture from European reference honey bee populations finding
  • Genetic variation in Algerian honey bees is structured mainly along an East-West geographic axis rather than by the A. m. intermissa/A. m. sahariensis subspecies split finding
  • Correlation between genetic and geographic distance is higher in the Western genetic cluster than the Eastern one finding
  • Close family relationships are mostly detected in the Eastern cluster, sometimes over long distances finding
  • Differences between the two main genetic clusters reflect differential breeding management practices between eastern and western Algeria mechanism
  • A panel of 96 ancestry-informative SNP markers, selected via random forest classification, effectively distinguishes Algerian from European honey bees, including in simulated admixed/backcross individuals resource
  • Combining high-throughput sequencing with machine learning for ancestry-informative marker selection is a promising general approach for taxa with weak genetic differentiation method
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome sequencing (WGS) 151 haploid drones (108 A. m. intermissa, 43 A. m. sahariensis), Algeria none genome-wide SNP genotypes Illumina NovaSeq 6000, TruSeq Nano DNA HT Library Prep
Variant calling / genotyping 379 combined samples (151 Algerian + 228 European reference drones) none SNP calls and quality-filtered genotypes BWA-MEM, GATK HaplotypeCaller/GenotypeGVCFs
Principal component analysis (PCA) LD-pruned SNP dataset from Algerian and reference drone genomes none principal components of genetic variation SMARTPCA (EIGENSOFT v7.2.1)
ADMIXTURE analysis Algerian and reference drone genomes, K=2 to 12 none ancestry proportions per sample ADMIXTURE v1.3.0, pong
Identity-by-descent (IBD) kinship analysis pairwise Algerian drone samples none pairwise IBD kinship coefficients hmmIBD
Genetic-geographic distance correlation 102 filtered Algerian samples none Pearson correlation between PC-based/IBD genetic distances and Haversine geographic distances MATLAB
Random forest machine learning classification for AIM SNP selection 99,274 candidate SNPs from Algerian and M/C lineage reference samples none SNP feature importance, classification accuracy of Algerian vs non-Algerian labels scikit-learn RandomForestClassifier
Simulated hybrid/backcross genotype validation 640 simulated F1 and backcross (BC1-BC3) genotypes between Algerian and M/C lineage bees simulated admixture (in silico crosses) RF classification probability of Algerian ancestry gscramble R package
Key results
  • No significant admixture detected between Algerian honey bees and European reference populations
  • Two main genetic clusters found along an East-West axis, not corresponding to A. m. intermissa/A. m. sahariensis subspecies boundaries
  • Genetic-geographic distance correlation higher in the Western cluster than Eastern cluster
  • Close-family relationships mostly detected in the Eastern cluster, sometimes at long distances
  • 96-SNP AIM panel maximized separation between Algerian samples and European (M and C) lineages in PCA validation, outperforming random 96-SNP panels
  • 16 putative full-sib families identified among initial 151 samples, requiring removal of related individuals before downstream analyses 64 of 151 samples discarded as relatives
  • Final filtered dataset retained 102 unrelated Algerian samples (86 A. m. intermissa, 16 A. m. sahariensis) from 48 locations
Key statistics
  • count 17,817,395 raw variants (joint genotyping of Algerian and reference samples before filtering)
  • count 12,302,217 SNPs (SNPs retained after removing indels)
  • count 8,865,912 SNPs (final SNP dataset after quality control filtering)
  • count 1,374,698 SNPs (SNPs retained after LD pruning of Algerian-only dataset)
  • other IBD kinship threshold set at 30% (used to identify full-sib families (~50% expected for first-degree relatives))
  • count 96 SNPs (size of selected ancestry-informative marker (AIM) panel)
  • count 640 simulated hybrid genotypes (160 F1, 160 BC1, 160 BC2, 160 BC3) (simulated admixture individuals used to validate AIM panel)
  • count 102 final samples (86 A. m. intermissa, 16 A. m. sahariensis) (samples retained after relatedness and call-rate filtering)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used whole-genome sequencing of 151 haploid Algerian drone bees to characterise population structure through PCA (SMARTPCA/EIGENSOFT), unsupervised ADMIXTURE clustering (K=2–12, 50 runs per K, CV-error K selection), and pairwise IBD kinship estimation via a hidden Markov model (hmmIBD). Isolation-by-distance was assessed with Pearson correlations between PC-derived genetic distances (or IBD kinships) and Haversine geographic distances. A random forest classifier (scikit-learn) was used to select a 96-SNP ancestry-informative panel distinguishing Algerian from European lineages, validated against 1,000 random panels and tested on simulated hybrid/backcross genotypes generated with the gscramble R package.

Replicationbiological Sample size151 drone bees from 20 Algerian regions across two field campaigns (2017–2018 and 2021–2022), reduced to 102 after IBD-based family filtering and call-rate filtering; reference panel N=228 from a prior published study (M lineage N=63, C lineage N=148, O lineage N=17) GroupsA.m. intermissa vs A.m. sahariensis; Eastern vs Western Algerian genetic clusters; Algerian vs European lineages (M: mellifera/iberiensis; C: ligustica/carnica/Royal Jelly; O: caucasia) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Principal Component Analysis (SMARTPCA/EIGENSOFT v7.2.1) Population structure visualisation — full dataset (Algerian + reference) and Algerian-only; also used to validate AIM panel separation 330 (whole dataset after filtering) or 102 (Algerian-only) not stated
ADMIXTURE unsupervised maximum-likelihood clustering (v1.3.0), K selected by lowest mean CV error across 50 runs Admixture analysis with reference populations and within Algerian samples, K=2–12 330 (with reference) or 102 (Algerian-only) not stated
hmmIBD hidden Markov model — pairwise IBD kinship estimation Full 151 samples for first-degree relatedness screening; 102 retained samples for isolation-by-distance analysis; two genetic clusters separately for close-kin comparison (kinship ≥10% threshold for 3rd-degree relatives) 151 initially; 102 after filtering not stated
Pearson correlation Top 7 PCs vs. three spatial coordinates (21 correlations); pairwise PC-derived Euclidean genetic distances vs. pairwise Haversine geographic distances; pairwise IBD kinships vs. pairwise geographic distances 102 (Algerian samples) not stated
Random Forest classifier (scikit-learn RandomForestClassifier, 100 independent runs per configuration) Selection and validation of 96 AIMs distinguishing Algerian from M-lineage and C-lineage European bees; performance benchmarked against 1,000 random 96-SNP panels (each run 50 times) Training N=250 (82 Algerian + 50 M-lineage or 118 C-lineage); testing N=63; extended to N=570/383 with simulated hybrids not stated
Approaches that could also have been used
  • Pearson correlation was used to assess the association between pairwise genetic and geographic distance matrices
    Could also: A Mantel test (or partial Mantel test controlling for a third matrix, e.g. shared subspecies) could also be used to test isolation-by-distance — Pairwise distance matrices share observations across rows and columns, violating the independence assumption of standard Pearson correlation; the Mantel test uses permutation to generate a null distribution appropriate for matrix correlations, and the partial Mantel extension allows geographic effects to be separated from subspecies membership effects
  • The number of genetic clusters K was selected using ADMIXTURE's cross-validation (CV) error across 50 runs per K
    Could also: The Evanno ΔK method, BIC from sparse non-negative matrix factorization (sNMF in the LEA R package), or STRUCTURE HARVESTER could also be used to guide K selection — CV error identifies K that minimises predictive error on withheld genotypes; ΔK captures the rate-of-change in the log-likelihood and can highlight hierarchical structure at a different level; sNMF provides a regularised likelihood criterion and runs efficiently on large SNP datasets — comparing multiple criteria can triangulate a stable K
  • A random forest classifier was used to rank and select 96 ancestry-informative SNPs based on feature importance (Gini impurity or entropy)
    Could also: FST-based ranking, informativeness-for-assignment (In) scores, or penalised logistic regression (LASSO) could also be used to identify ancestry-informative markers — FST ranking is a classical, transparent one-locus approach for AIM selection; In scores are information-theoretic measures directly interpretable as assignment power per locus; LASSO provides built-in sparsity with a regularisation path tuned by cross-validated accuracy — these methods offer complementary perspectives to tree-based importance scores and can serve as benchmarks
  • IBD kinship was estimated with hmmIBD, a tool originally developed for haploid P. falciparum genomes and adapted here by changing chromosome number and recombination rate parameters
    Could also: KING-robust or a method-of-moments estimator (PLINK2 --make-king) applied to the diploid-model genotype calls could also estimate pairwise relatedness in structured populations — hmmIBD's HMM framework explicitly models recombination breakpoints and segment lengths, which is well-suited to haploid data; KING is designed to be robust to allele-frequency differences between subpopulations — comparing estimates from both approaches would help assess the sensitivity of close-kin calls to the choice of estimator under the observed population structure
  • Genetic distance between pairs of samples was operationalised as a weighted Euclidean distance on the top two PCs (weighted by the eigenvalue ratio of PC1/PC2)
    Could also: Proportion of allele-sharing differences (1 − IBS) computed directly from all SNPs, or Reynolds' co-ancestry distance from allele frequencies, could also serve as pairwise genetic distance inputs — PC-based distance compresses genome-wide information efficiently but depends on the number of PCs retained and the chosen weighting; allele-sharing distances use the full SNP complement without dimensionality reduction and are a standard input for Mantel-type isolation-by-distance analyses, providing a complementary summary of genome-wide differentiation
  • Admixture proportions from ADMIXTURE were used to characterise and describe the absence of large-scale introgression from European lineages into Algerian bees
    Could also: D-statistics (ABBA-BABA tests) or f3/f4-statistics from the ADMIXTOOLS package could also formally test for gene flow between specific population pairs — ADMIXTURE proportions reflect the proportion of ancestry attributable to each cluster but do not assign a formal test statistic or p-value to introgression; D- and f-statistics use the site-frequency spectrum of four-taxon configurations to detect asymmetric gene flow and provide Z-scores against a permutation null, offering a complementary and more formal inference of admixture direction and magnitude
Software: BWA-MEM 0.7.15 · GATK (HaplotypeCaller, CombineGVCFs, GenotypeGVCFs, SelectVariants) 4.1.2 · PICARD MarkDuplicates · PLINK (indep-pairwise LD pruning) · SMARTPCA / EIGENSOFT 7.2.1 · ADMIXTURE 1.3.0 · pong (run alignment and clustering) · hmmIBD · scikit-learn / Python (RandomForestClassifier) · gscramble (R package, for simulated hybrids) · MATLAB (geographic maps)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
3
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

PRJNA1044268 BioProject in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet

Downstream reach in the literature

0 downstream papers · 1 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

PRJNA1044268 BioProject reused by 1 papers in the literature

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-38114899

Paper: Salvatore et al. 2023, Natural clines and human management impact the genetic structure of Algerian honey bee populations. Genet Sel Evol. DOI 10.1186/s12711-023-00864-5 · PMCID PMC10729559.

Data: SRA BioProject PRJNA1044268 — raw WGS reads (FASTQ) of 151 haploid drones (Illumina NovaSeq 6000, 2×150 bp). Public and resolvable, but raw reads only — no intermediate genotype matrix / VCF is deposited.

Code (per the paper's "Availability"): only gscramble (https://github.com/eriqande/gscramble) is cited — a generic third-party R package (eriqande, CRAN v1.0.1) used for one validation step (simulating F1/backcross genotypes). The authors' own analysis pipeline (alignment, variant-calling, ADMIXTURE/PCA/IBD wrappers, Random-Forest AIM selection, gscramble driver, cline analysis) is NOT deposited anywhere — no authors' GitHub/Zenodo analysis repo was found (only SRA raw reads + a Figshare sample table, Additional file 2).

Pipeline map per reported result

Reported result Pipeline / tool In scope? Why
8,865,912 SNPs (post-QC); 1,305,303 / 1,374,698 LD-pruned BWA-MEM 0.7.15 → GATK 4.1.2 HaplotypeCaller + joint genotyping → PLINK LD-prune OUT (80/20) Requires aligning + joint-calling 151 WGS samples (≈TB-scale, many CPU-days) and the wrapper scripts/parameters are not deposited; intermediate VCF not deposited. Heaviest + least-specified part.
ADMIXTURE K=4 (ref+Alg), K=2 (Alg only); 50 runs/K ADMIXTURE 1.3.0 + pong OUT Depends on the (undeposited) SNP VCF and an external reference panel of 228 M/C/O samples that is not part of PRJNA1044268.
PCA PC1 28.23 %, PC2 16.07 % SMARTPCA (EIGENSOFT 7.2.1) OUT Depends on the undeposited VCF + external reference panel.
Geo-genetic r (45 % / 71.6 % / 25.9 %); IBD kinship r (−25.5 % / −63.1 % / −21.3 %); 33/5151 close pairs hmmIBD + custom correlation/cline scripts OUT Depends on the undeposited VCF; correlation/cline scripts not deposited.
96-SNP AIM panel; 100 % test accuracy on 63 samples Random-Forest (scikit-learn) OUT Selection script not deposited; depends on the undeposited VCF.
Hybrid misclassification BC2M 1.4 %, BC3M 7.25 % gscramble (F1/BC gene-drop) + RF classifier PARTIAL (tool only) gscramble itself is deposited & runnable. But reproducing the paper's number needs the real founder genotypes, an Amel_HAv3.1 recombination map, the 96-SNP panel, the exact pedigree design and the trained RF — none of these gscramble inputs are deposited.

What we actually attempt (the auditable 80/20 floor)

The only deposited, executable code artifact is gscramble. We therefore:

  1. Install gscramble pinned to v1.0.1 (commit 7018bf787f741ec361b9f526a107a6d9fd2d5b06) in a conda R env on «our HPC».
  2. Run the package's own R CMD check / testthat suite (its baked-in expected values) → deterministic confirmation the cited code builds and behaves as published.
  3. Compute a fingerprint of the shipped example data + run a set.seed() gene-drop example → a concrete, re-runnable reproduced value.

This reproduces the cited tool, not the paper's bee-specific gscramble numbers (inputs not deposited). All paper headline numbers above are graded partial/not-reproducible-from-shipped-artifacts, with the reason recorded.

Honest verdict direction: PARTIAL. The cited code runs and is deterministic; none of the paper's reported numbers are regenerable from the deposited artifacts (authors' pipeline + intermediate genotypes not deposited; raw→SNP step out of 80/20 and needs an external reference panel).

C1
Reported
8,865,912 SNPs post-QC
Reproduced
not attempted
partial
C2
Reported
1,305,303 / 1,374,698 LD-pruned SNPs
Reproduced
not attempted
partial
C3
Reported
ADMIXTURE K=4 / K=2
Reproduced
not attempted
partial
C4
Reported
PCA PC1 28.23% / PC2 16.07%
Reproduced
not attempted
partial
C5
Reported
geo-genetic r 45 / 71.6 / 25.9%
Reproduced
not attempted
partial
C6
Reported
IBD kinship r -25.5 / -63.1 / -21.3%
Reproduced
not attempted
partial
C7
Reported
33/5151 close family pairs (31 Eastern)
Reproduced
not attempted
partial
C8
Reported
96-SNP AIM panel
Reproduced
not attempted
partial
C9
Reported
100% accuracy on 63 test samples
Reproduced
not attempted
partial
C10
Reported
hybrid misclassification BC2M 1.4% / BC3M 7.25%
Reproduced
gscramble tool reproduces deterministically; paper's bee-specific number not regenerable (founder genotypes + recomb map + 96-SNP panel + RF not deposited)
partial
C11
Reported
gscramble cited code builds & runs deterministically
Reproduced
v1.0.1 @ 7018bf7 builds+installs (R 4.3.3); all functional R CMD check stages OK (only cosmetic vignette image-include ERROR; no testthat suite ships); set.seed(15) segregate()=36-row tbl (digest cedbcccf) -> segments2markers()=78x200 geno matrix (digest d31fd8b7)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 55/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🔴6. Severity of the deviation
🔴7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

This study is not reproducible from deposited artifacts: only raw FASTQ sits in SRA (no VCF, no reference panel) and the authors' entire analysis pipeline is deposited nowhere, so claims C1–C9 (SNP counts, ADMIXTURE K, PCA variance, IBD/geo correlations, the 96-SNP AIM panel and its 100%/63 accuracy) are all unverifiable — non-derivable from shared data+code rather than demonstrably wrong. The failure is squarely on the authors' deposit side (missing scripts + intermediate genotypes), not a methodology choice or a legitimate data restriction. The lone deterministic success is the cited third-party package gscramble (v1.0.1 @7018bf7, gene-drop digest d31fd8b7), a tool-level meta-claim that reproduces no paper number; C9's 100% accuracy is additionally flagged as a too-perfect, uncheckable figure.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

168.9 k
tokens (I/O) · 12.5 M incl. cache
35 min
runtime · 0.04 CPU-h
2.5 GB
peak RAM
1
HPC jobs
hummel
machine