Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Natural clines and human management impact the genetic structure of Algerian honey bee populations.

Genet Sel Evol · 2023
L1 55/100 PQI 87
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🔴The deviation was non-trivial in magnitude
  • 🔴The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
How its reproducibility compares
55/100
Reproducibility score
1.1 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 15% of all assessed papers rank 986 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL. The paper (Algerian honey bee population genomics) is described well at the methods level, but its computational results are NOT reproducible from deposited artifacts: the authors' own analysis pipeline (BWA-MEM->GATK->ADMIXTURE/SMARTPCA/hmmIBD/Random-Forest/cline scripts) is not deposited anywhere (no authors' GitHub/Zenodo found), and only raw FASTQ is in SRA PRJNA1044268 (no genotype matrix/VCF). The ONLY deposited runnable code is the cited third-party R package gscramble (eriqande); per P16 we reproduced it faithfully: pinned to v1.0.1 (commit 7018bf7) on «our HPC», it builds, installs and runs a deterministic gene-drop (set.seed(15) -> 78x200 simulated genotype matrix, digest d31fd8b7). NOT attempted: the full raw-reads->SNP pipeline (out of 80/20, TB-scale, needs an external 228-sample reference panel not in this accession) and all cohort-level numbers (C1-C9) which depend on the undeposited VCF + undeposited scripts. The paper's gscramble-based hybrid-misclassification numbers (C10) are not regenerable because the gscramble inputs (real founder genotypes, Amel_HAv3.1 recomb map, 96-SNP panel, trained RF) are not deposited. Not a clean drop (cited code runs); not a full reproduction (no paper number regenerable from shipped artifacts).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 55
    assessed: 2026-06-15 ⛓ b6c86286075c
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests whether the two Algerian honey bee subspecies (A. m. intermissa and A. m. sahariensis) are genetically differentiated and whether Algerian honey bee populations have been admixed by imported European subspecies, while also evaluating whether a reduced SNP panel can distinguish Algerian from European bees.

Core claims
  • No significant admixture from European subspecies was detected in Algerian honey bees, suggesting large-scale queen imports have not occurred in Algeria. finding
  • Most genetic variation lies not between the intermissa and sahariensis subspecies but along an East–West axis forming two main genetic clusters. finding
  • Correlation between genetic and geographic distances was higher in the Western cluster, while close-family relationships were mostly detected in the Eastern cluster, sometimes over long distances, indicating differential breeding management. finding
  • A panel of 96 ancestry-informative SNPs selected via a random forest classifier effectively distinguishes Algerian from European (M and C lineage) honey bees. resource
  • Whole-genome sequencing of 151 haploid Algerian drones extends the available haploid genome dataset. resource
  • Combining high-throughput sequencing and machine learning for ancestry-informative marker selection is promising for genetic characterization across taxa with difficult-to-capture differentiation. method
  • The 96-SNP panel was validated in simulated F1 and reciprocal backcross admixture scenarios using gscramble. method
Experimental setups
Assay System Perturbation Readout Platform
whole-genome DNA sequencing Apis mellifera haploid drones (108 A. m. intermissa, 43 A. m. sahariensis) from Algeria none genome-wide SNP genotypes Illumina NovaSeq 6000 S4, TruSeq Nano DNA HT Library Prep Kit, 2 × 150 bp paired-end
read mapping and variant calling 151 Algerian drones + 228 European reference samples (379 total) none called SNPs after QC BWA-MEM v0.7.15, PICARD, GATK v4.1.2 HaplotypeCaller
identity-by-descent kinship analysis Algerian honey bee drone samples none pairwise IBD kinship / IBD segments hmmIBD (nchrom=16, rec_rate=9.04e-7)
principal component analysis LD-pruned Algerian + reference samples none principal components of genetic variation SMARTPCA, EIGENSOFT v7.2.1
admixture analysis reference and Algerian samples, K=2 to 12 none ancestral proportions / cluster assignment ADMIXTURE v1.3.0; pong
genetic-geographic distance correlation 102 filtered Algerian samples none Pearson correlation between genetic and geographic (Haversine) distances MATLAB
random forest ancestry-informative marker selection 82 Algerian vs 50 M-lineage or 118 C-lineage samples none 96-SNP panel and classification accuracy scikit-learn RandomForestClassifier (gini/entropy)
simulated hybrid/backcross validation simulated F1, BC1, BC2, BC3 hybrids between Algerian and M/C lineages in silico simulated admixture/backcrossing class assignment probability with 96-AIM panel gscramble R package
Key results
  • No significant admixture detected between Algerian honey bees and European reference populations
  • Two main genetic clusters found along an East–West axis rather than between the two subspecies
  • Correlation between genetic and geographic distances higher in the Western cluster
  • Close-family (full-sib/related) pairs mostly detected in the Eastern cluster, sometimes at long distances
  • 16 putative families of full sibs identified (2 to 26 samples per family; 64 samples involved) leading to removal of 48 samples
  • 96-SNP panel selected and validated to distinguish Algerian from European bees
Key statistics
  • count 151 haploid drones sequenced (108 A. m. intermissa, 43 A. m. sahariensis) (Algerian samples collected from 20 regions, 2017–2022)
  • count 17,817,395 raw variants (joint genotyping of 379 samples (228 reference + 151 Algerian))
  • count 12,302,217 SNPs after removing indels (SNP QC starting set)
  • count 8,865,912 SNPs in final dataset (after quality control filters)
  • count 102 samples (86 A. m. intermissa, 16 A. m. sahariensis) from 48 sampling locations (final samples after family and call-rate filtering)
  • count 1,374,698 SNPs after LD pruning (Algerian only); 1,305,303 SNPs (whole dataset) (LD-pruned datasets for analyses)
  • count 99,274 SNPs fed to RF as AIM candidates; 96-SNP final panel (random forest marker selection)
  • count 640 hybrids simulated (160 each F1, BC1, BC2, BC3) (gscramble simulated admixture for panel validation)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used whole-genome sequencing of 151 haploid Algerian drone bees to characterise population structure through PCA (SMARTPCA/EIGENSOFT), unsupervised ADMIXTURE clustering (K=2–12, 50 runs per K, CV-error K selection), and pairwise IBD kinship estimation via a hidden Markov model (hmmIBD). Isolation-by-distance was assessed with Pearson correlations between PC-derived genetic distances (or IBD kinships) and Haversine geographic distances. A random forest classifier (scikit-learn) was used to select a 96-SNP ancestry-informative panel distinguishing Algerian from European lineages, validated against 1,000 random panels and tested on simulated hybrid/backcross genotypes generated with the gscramble R package.

Replicationbiological Sample size151 drone bees from 20 Algerian regions across two field campaigns (2017–2018 and 2021–2022), reduced to 102 after IBD-based family filtering and call-rate filtering; reference panel N=228 from a prior published study (M lineage N=63, C lineage N=148, O lineage N=17) GroupsA.m. intermissa vs A.m. sahariensis; Eastern vs Western Algerian genetic clusters; Algerian vs European lineages (M: mellifera/iberiensis; C: ligustica/carnica/Royal Jelly; O: caucasia) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Principal Component Analysis (SMARTPCA/EIGENSOFT v7.2.1) Population structure visualisation — full dataset (Algerian + reference) and Algerian-only; also used to validate AIM panel separation 330 (whole dataset after filtering) or 102 (Algerian-only) not stated
ADMIXTURE unsupervised maximum-likelihood clustering (v1.3.0), K selected by lowest mean CV error across 50 runs Admixture analysis with reference populations and within Algerian samples, K=2–12 330 (with reference) or 102 (Algerian-only) not stated
hmmIBD hidden Markov model — pairwise IBD kinship estimation Full 151 samples for first-degree relatedness screening; 102 retained samples for isolation-by-distance analysis; two genetic clusters separately for close-kin comparison (kinship ≥10% threshold for 3rd-degree relatives) 151 initially; 102 after filtering not stated
Pearson correlation Top 7 PCs vs. three spatial coordinates (21 correlations); pairwise PC-derived Euclidean genetic distances vs. pairwise Haversine geographic distances; pairwise IBD kinships vs. pairwise geographic distances 102 (Algerian samples) not stated
Random Forest classifier (scikit-learn RandomForestClassifier, 100 independent runs per configuration) Selection and validation of 96 AIMs distinguishing Algerian from M-lineage and C-lineage European bees; performance benchmarked against 1,000 random 96-SNP panels (each run 50 times) Training N=250 (82 Algerian + 50 M-lineage or 118 C-lineage); testing N=63; extended to N=570/383 with simulated hybrids not stated
Approaches that could also have been used
  • Pearson correlation was used to assess the association between pairwise genetic and geographic distance matrices
    Could also: A Mantel test (or partial Mantel test controlling for a third matrix, e.g. shared subspecies) could also be used to test isolation-by-distance — Pairwise distance matrices share observations across rows and columns, violating the independence assumption of standard Pearson correlation; the Mantel test uses permutation to generate a null distribution appropriate for matrix correlations, and the partial Mantel extension allows geographic effects to be separated from subspecies membership effects
  • The number of genetic clusters K was selected using ADMIXTURE's cross-validation (CV) error across 50 runs per K
    Could also: The Evanno ΔK method, BIC from sparse non-negative matrix factorization (sNMF in the LEA R package), or STRUCTURE HARVESTER could also be used to guide K selection — CV error identifies K that minimises predictive error on withheld genotypes; ΔK captures the rate-of-change in the log-likelihood and can highlight hierarchical structure at a different level; sNMF provides a regularised likelihood criterion and runs efficiently on large SNP datasets — comparing multiple criteria can triangulate a stable K
  • A random forest classifier was used to rank and select 96 ancestry-informative SNPs based on feature importance (Gini impurity or entropy)
    Could also: FST-based ranking, informativeness-for-assignment (In) scores, or penalised logistic regression (LASSO) could also be used to identify ancestry-informative markers — FST ranking is a classical, transparent one-locus approach for AIM selection; In scores are information-theoretic measures directly interpretable as assignment power per locus; LASSO provides built-in sparsity with a regularisation path tuned by cross-validated accuracy — these methods offer complementary perspectives to tree-based importance scores and can serve as benchmarks
  • IBD kinship was estimated with hmmIBD, a tool originally developed for haploid P. falciparum genomes and adapted here by changing chromosome number and recombination rate parameters
    Could also: KING-robust or a method-of-moments estimator (PLINK2 --make-king) applied to the diploid-model genotype calls could also estimate pairwise relatedness in structured populations — hmmIBD's HMM framework explicitly models recombination breakpoints and segment lengths, which is well-suited to haploid data; KING is designed to be robust to allele-frequency differences between subpopulations — comparing estimates from both approaches would help assess the sensitivity of close-kin calls to the choice of estimator under the observed population structure
  • Genetic distance between pairs of samples was operationalised as a weighted Euclidean distance on the top two PCs (weighted by the eigenvalue ratio of PC1/PC2)
    Could also: Proportion of allele-sharing differences (1 − IBS) computed directly from all SNPs, or Reynolds' co-ancestry distance from allele frequencies, could also serve as pairwise genetic distance inputs — PC-based distance compresses genome-wide information efficiently but depends on the number of PCs retained and the chosen weighting; allele-sharing distances use the full SNP complement without dimensionality reduction and are a standard input for Mantel-type isolation-by-distance analyses, providing a complementary summary of genome-wide differentiation
  • Admixture proportions from ADMIXTURE were used to characterise and describe the absence of large-scale introgression from European lineages into Algerian bees
    Could also: D-statistics (ABBA-BABA tests) or f3/f4-statistics from the ADMIXTOOLS package could also formally test for gene flow between specific population pairs — ADMIXTURE proportions reflect the proportion of ancestry attributable to each cluster but do not assign a formal test statistic or p-value to introgression; D- and f-statistics use the site-frequency spectrum of four-taxon configurations to detect asymmetric gene flow and provide Z-scores against a permutation null, offering a complementary and more formal inference of admixture direction and magnitude
Software: BWA-MEM 0.7.15 · GATK (HaplotypeCaller, CombineGVCFs, GenotypeGVCFs, SelectVariants) 4.1.2 · PICARD MarkDuplicates · PLINK (indep-pairwise LD pruning) · SMARTPCA / EIGENSOFT 7.2.1 · ADMIXTURE 1.3.0 · pong (run alignment and clustering) · hmmIBD · scikit-learn / Python (RandomForestClassifier) · gscramble (R package, for simulated hybrids) · MATLAB (geographic maps)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
3
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

PRJNA1044268 BioProject in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet

Downstream reach in the literature

0 downstream papers · 1 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

PRJNA1044268 BioProject reused by 1 papers in the literature

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-38114899

Paper: Salvatore et al. 2023, Natural clines and human management impact the genetic structure of Algerian honey bee populations. Genet Sel Evol. DOI 10.1186/s12711-023-00864-5 · PMCID PMC10729559.

Data: SRA BioProject PRJNA1044268 — raw WGS reads (FASTQ) of 151 haploid drones (Illumina NovaSeq 6000, 2×150 bp). Public and resolvable, but raw reads only — no intermediate genotype matrix / VCF is deposited.

Code (per the paper's "Availability"): only gscramble (https://github.com/eriqande/gscramble) is cited — a generic third-party R package (eriqande, CRAN v1.0.1) used for one validation step (simulating F1/backcross genotypes). The authors' own analysis pipeline (alignment, variant-calling, ADMIXTURE/PCA/IBD wrappers, Random-Forest AIM selection, gscramble driver, cline analysis) is NOT deposited anywhere — no authors' GitHub/Zenodo analysis repo was found (only SRA raw reads + a Figshare sample table, Additional file 2).

Pipeline map per reported result

Reported result Pipeline / tool In scope? Why
8,865,912 SNPs (post-QC); 1,305,303 / 1,374,698 LD-pruned BWA-MEM 0.7.15 → GATK 4.1.2 HaplotypeCaller + joint genotyping → PLINK LD-prune OUT (80/20) Requires aligning + joint-calling 151 WGS samples (≈TB-scale, many CPU-days) and the wrapper scripts/parameters are not deposited; intermediate VCF not deposited. Heaviest + least-specified part.
ADMIXTURE K=4 (ref+Alg), K=2 (Alg only); 50 runs/K ADMIXTURE 1.3.0 + pong OUT Depends on the (undeposited) SNP VCF and an external reference panel of 228 M/C/O samples that is not part of PRJNA1044268.
PCA PC1 28.23 %, PC2 16.07 % SMARTPCA (EIGENSOFT 7.2.1) OUT Depends on the undeposited VCF + external reference panel.
Geo-genetic r (45 % / 71.6 % / 25.9 %); IBD kinship r (−25.5 % / −63.1 % / −21.3 %); 33/5151 close pairs hmmIBD + custom correlation/cline scripts OUT Depends on the undeposited VCF; correlation/cline scripts not deposited.
96-SNP AIM panel; 100 % test accuracy on 63 samples Random-Forest (scikit-learn) OUT Selection script not deposited; depends on the undeposited VCF.
Hybrid misclassification BC2M 1.4 %, BC3M 7.25 % gscramble (F1/BC gene-drop) + RF classifier PARTIAL (tool only) gscramble itself is deposited & runnable. But reproducing the paper's number needs the real founder genotypes, an Amel_HAv3.1 recombination map, the 96-SNP panel, the exact pedigree design and the trained RF — none of these gscramble inputs are deposited.

What we actually attempt (the auditable 80/20 floor)

The only deposited, executable code artifact is gscramble. We therefore:

  1. Install gscramble pinned to v1.0.1 (commit 7018bf787f741ec361b9f526a107a6d9fd2d5b06) in a conda R env on «our HPC».
  2. Run the package's own R CMD check / testthat suite (its baked-in expected values) → deterministic confirmation the cited code builds and behaves as published.
  3. Compute a fingerprint of the shipped example data + run a set.seed() gene-drop example → a concrete, re-runnable reproduced value.

This reproduces the cited tool, not the paper's bee-specific gscramble numbers (inputs not deposited). All paper headline numbers above are graded partial/not-reproducible-from-shipped-artifacts, with the reason recorded.

Honest verdict direction: PARTIAL. The cited code runs and is deterministic; none of the paper's reported numbers are regenerable from the deposited artifacts (authors' pipeline + intermediate genotypes not deposited; raw→SNP step out of 80/20 and needs an external reference panel).

C1
Reported
8,865,912 SNPs post-QC
Reproduced
not attempted
partial
C2
Reported
1,305,303 / 1,374,698 LD-pruned SNPs
Reproduced
not attempted
partial
C3
Reported
ADMIXTURE K=4 / K=2
Reproduced
not attempted
partial
C4
Reported
PCA PC1 28.23% / PC2 16.07%
Reproduced
not attempted
partial
C5
Reported
geo-genetic r 45 / 71.6 / 25.9%
Reproduced
not attempted
partial
C6
Reported
IBD kinship r -25.5 / -63.1 / -21.3%
Reproduced
not attempted
partial
C7
Reported
33/5151 close family pairs (31 Eastern)
Reproduced
not attempted
partial
C8
Reported
96-SNP AIM panel
Reproduced
not attempted
partial
C9
Reported
100% accuracy on 63 test samples
Reproduced
not attempted
partial
C10
Reported
hybrid misclassification BC2M 1.4% / BC3M 7.25%
Reproduced
gscramble tool reproduces deterministically; paper's bee-specific number not regenerable (founder genotypes + recomb map + 96-SNP panel + RF not deposited)
partial
C11
Reported
gscramble cited code builds & runs deterministically
Reproduced
v1.0.1 @ 7018bf7 builds+installs (R 4.3.3); all functional R CMD check stages OK (only cosmetic vignette image-include ERROR; no testthat suite ships); set.seed(15) segregate()=36-row tbl (digest cedbcccf) -> segments2markers()=78x200 geno matrix (digest d31fd8b7)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 55/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🔴6. Severity of the deviation
🔴7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

This study is not reproducible from deposited artifacts: only raw FASTQ sits in SRA (no VCF, no reference panel) and the authors' entire analysis pipeline is deposited nowhere, so claims C1–C9 (SNP counts, ADMIXTURE K, PCA variance, IBD/geo correlations, the 96-SNP AIM panel and its 100%/63 accuracy) are all unverifiable — non-derivable from shared data+code rather than demonstrably wrong. The failure is squarely on the authors' deposit side (missing scripts + intermediate genotypes), not a methodology choice or a legitimate data restriction. The lone deterministic success is the cited third-party package gscramble (v1.0.1 @7018bf7, gene-drop digest d31fd8b7), a tool-level meta-claim that reproduces no paper number; C9's 100% accuracy is additionally flagged as a too-perfect, uncheckable figure.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

168.9 k
tokens (I/O) · 12.5 M incl. cache
35 min
runtime · 0.04 CPU-h
2.5 GB
peak RAM
1
HPC jobs
hummel
machine