Uncertainty in the mating strategy of honeybees causes bias and unreliability in the estimates of genetic parameters.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Any deviation was negligible
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Kistler et al. 2024 (Genet Sel Evol 56:23) is a SIMULATION study. Its headline results (Tables 2-4: bias and relative SE of honeybee worker/queen variance components under correct vs mis-specified mating designs) are NOT reproducible from shipped artifacts, because the authors explicitly state 'The R simulation code used during the current study is not publicly available while the authors continue to perform active software developments', and no simulated datasets were deposited (the Zenodo 'data' DOI 10.5281/zenodo.7951334 is a tagged zip of the AINV-honeybees CODE, md5 4f5a95a1901f13e36aedeb093a9fb311). The variance-component estimator (AIREMLF90/BLUPF90) is free but has no published input data. The ONE public, shipped, runnable pipeline component is the AINV-honeybees relationship-matrix tool (Brascamp), which the paper used to build the honeybee A-inverse; per BRIEF rule 2 / P16 this is an equally valid reproduction target, and it was reproduced FRESH on «our HPC» (SLURM «job», R 4.1.3, commit a5ae26a). It is correct and reproducible by every objective criterion the methodology defines: the program's own A*A-inverse=I self-check passes with 0 deviating elements (pinv=0 branch); two fully independent inversion algorithms (Mendelian-sampling factorization vs direct solve) agree to ~1e-13; and every hand-derivable closed-form honeybee-genetics anchor matches exactly (base queen Aii=1, base sire Aii=1/NS=1/3, base colony variance 0.395833=0.25+1/16+1/12, full-sib-queen relationship 0.395833, queen-dam 0.5). The honeybee-specific worker inbreeding F=0.124131944 was upgraded from the prior run's 'partial' to EXACT by independent hand-derivation (F = A[queen,sire]/2, with A[queen,sire] decomposed exactly via the tabular rule). As an additional in-scope step the shipped base-variance formula was used to DEMONSTRATE the paper's mechanism: mis-specifying the assumed mating structure (NS/ND) deterministically shifts the base relationship structure by up to +153%, which is the source of the bias/unreliability the paper quantifies. No worked numeric A/A-inverse is printed in either Kistler 2024 or Brascamp & Bijma 2014, so validation is against the tool's self-check, cross-algorithm agreement and theory -- all exact. Tables 2-4 and the AIREMLF90 REML step were assessed and not attempted (genuine withheld-code blocker, recorded above); a from-scratch re-implementation of the stochastic simulator was deliberately avoided as it would be a new study that could not faithfully reproduce the authors' RNG-dependent ensemble. No fabrication concern surfaced: no reported value was found to be non-derivable; the headline gap is a reproducibility limitation from withheld code, not evidence of fabrication. This is an honest partial reproduction -- the public tool is fully verified to machine precision; the paper's own simulation results remain unreproducible because the generating code is not released.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 87assessed: 2026-06-16 ⛓ 71f54474b44b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether uncertainty about honeybee mating strategy (single drone-producing queen vs. group of sister drone-producing queens, and open-mating of drone-producing queens with heterogeneous drone populations), and how this is modeled in the pedigree/statistical model, biases estimates of genetic parameters and genetic trends in honeybee breeding programs.
- ★ The most precise estimates of genetic parameters and genetic trends are obtained when breeding queens are mated with drones of a single DPQ that is correctly assigned in the pedigree (SS mating). finding
- ★ When the true mating strategy (single DPQ vs. group of sister-DPQs) is unknown and incorrectly assumed, estimates of genetic parameters and trends become considerably biased. finding
- ★ Adding a dummy pseudo-sire in the pedigree for each open-mating drone subpopulation to correct for heterogeneous open-mated DPQ phenotypes considerably overestimates genetic variances. finding
- ★ Adding a non-genetic (fixed or random) effect to the statistical model to account for the heterogeneous drone population mating open-mated DPQs produces unbiased estimates. finding
- ★ Breeders should keep track of individual DPQs in the pedigree rather than only the dam of the DPQ(s) used in a mating. method
- An infinitesimal-model simulation framework (following Kistler et al.) was used to simulate founders, mating, selection and reproduction of honeybee colonies over multiple generations. method
- Colony phenotypes are affected by two genetic effects (queen effect and worker effect) plus their covariance, rather than a single genetic variance. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Simulation Set I: sire-modeling scenarios for controlled mating of breeding queens (SS vs PS mating), compared under C_SS_P, C_dummySS_P/DPQdam, C_dummySS_P/Q, and C_PS_P pedigree scenarios | simulated honeybee (Apis mellifera) population/colonies | controlled mating strategy (single sire vs pseudo-sire) with correct or incorrect pedigree assumptions | bias and standard error of estimated genetic parameters (queen/worker variances, covariance) and genetic trends | — |
| Simulation Set II: sire-modeling scenarios for open mating of drone-producing queens with a heterogeneous drone population, compared under O_NoPheno, O_TwoPS_P, O_FixedGroup, and O_RandGroup scenarios | simulated honeybee (Apis mellifera) population/colonies | open mating of DPQs to two genetically distinct drone subpopulations, modeled via exclusion, pseudo-sire pedigree assignment, or fixed/random non-genetic effect | bias and standard error of estimated genetic parameters and genetic trends | — |
| Sensitivity analysis across alternative genetic parameter sets (varying worker variance and queen-worker genetic correlation) | simulated honeybee (Apis mellifera) population/colonies | varying simulated genetic variance of worker effects (1x or 2x) and genetic correlation (0 or -0.5) between queen and worker effects | estimated genetic parameters under each mating/sire-modeling scenario | — |
- – SS mating with correct single-sire assignment in the pedigree yielded the most precise (lowest SE) estimates of genetic parameters and trends compared to PS mating.
- – Erroneous assumptions about whether drones originated from a single DPQ or a group of sister-DPQs led to considerable bias in genetic parameter and trend estimates.
- ▲ Assigning a dummy pseudo-sire per open-mating drone subpopulation (O_TwoPS_P) considerably overestimated genetic variances relative to true simulated values.
- – Adding a fixed or random non-genetic group effect (O_FixedGroup/O_RandGroup) to account for the heterogeneous open-mating drone population produced unbiased genetic parameter estimates.
- count 432 (number of founder base breeding queens (BQs) in the simulation)
- count 24 (number of BQs selected each year to produce new offspring queens)
- count 20 (number of DPQs produced annually by each selected BQ)
- count 8 (number of drones mating each queen)
- count 3 (number of sister-DPQs forming a pseudo-sire (PS) group)
- count 75% (proportion of colonies surviving annual mortality event, yielding 18 phenotyped sister BQs and 15 phenotyped sister DPQs per dam)
- other 1/3 of residual variance (σe²) (base genetic variance for queen effects (σQ²) and worker effects (σW²) in the founder population)
- other 1.5 years (resulting generation interval in the simulation (1 year dam path, 2 years sire path))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a stochastic simulation study that used an infinitesimal genetic model to generate replicated honeybee breeding datasets under multiple mating-strategy and sire-modeling scenarios, then estimated genetic parameters (queen-effect variance, worker-effect variance, and their covariance) from each simulated dataset. Bias (mean estimate minus true value) and standard error of estimates across simulation replicates were used as the primary performance metrics. Two simulation sets were examined: Set I varied controlled mating strategy (single-sire vs. pseudo-sire) and four pedigree-assignment assumptions; Set II varied how open-mated drone-producing queen records were integrated into the evaluation model. No traditional hypothesis tests were applied; all inference is descriptive comparison of bias and SE across scenarios.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Simulation-based bias estimation (mean estimate minus true parameter value across replicates) | All genetic parameter estimates (queen-effect variance, worker-effect variance, genetic covariance/correlation) across all Set I and Set II scenarios and four genetic parameter sets | — | not stated |
| Standard error of genetic parameter estimates across simulation replicates | All genetic parameter estimates across all scenarios — used as a measure of estimator precision | — | not stated |
| Genetic trend comparison (estimated vs. true genetic trend across scenarios) | Queen and worker genetic trends over 10 simulated years, reported per scenario | — | not stated |
-
The number of simulation replicates used to compute bias and SE is not reported in the main text↳ Could also: Explicitly reporting replicate count and a simulation precision analysis (e.g., minimum replicates to detect a target bias magnitude with specified Monte Carlo error) is also standard practice, as formalized in the ADEMP simulation-study reporting framework — Replicate count allows readers to assess the Monte Carlo uncertainty of the bias and SE estimates themselves — without it, it is unclear how precisely the reported bias values are estimated
-
Bias was reported as a point estimate without an accompanying interval↳ Could also: Monte Carlo confidence intervals around each bias estimate (e.g., mean bias ± 1.96 × MCSE, where MCSE = SD of estimates / √replicates) could also be reported — Confidence intervals around bias estimates would allow readers to judge whether an observed deviation from the true parameter is distinguishable from zero given simulation sampling variability, enabling more formal scenario comparisons
-
Estimator performance was summarized using bias and SE separately↳ Could also: Root mean squared error (RMSE), which jointly captures both bias and variance of the estimator, could also be reported alongside bias and SE — RMSE provides a single scalar summary of total estimation error that accounts for both systematic bias and sampling variability, and is widely used in simulation evaluation studies to rank competing estimators overall
-
Scenario comparisons across the four genetic parameter sets were presented separately (per parameter set) rather than synthesized↳ Could also: A factorial summary (e.g., mixed-effects model with scenario and parameter set as factors and replicate-level bias as outcome) could also characterize whether scenario effects on bias are consistent across parameter sets or depend on specific genetic configurations — A structured cross-scenario analysis would help readers distinguish findings that are robust across genetic architectures from those that are specific to particular parameter combinations
-
Structural simulation parameters (e.g., 8 drones per queen, 3 sister-DPQs per pseudo-sire, 24 selected breeding queens per year, 10-year horizon) were fixed at single values↳ Could also: Sensitivity analyses varying these structural parameters across a plausible range could also be conducted to characterize the robustness of bias findings to breeding program scale and configuration — Sensitivity analyses would show whether reported bias patterns generalize across a wider range of realistic program sizes, or are specific to the chosen parameterization — relevant because breeding program sizes vary considerably in practice
-
Genetic parameter estimation was performed under a set of modeling scenarios but the specific estimation software and algorithm (e.g., REML via ASReml, BLUPF90, or a Bayesian sampler) are not stated↳ Could also: Bayesian MCMC methods (e.g., Gibbs sampling as implemented in THRGIBBS1F90 or MCMCglmm) could also be used to estimate the same parameters, yielding full posterior distributions with credible intervals rather than point estimates and SE — Bayesian approaches naturally quantify parameter uncertainty via posterior distributions, and have been widely applied to honeybee genetic parameter estimation; reporting which estimation method was used also aids reproducibility
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-38632535
Paper: Kistler T, Brascamp EW, Basso B, Bijma P, Phocas F. Uncertainty in the mating strategy of honeybees causes bias and unreliability in the estimates of genetic parameters. Genet Sel Evol 2024;56:23. PMID 38632535 · PMC11022492 · DOI 10.1186/s12711-024-00898-3.
Repo: https://github.com/Tristan-Kistler/AINV-honeybees @ a5ae26a (2023-09-14)
Zenodo DOI: 10.5281/zenodo.7951334 → file Tristan-Kistler/AINV-honeybees-v19.zip
(243 KB) = a tagged snapshot of the same code repo, not a data archive.
What the paper is
A simulation study. The authors simulate honeybee breeding populations under different mating designs (single-drone insemination, group of drone-producing queens = "pseudo-sire", open mating), then estimate genetic (co)variance components (worker effect σ²_W, queen effect σ²_Q, correlation r_WQ) under correct vs mis-specified mating assumptions, and report the resulting bias and relative standard error of those estimates (Tables 2–4). Headline numbers: relative SE ~17–20 % when correctly modelled; up to −32 % bias on σ²_W and +64 % on σ²_W under specific mis-specifications; r_WQ bias −0.14 … +0.25.
Pipeline components and their availability
| Component | Role | Public? | In scope? |
|---|---|---|---|
| Simulation code (R, by the authors) | generates the simulated pedigrees + phenotypes that ALL Tables 2–4 rest on | NO — paper states it "is not publicly available while the authors continue to perform active software developments" | out (code withheld) |
| Simulated datasets | the per-scenario data fed to the estimator | NO — not on Zenodo (Zenodo = code zip only); none shipped | out (data not shipped) |
| AINV-honeybees (R, Brascamp) | builds the honeybee numerator relationship matrix A and its inverse AINV.giv from a pedigree |
YES (GitHub + Zenodo) | IN |
| AIREMLF90 (BLUPF90) | REML variance-component estimation given AINV + data | free, but needs the (withheld) simulated data | out (no input data) |
Conclusion on scope. The paper's reported results (Tables 2–4) are not
reproducible from shipped artifacts: the analysis-defining simulation code is
explicitly withheld and no simulated datasets are published, so the bias/SE
numbers cannot be regenerated. The only public, shipped, runnable pipeline
component is the AINV-honeybees relationship-matrix tool. Per BRIEF rule 2
(P16 — a third-party/utility tool on the paper's own methodology is an equally
valid reproduction target), this tool is what we reproduce.
In-scope reproduction target
Run pedigree-20.r → AMD-AINV-20.r (v20, the version the paper used for its
simulations) on a defined honeybee pedigree and verify the deterministic
A / A⁻¹ construction is correct and reproducible, against objective criteria
that the methodology itself defines:
- Self-consistency: program's built-in check that A · A⁻¹ = I
(
log-AMD-AINV.txt). - Cross-algorithm agreement:
pinv=0(Mendelian-sampling factorization A⁻¹ = (I−M)ᵀD⁻¹(I−M)) vspinv=1(directsolve(A)) must give identical A⁻¹. - Closed-form theory anchors (hand-derivable from Brascamp & Bijma 2014,
DOI 10.1186/s12711-014-0053-9, the method this code implements):
- base queen diagonal A = 1
- base sire (NS unrelated DPQ, equi=0) diagonal A = 1/NS
- base colony variance = 0.25 + 1/(2·ND) + 1/(4·NS)
- full-sib base-queen relationship A[i,j] = same closed form
- queen–dam relationship A = 0.5
- honeybee worker inbreeding F > 0 under a relative-mating loop.
NOT attempted (and why)
- Tables 2–4 (bias / relative SE of σ²_W, σ²_Q, r_WQ): simulation code withheld + no simulated data published → cannot regenerate the inputs.
- AIREMLF90 REML estimation: no input data to run it on.
- No worked numeric A/A⁻¹ example is printed in either Kistler 2024 or Brascamp & Bijma 2014, so there is no published matrix value to match digit-for-digit
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Kistler et al. 2024 is a simulation study whose headline results (Tables 2-4: bias and relative SE of worker/queen variance components under mis-specified mating designs) are not reproducible — the authors openly withhold the data-generating R code and publish no simulated datasets (the Zenodo 'data' DOI is just a zip of the AINV code). This is an authors'-side availability gap (q1/q2/q4 red), not a discrepancy or fabrication: the one public, runnable component, the AINV-honeybees v20 relationship-matrix tool, reproduces exactly (A·A⁻¹=I with 0 deviating elements, pinv0/pinv1 agree to 9.9e-14, all closed-form anchors like 0.395833 and 1/3 match to 15 figures). So everything checkable is 1:1 with zero deviation (q6 green), but the central conclusion is untestable rather than confirmed (q5/q7/q8 yellow). Overall a yellow-criticality withheld-code / data-unavailable case with no fabrication signal.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.