Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Optimisation of the core subset for the APY approximation of genomic relationships.

Genet Sel Evol · 2022
L1 85/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce -> YES for the in-scope target; result is 1:1 (within tolerance). This is a pure AlphaSimR SIMULATION paper: the Zenodo 'data' deposit ships only the 3 code files, no external dataset. PRIMARY in-scope target reproduced: the 8 simulated-cattle core subset sizes (n largest eigenvalues of the centred 15000x10000 marker matrix capturing 10/30/50/70/90/95/98/99% of variance). Faithful re-run of OptimisedCore4APY.R L19-126 on «our HPC» (SLURM 2175425; conda r-base 4.5.3 + AlphaSimR 2.1.0 compiled from CRAN). Reproduced 10,51,136,323,937,1463,2307,3035 vs reported 10,50,135,326,968,1516,2386,3129 -> all 8 within ~3.5% relative (one exact); structural counts 15000 genotyped x 10000 SNP exact. CRITICAL HONESTY NOTE: the original script has NO set.seed(), so this is an independent fresh realisation and an exact match is impossible by design -- the close agreement is strong evidence the reported core sizes are genuine and derivable from the shipped code, but the human auditor should treat exactness as not-applicable. NOT ATTEMPTED (the hard 20%): (1) GBLUP prediction-accuracy curves per core-selection method (Fig 2-7) require the external proprietary BLUPF90 binary (UGA) and are stochastic over 5 reps -> deferred; (2) all real-pig results (Fig 8) use proprietary on-request PIC data (data_restricted) -> out of scope. Status 'partial': the cleanly-specified low-hanging pipeline output reproduced within tolerance; the accuracy figures were deliberately not chased.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.7181323

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 85
    assessed: 2026-06-14 ⛓ 6e2d32b7dc75
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Whether the core subset of genotyped animals used in the APY approximation of genomic relationships can be chosen deliberately (rather than at random) to make the resulting genomic predictions more accurate and stable/repeatable.

Core claims
  • APY approximates the full genomic relationship matrix by splitting genotyped animals into a core subset (fully dependent, direct inverse) and a non-core subset (conditionally independent given core), reducing inversion cost. method
  • Random selection of the core subset can make GEBV results unstable and can cause animals to be reranked between runs on the same data. finding
  • A novel conditional algorithm for optimising the core subset is derived, based on a conditional genomic relationship matrix or conditional SNP genotype matrix. method
  • All core subset construction methods (random, diagonal-based, weighted-random, conditional) perform equally well once the core subset size captures most of the variation in genomic relationships. finding
  • When the core subset is not sufficiently large, random core selection produces substantial variability in results, whereas the conditional algorithm produces no variability. finding
  • The conditional core subset construction spreads core animals across the whole domain of genotyped animals (visualised via PCA and UMAP) in a repeatable manner. finding
  • The size of the core subset in APY is a critical determinant of accuracy. finding
Experimental setups
Assay System Perturbation Readout Platform
GBLUP genomic prediction / APY core subset comparison simulated cattle population (AlphaSimR, 15,000 genotyped/phenotyped animals, generations 15-20) none (simulated selection breeding program) accuracy of GEBV in 3000-animal validation subset (generation 20) AlphaSimR R package
GBLUP genomic prediction / APY core subset comparison real pig data set (PIC; 42,868 phenotyped pigs, 49,788 genotyped pigs across purebred L1, crossbred F1, backcross BC1/BC2) none accuracy of GEBV in 478-animal validation subset (youngest animals born 2021) 42,707 SNP genotyping panel after quality control
Population structure visualisation (linear PCA and non-linear UMAP) genotyped animals from both simulated cattle and real pig data sets none spatial distribution/spread of chosen core animals across the domain of genotyped animals
Key results
  • All core subset constructions performed equally well when the number of core animals captured most of the variation in genomic relationships, in both simulated and real data sets
  • With insufficiently large core subsets, random construction showed substantial variability in accuracy across replicate runs
  • With insufficiently large core subsets, the conditional algorithm showed no variability in results across replicate runs
  • Conditional construction spread core animals across the entire domain of genotyped animals in a repeatable pattern, as shown by PCA/UMAP visualisation
Key statistics
  • other heritability = 0.30 (simulated cattle phenotype simulation)
  • other sigma_a^2 = 1.00, sigma_e^2 = 2.33 (GBLUP model variance components for simulated cattle data)
  • count 15,000 genotyped and phenotyped animals (simulated cattle data set (generations 15-20))
  • count 3000 genotyped validation animals (simulated cattle data, generation 20 validation subset)
  • count 42,868 phenotyped pigs (33,544 L1, 114 F1, 9210 BC1/BC2) (real pig data set phenotypic records, 1999-2021)
  • count 49,788 genotyped pigs (37,598 L1, 486 F1, 11,704 BC1/BC2) at 42,707 SNPs (real pig data set genotyping)
  • count 478 validation animals (phenotyped and genotyped, born 2021) (real pig data validation subset)
  • other core subset sizes set to number of eigenvalues capturing 10, 30, 50, 70, 90, 95, 98, and 99% of variation in G (core subset size levels compared across both data sets)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This methodological paper compares four approaches to constructing core subsets for the APY (Algorithm for Proven and Young) sparse approximation of the genomic relationship matrix inverse, evaluated on a simulated cattle dataset (15,000 genotyped animals, 3,000 holdout validation) and a real pig dataset (49,788 genotyped animals, 478 holdout validation). Both datasets were analyzed using GBLUP mixed models; the pig model additionally included contemporary group fixed effects and shared litter random effects. Performance was compared across eight levels of explained genomic variation (10–99%), and population structure was visualized with linear PCA and non-linear UMAP.

Replicationunclear Sample sizeSimulated cattle: 15,000 genotyped animals across 6 generations (training: generations 15–19; validation: 3,000 in generation 20); real pig: 49,788 genotyped pigs total, 478 youngest (born 2021) held out; 4 construction methods × 8 core sizes defined by 10/30/50/70/90/95/98/99% of variance in G Groups4 core subset construction methods (random; diagonal-of-G-based; weighted random using diagonal; novel conditional algorithm) × 8 levels of genomic variation explained Pairingna Randomization/blindingnot stated Dispersionunclear Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
GBLUP (Genomic Best Linear Unbiased Prediction) via Henderson's mixed model equations Simulated cattle dataset: genomic prediction across all core subset construction methods and sizes 15,000 genotyped+phenotyped animals (training, generations 15–19); 3,000 validation (generation 20) stated
GBLUP with contemporary group fixed effects and shared litter random effects Real pig dataset: genomic prediction across all core subset construction methods and sizes up to 49,788 genotyped pigs (training); 478 youngest animals held out (born 2021) stated
GEBV accuracy (Pearson correlation between predicted GEBV and true or adjusted phenotypic value in the validation set — standard metric in this field) Primary comparison metric for all core subset methods and core sizes in both datasets 3,000 (simulated validation); 478 (pig validation) not stated
Linear Principal Component Analysis (PCA) Visualization of population structure of genotyped animals to illustrate placement of core animals na
Non-linear UMAP (Uniform Manifold Approximation and Projection) Visualization of population structure of genotyped animals to illustrate placement of core animals na
Approaches that could also have been used
  • The variability of the random core subset construction is described qualitatively, but the number of independent random draws used to characterize this variability is not stated in the available text
    Could also: Reporting the distribution of GEBV accuracy over a fixed number of random replicate draws (e.g., 50–100) with mean ± SD or IQR would also quantify variability — An explicit distribution over random draws directly measures the instability that motivates the conditional algorithm and provides a numerical basis for comparing the magnitude of variability across methods
  • A single temporally structured holdout set (most recent generation or birth year) was used to assess GEBV accuracy in each dataset
    Could also: Multiple cross-validation folds structured by generation or year (e.g., rolling-window or leave-one-year-out) could also be used to estimate prediction accuracy — Multiple holdout partitions yield a distribution of accuracy estimates, reducing dependence on the particular characteristics of a single split and enabling uncertainty quantification around differences between methods
  • Variance components for the simulated cattle GBLUP model were fixed at known values (σ²_a = 1.00, σ²_e = 2.33) rather than estimated from data
    Could also: Estimating variance components via REML from the training data is standard practice and could also be applied here — REML estimation is the default approach in applied genetic evaluations; understanding whether results differ between fixed-true and REML-estimated components would clarify how closely the simulation setting approximates real-data conditions
  • The number of core animals was determined by selecting the number of eigenvectors capturing predefined discrete percentages (10, 30, 50, 70, 90, 95, 98, 99%) of variation in G
    Could also: A formal eigenvalue gap (scree-plot elbow) analysis or an information-theoretic criterion applied to the spectral decomposition could also be used to identify the minimum sufficient core size — Analytic data-driven criteria for selecting the number of principal components are well established and would complement the discrete percentage thresholds by precisely locating where accuracy stabilizes without requiring a predefined grid
  • Accuracy differences between core subset construction methods were compared numerically without formal statistical tests or uncertainty bounds
    Could also: Bootstrap confidence intervals around Pearson correlations (or Fisher z-transformed differences between correlations) could also be reported to characterize the precision of accuracy estimates — Bootstrapping the validation set provides uncertainty bounds on accuracy differences, helping readers judge whether observed differences between methods exceed what would be expected from sampling variability in the holdout set
  • Population structure was visualized with PCA (linear) and UMAP (non-linear) to show how conditional core animals spread across the genotypic space
    Could also: t-SNE (t-distributed Stochastic Neighbor Embedding) is another widely used non-linear embedding that could also be used for this visualization — PCA, UMAP, and t-SNE differ in how they balance preservation of global versus local structure; including more than one non-linear method would show whether the observed spread pattern of conditional core animals is robust to the choice of embedding
Software: AlphaSimR (R package)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
13
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.
Cited by (assessed papers) (1)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36418945

Title: Optimisation of the core subset for the APY approximation of genomic relationships. Pocrnic, Lindgren, Tolhurst, Herring & Gorjanc (2022), Genet Sel Evol 54:76. DOI 10.1186/s12711-022-00767-x · Code: github.com/HighlanderLab/ipocrnic_OptimisedCore4APY · Data: zenodo 10.5281/zenodo.7181323

Nature of the study

This is a methods / simulation paper. The Zenodo "data" deposit contains only the 3 code files (OptimisedCore4APY.R, functions.R, README.md) — there is no external dataset to download. The simulated-cattle data is generated in code by AlphaSimR (runMacs, CATTLE). A second, real pig dataset (49,788 genotyped animals, Pig Improvement Company) is used but is proprietary / not shipped → out of scope (data_restricted).

Pipeline map (which reported results come from a pipeline)

Result Pipeline In scope?
Core subset sizes (n eigenvalues capturing 10/30/50/70/90/95/98/99 % of marker variance), simulated cattle AlphaSimR sim → centred marker matrix Z → svd(Z) / eigen(ZZ') → count dims by cumulative variance (coresize.txt) YES — primary target. Pure R, no external binary, fully derivable from shipped code.
GBLUP prediction accuracies per core-selection method (random / conditional-prior / diagonal / weighted-diagonal / SVD-prior) vs full inverse — Fig 2–7 (sim) AlphaSimR sim → APY G-inverse (R) → BLUPF90 GBLUP → cor(EBV, TBV) Partial / deferred (the hard 20%). Requires external proprietary BLUPF90 binary (UGA), is stochastic (5 reps, no set.seed), and is many runs. Assessed for feasibility; not the primary deliverable.
Real-pig core sizes & accuracies (Fig 8) same, on PIC pig data NO — out of scope. Data is proprietary/on-request (data_restricted).

Reproducibility caveats (honesty)

  • No set.seed() anywhere in OptimisedCore4APY.R. Every run is a fresh AlphaSimR realisation, so an exact bit-for-bit match is impossible by design. The eigenvalue spectrum of a 15,000×10,000 simulated marker matrix is, however, statistically stable, so the core sizes should land close to the reported values. Realistic best grade = within-tolerance, not exact.
  • Faithful pipeline: 3000 founders → select 50 sires + 1500 dams → 20 generations of randCross2 (1500 crosses × 2 progeny = 3000/gen), genotypes pulled for the last 5 generations (gen 16–20) = 15,000 animals, 10 chr × 1000 SNP = 10,000 SNP, h² = 0.30. Matches paper Methods.
  • OptimisedCore4APY.R line 13 sources functions_opticore.R but the shipped file is functions.R (minor filename drift; same content) — handled.

What we attempt (80/20)

Primary (low-hanging, clearly specified): regenerate the 8 simulated-cattle core sizes and compare to the paper's reported 10, 50, 135, 326, 968, 1516, 2386, 3129. Not attempted: BLUPF90 accuracy curves (external binary + stochastic + heavy) and all real-pig results (proprietary data).

Figures / tables: Fig 2
coresize_e10
Reported
10
Reproduced
10
exact
coresize_e30
Reported
50
Reproduced
51
within tolerance
coresize_e50
Reported
135
Reproduced
136
within tolerance
coresize_e70
Reported
326
Reproduced
323
within tolerance
coresize_e90
Reported
968
Reproduced
937
within tolerance
coresize_e95
Reported
1516
Reproduced
1463
within tolerance
coresize_e98
Reported
2386
Reproduced
2307
within tolerance
coresize_e99
Reported
3129
Reproduced
3035
within tolerance
n_genotyped
Reported
15000
Reproduced
15000
exact
acc_full
Reported
0.74
Reproduced
not-attempted
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

The in-scope target — the 8 simulated-cattle core subset sizes — reproduces cleanly: all eight land within ~3.5% relative (one exact) from the shipped code, and since the original has no set.seed() this close agreement is strong evidence the reported values (10,50,135,326,968,1516,2386,3129) are genuine and derivable. The small deviations sit entirely on the input/stochastic side, not in the computation — a technical/expected cause, no fabrication concern. However, the paper's central scientific claim (GBLUP/APY prediction-accuracy per core-selection method, Fig 2–7, and the real-pig Fig 8) was not attempted because it needs the proprietary BLUPF90 binary and restricted PIC data, so the overall conclusion is only partially verified — hence a solid-but-limited yellow rather than full green.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

131 k
tokens (I/O) · 12.9 M incl. cache
19 min
runtime · 0.88 CPU-h
11.6 GB
peak RAM
1
HPC jobs
hummel
machine