Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Estimates of recent and historical effective population size in turbot, seabream, seabass and carp selective breeding programmes.

Genet Sel Evol · 2021
L1 53/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
How its reproducibility compares
53/100
Reproducibility score
1.2 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 14% of all assessed papers rank 1005 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

GONE (third-party tool, github.com/esrud/GONE) reproduced on «our HPC». The paper's HEADLINE result (per-species contemporary Ne 31/46/32/40/33; parents 26/50/30/32/15; Fig 2 historical Ne) is NOT reproducible from deposited artifacts: GONE's required input is a PLINK .ped/.map SNP genotype panel, and those panels were never deposited -- only raw RAD reads exist (carp PRJNA414021 [20 PE FASTQ, verified ENA], seabass PRJNA407892; turbot/seabream have NO accession), and the read->SNP-panel step (exact aligner/caller/filters yielding 18097/15184/21701/8014/12311 SNPs) is unspecified. Reconstructing panels is the underspecified, high-cost last >=80% with no expected 1:1 match -> not attempted per 80/20. What WAS reproduced: GONE builds+runs end-to-end on «our HPC» on its own shipped EXAMPLE, emitting the full 676-generation Ne trajectory (structure EXACT), internally consistent across 10 runs of 2 code versions (gen-1 Ne ~63 +/- 2). Notable finding: the repo's own shipped reference value Output_Ne_example gen-1=87.4605 is NOT regenerable from the bundled data+code in EITHER the current (2026, commit 2288c61) or the contemporaneous 2021 build (commit 727ae5f) -- both give ~63, a ~27% systematic shortfall (~12 SD, not Monte-Carlo noise; params unchanged). This is a reproducibility gap in the GONE repo, NOT a paper claim; no fabrication asserted against the paper (its Ne values are simply unverifiable from shipped data). NOT bit-deterministic even with a pinned seed. Outcome: partial -- tool pipeline reproduced structurally + characterised; paper's species Ne out of reach because analysis-ready genotypes were never deposited.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 53
    assessed: 2026-06-15 ⛓ f9db48179c0e
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests whether commercial selective-breeding populations of turbot, gilthead seabream, European seabass and common carp have small effective population sizes (Ne) and whether historical Ne shows drops associated with domestication and the onset of selective breeding in European aquaculture.

Core claims
  • Current effective population size for all four farmed fish populations is small (≤50 fish), potentially threatening breeding-programme sustainability finding
  • Important drops in effective population size occurred about five to nine generations ago, likely due to domestication and the start of selective breeding programmes finding
  • The GONE method (Santiago et al.) uses the LD spectrum across the whole range of genetic distances to detect drastic (non-linear) changes in historical Ne, unlike prior linear-only methods method
  • The Hayes et al. LD method only holds for linear changes in population size and fails to reflect sudden drops, giving downwardly biased historical estimates finding
  • Genetic composition of base populations should be broadened and measures to increase Ne implemented to ensure breeding-programme sustainability finding
  • Turbot and seabass show moderately low LD at short distances with rapid LD decay over physical distance finding
  • Estimates from the reduced parental samples are as reliable as those from the more extensive offspring samples finding
Experimental setups
Assay System Perturbation Readout Platform
RAD sequencing (reduced representation genotyping) for SNP genotyping turbot (Scophthalmus maximus) experimental population of Atlantic origin, broodstock and offspring none SNP genotypes used to estimate LD and effective population size RAD-seq; BEAGLE 4.1 for imputation; reference genome GCA_003186165.1
RAD sequencing for SNP genotyping gilthead seabream (Sparus aurata), Andromeda Group (mass spawning) and Ferme Marine de Douhet (partial factorial mating) cohorts none SNP genotypes used to estimate LD and Ne RAD-seq
RAD sequencing for SNP genotyping European seabass (Dicentrarchus labrax), FMD breeding nucleus cohort none SNP genotypes used to estimate LD and Ne RAD-seq; reference genome GCA_000689215
RAD sequencing for SNP genotyping common carp (Cyprinus carpio), Amur Mirror strain (Vodňany line), full factorial crosses none (admixed origin from cultured and wild strains) SNP genotypes used to estimate contemporary Ne RAD-seq
Linkage disequilibrium and effective population size estimation (GONE) all four species' farmed populations none r2/d2 LD measures and temporal series of Ne across generations GONE software (Santiago et al.)
Linkage disequilibrium-based historical Ne estimation (Hayes et al. method) all four species' farmed populations none linear historical Ne trends for comparison method of Hayes et al. as implemented by Saura et al.
Computer simulation of population size drop simulated population (N=1000 or 10,000 dropping to N=100 or 50) sudden population size reduction in last 10 or 5 generations comparison of GONE vs Hayes et al. Ne estimates over 20 replicates
Key results
  • Recent Ne (offspring data): turbot 31, seabream_A 46, seabream_F 32, seabass 40, carp 33 ≤50 fish
  • Recent Ne (parents data): turbot 26, seabream_A 50, seabream_F 30, seabass 32, carp 15 15-50
  • Historical Ne larger than 1000 fish about 20 generations ago in all species, with important drops about five generations ago for turbot and seabream and eight to nine generations ago for seabass >1000 to <50
  • Average LD (r2) between SNPs at short distances (<0.01 kb) was 0.15 for turbot and 0.24 for seabass, halving within distances shorter than 5 kb 0.15 / 0.24
  • At distances longer than 10 Mb, r2 reached values lower than 0.05 r2<0.05
  • Hayes et al. recent Ne (turbot 44, seabass 33, seabream_A 51, seabream_F 49) were of the same order of magnitude as GONE estimates, but its historical estimates at generation 100 were <1000
  • Simulations show Hayes et al. method fails to reflect sudden drops and gives downwardly biased historical sizes, while GONE detects drops but can overestimate large ancestral sizes
Key statistics
  • count Ne ≤ 50 fish (recent estimates, all populations) (current effective population size considered critical threshold)
  • correlation r2 = 0.15 (average LD between turbot SNPs separated by <0.01 kb)
  • correlation r2 = 0.24 (average LD between seabass SNPs separated by <0.01 kb)
  • count Ne = 88 (reported Ne in GIFT tilapia programme after seven generations of selection (comparison))
  • count >1000 (historical Ne about 20 generations ago in all species)
  • count minimum sample×√markers/Ne value of 100 for accurate estimation (power requirement of the GONE method)
  • count Noff turbot 1391, seabream_A 724, seabream_F 881, seabass 1308, carp 1349 (number of offspring genotyped per population (Table 1))
  • count Npar turbot 46, seabream_A 117, seabream_F 107, seabass 65, carp 60 (number of parents genotyped per population (Table 1))

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a population-genomics short communication that estimates current and historical effective population size (Ne) in four farmed fish species from RAD-seq SNP genotypes. The core 'statistical' approach is not hypothesis testing but model-based estimation: linkage disequilibrium (squared correlation r^2 and the variance-weighted d^2 statistic) is computed across genetic-distance bins, and a genetic-algorithm optimization (software GONE) infers a temporal Ne series by minimizing the squared differences between observed and predicted d^2. Results are reported as point estimates of Ne by species and sample type (parents vs offspring), compared against an older linear-trend LD method (Hayes et al.) and supported by 20 replicate forward simulations.

Replicationmixed Sample sizeSample sizes (offspring, parents, SNPs, linkage groups) tabulated per population in Table 1; a method-accuracy threshold (sample size x sqrt(markers)/Ne >= 100) is stated rather than a formal power calculation Groupsfour/five farmed fish populations; parents vs offspring; GONE vs Hayes et al. method Pairingna Randomization/blindingna Dispersionnone Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Linkage disequilibrium estimation via squared allele-frequency correlation r^2 (Hill & Robertson) and variance-weighted average d^2 all pairs of SNPs within each linkage group; LD-decay curves in Fig. 1 (turbot, seabass) and the inputs to Ne estimation per-population SNP and sample counts in Table 1 (e.g. turbot 18,097 SNPs, 1391 offspring, 46 parents) stated
Genetic-algorithm optimization (GONE) minimizing squared differences between observed and predicted binned d^2 to infer a temporal Ne series recent and historical Ne for turbot, seabream_A, seabream_F, seabass, carp (Fig. 2, Additional file 1: Fig. S1) based on parents and offspring samples per Table 1; method requires sample-size x sqrt(markers)/Ne >= 100 for accuracy stated
Hayes et al. LD-based Ne method (assuming constant or linear changes in Ne) for comparison historical and recent Ne re-analysis (Additional file 2: Fig. S2) stated
Forward population simulation (20 replicates) under defined demographic scenarios validation of GONE vs Hayes behaviour under a sudden Ne drop (Additional file 3: Fig. S3) 20 simulation replicates; N constant 1000 or 10,000 dropping to 100 or 50 stated
Approaches that could also have been used
  • Ne is reported as single point estimates per population and sample type (e.g. 31 for turbot offspring).
    Could also: Reporting could ALSO include bootstrap or replicate-based confidence intervals (e.g. resampling SNPs/chromosomes) around each Ne estimate. — Interval estimates would convey the precision of each Ne value alongside the point estimate, which is often informative when estimates are near a decision threshold such as 50.
  • Agreement between parents-based and offspring-based estimates, and between GONE and the Hayes et al. method, is assessed descriptively by comparing values and trend shapes.
    Could also: One could ALSO summarize concordance quantitatively, for example with correlation or mean absolute difference across the matched estimates. — A numeric concordance measure would complement the visual/qualitative comparison and give a single summary of how closely the two data sources or two methods agree.
  • Method accuracy and potential over/underestimation are explored with 20 forward simulations under a few fixed demographic scenarios.
    Could also: The simulation study could ALSO span a wider grid of ancestral sizes, drop timings, marker densities, and more replicates, with summary distributions of estimated vs true Ne. — A broader simulation design would more fully characterize bias and variance of the estimator across the parameter space relevant to these populations.
  • LD decay with physical distance is presented graphically (Fig. 1) for the two species with a physical map.
    Could also: A fitted LD-decay model (e.g. a nonlinear regression of r^2 on distance) with parameter estimates could ALSO be reported. — Fitted decay parameters would provide a compact, comparable numeric summary of LD decay across species in addition to the visual curves.
  • For seabream and seabass the authors note estimates may be slightly affected because the schemes have overlapping generations while the method assumes discrete generations.
    Could also: A sensitivity analysis or a method/correction accommodating overlapping generations could ALSO be applied. — Quantifying how much the discrete-generation assumption shifts the estimates would help bound the noted potential underestimation.
Software: GONE and auxiliary programs (Santiago et al.) · BEAGLE (genotype imputation, turbot only) 4.1

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
49
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GCA_000689215 GCA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GCA_003186165.1 GCA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34742227 (Saura et al. 2021, GenSelEvol)

Paper: Estimates of recent and historical effective population size (Ne) in turbot, seabream, seabass and carp selective breeding programmes. Tool reproduced (P16, third-party): GONE — https://github.com/esrud/GONE @ commit 2288c61d21a1fd21ad01b693f59990f117566448 (master, latest).

GONE (Santiago/Caballero, MBE 2020) estimates the temporal Ne trajectory from linkage-disequilibrium between SNP pairs binned by genetic distance. Input = PLINK .ped/.map. It runs REPS=40 replicate estimates, each on a random SNP subsample (seed = $RANDOM, regenerated every run), and reports the geometric mean Ne per generation → output is stochastic, not bit-reproducible.

In scope (pipeline-derived, attempted)

# result pipeline reproducible?
R1 GONE Ne trajectory on the shipped EXAMPLE (example.ped/.map, default INPUT_PARAMETERS_FILE). Shipped reference: EXAMPLE/Output_Ne_example (gen-1 Ne = 87.4605, plateau ≈ 110). GONE script_GONE.sh example on «our HPC» YES — run the exact tool; compare gen-1 Ne + trajectory to the shipped reference, within Monte-Carlo tolerance (40-rep geom-mean, random SNP subsampling → not exact).
R2 GONE reproducibility character: re-running yields gen-1 Ne within a tight band around 87.46; fixing the seed makes it deterministic. repeat R1 ×5 YES — quantify spread.

R1 reproduces the exact tool and pipeline the paper used, on the tool's own reference data, with a shipped expected output — the cleanest auditable point.

Out of scope / NOT attempted (the hard ≥80%) — and why

The paper's headline numbers are species-specific contemporary Ne (offspring: turbot 31, seabream_A 46, seabream_F 32, seabass 40, carp 33; parents: 26/50/30/ 32/15; Fig. 2 historical Ne). Reproducing these 1:1 is not feasible from the deposited artifacts:

  • GONE's actual input (the per-species .ped/.map genotype panels) is NOT deposited. Only raw RAD-seq reads are: carp PRJNA414021 (20 paired FASTQ runs, Cyprinus carpio) and seabass PRJNA407892. Turbot and seabream have no data accession in the Availability statement at all.
  • The read→genotype step (reference genome build, aligner, variant caller, and the filters that yield exactly 18 097 / 15 184 / 21 701 / 8 014 / 12 311 SNPs) is not specified in Methods → the SNP panel cannot be reconstructed.
  • Even with a panel, GONE's stochastic subsampling means an exact match to 31/46/… is not expected; only a faithful approximate trajectory.

Attempting carp end-to-end (download ~20 FASTQ runs → align to C. carpio ref → call+filter SNPs → GONE) is the underspecified, high-cost last 20% with no expected 1:1 match. Per brief (80/20, "do not chase the last 20%"), not attempted; documented as the honest scope boundary.

Drop vs partial

Not a drop: the tool + a runnable reference dataset with shipped expected output exist and reproduce. Outcome = partial — GONE pipeline reproduced on its reference data (within Monte-Carlo tolerance); the paper's species Ne values are out of reach because the genotype inputs were never deposited.

Figures / tables: TableFig 1Fig 2
TOOL-struct
Reported
GONE shipped-example output = 676-generation geometric-mean Ne trajectory
Reproduced
676 generations, identical format, on «our HPC»
exact
TOOL-ref-gen1
Reported
87.4605 (EXAMPLE/Output_Ne_example, gen-1 Ne)
Reproduced
~63 (range 59.17-66.14 over 10 runs / 2 code versions)
did not match
TOOL-determinism
Reported
implicit reproducibility of GONE
Reproduced
run-to-run consistent (~63+/-2) but not bit-identical; pinned seed still differs (62.06 vs 59.17)
partial
PAPER-Ne-species
Reported
contemporary Ne: turbot 31, seabream_A 46, seabream_F 32, seabass 40, carp 33 (offspring); 26/50/30/32/15 (parents); Fig 2 historical Ne
Reproduced
NOT_ATTEMPTED
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 53/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

The paper's headline result — per-species contemporary Ne (offspring 31/46/32/40/33; parents 26/50/30/32/15) and Fig.2 historical Ne — cannot be reproduced from deposited data: GONE requires .ped/.map SNP panels that were never deposited, only raw RAD reads exist for 2/5 datasets, and the read→panel filtering is unspecified. This places the blocker squarely on the authors'/data-availability side, so the central claims are neither confirmed nor refuted (q7 limited), not flagged as fabrication. Separately, the GONE tool runs end-to-end and reproduces the trajectory structure exactly, but the repo's own shipped reference (gen-1 Ne=87.4605) is not regenerable in either the 2021 or 2026 build (~63, ~27%/~12SD shortfall) and is non-deterministic even with a pinned seed — a tool-repo reproducibility gap, not a paper defect. Overall the study is effectively non-reproducible from what was shared.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

225 k
tokens (I/O) · 18.5 M incl. cache
36 min
runtime · 1.69 CPU-h
0.3 GB
peak RAM
2
HPC jobs
hummel
machine