Strong population differentiation in lingcod (Ophiodon elongatus) is driven by a small portion of the genome.
Part of the results reproduced; minor but material deviations remained.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
salvaged by watchdog from agreement.json (agent omitted ROOM_RESULT.json)
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18no human curator yet
- Last updated
- 2026-07-29
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusDoes lingcod (Ophiodon elongatus) exhibit genetic population structure across its west coast North American range when assessed with high-resolution genomic (RADseq) data, in contrast to prior studies that found high gene flow and little structure?
- ★ Lingcod comprise two distinct genetic clusters separated latitudinally at a break near Point Reyes off Northern California, with a high frequency of admixed individuals near the break. finding
- ★ Strong population differentiation is driven by a small portion of the genome, as most loci have low F_ST indicating high gene flow genome-wide. finding
- ★ Outlier and fixed loci are concentrated on a single genomic region/chromosome, a pattern consistent with chromosomal inversions seen in diverse taxa. mechanism
- ★ Outlier analyses identified 182 loci putatively under divergent selection; 71 loci were fixed between northern and southern clusters and all were among the outliers. finding
- RADseq across thousands of polymorphic loci provides increased power to detect population structure and adaptive loci compared with allozymes, mtDNA, and microsatellites. method
- The identified genetic clusters represent clear evolutionary units that could inform fisheries management. finding
- A draft lingcod genome assembly plus three chromosome-level teleost genomes were used to map RADseq loci and localize differentiated loci. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RADseq (restriction site-associated DNA sequencing) | Ophiodon elongatus (lingcod), 611 individuals from 42 sites, Southeast Alaska to Baja California | none | SNP genotypes / polymorphic loci for population structure and outlier analysis | PstI digestion; Illumina HiSeq 4000, 100 bp paired-end (Ali et al. 2016 protocol) |
| Whole genome sequencing (draft genome assembly) | single male lingcod, gill and liver tissue, Hood Canal (Salish Sea) | none | draft lingcod genome assembly for mapping RADseq loci | high molecular weight DNA, flash-frozen tissue |
| Bayesian clustering / population structure analysis | lingcod RADseq genotypes (611 individuals) | none | number of genetic clusters (K), membership coefficients (Q values) | STRUCTURE v2.3.4, Structure Harvester, CLUMPP v1.1.2, DISTRUCT v1.1 |
| PCA and DAPC (multivariate ordination) | lingcod RADseq genotypes | none | diversity/variation across loci, optimal cluster number | R package adegenet v2.1.1 |
| F-statistics / population genetic differentiation | lingcod grouped by 42 sampling sites and by cluster assignment | none | pairwise F_ST, locus-specific F-statistics, fixed loci | — |
| Outlier locus detection | lingcod RADseq loci | none | loci under putative divergent selection (182 outliers, 71 fixed) | — |
| mtDNA sequence diversity analysis | subset of strongly differentiated lingcod individuals | none | mtDNA sequence diversity | — |
| Comparative genome alignment | lingcod draft genome plus three teleost chromosome-level assemblies | none | genomic location/chromosomal distribution of outlier and fixed loci | — |
- – Two distinct genetic clusters separating northern and southern sampling sites with a break near Point Reyes, Northern California
- ▼ Most loci characterized by low F_ST, indicating high gene flow throughout most of the genome
- – 182 loci identified as putatively under divergent selection, most mapping to a single genomic region 182 loci
- – 71 loci fixed between northern and southern clusters, all identified in outlier scans 71 loci
- – Admixed individuals showed near 50:50 assignment to northern and southern clusters and were heterozygous for most fixed loci ~50:50
- – Outlier and fixed loci concentrated on a single chromosome, consistent with chromosomal inversion patterns
- count 16,749 RADseq markers (markers used in analysis)
- count 611 individuals (final analyzed samples across the species range)
- count 182 loci (outlier loci putatively under divergent selection)
- count 71 loci (loci fixed between northern and southern clusters)
- count 42 sites (sampling sites encompassing most of species range)
- count 940 samples extracted; 841 sequenced (DNA extracted from 940, 841 met concentration threshold for RADseq)
- other 1,628 metric tons in 2016 (U.S. lingcod landings, with average growth >200 metric tons/year since 2010)
- count up to 500,000 eggs (egg deposition per female during spawning)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This population genomics study applied 16,749 RADseq loci genotyped in 611 lingcod individuals from 42 sites spanning Southeast Alaska to Baja California to characterize range-wide genetic structure without a priori population assumptions. Bayesian clustering (STRUCTURE), PCA, and DAPC were used to identify genetic clusters, and pairwise FST was calculated by sampling site and by cluster assignment to quantify differentiation. Outlier analyses were conducted to identify loci putatively under divergent selection, with results mapped to a draft lingcod genome assembly and three reference teleost genomes; the paper's Methods section is truncated in the supplied text, so details of several statistical tests (outlier methods, significance thresholds, mtDNA analyses) are not fully assessable.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Bayesian clustering — STRUCTURE admixture model (MCMC) | Range-wide genetic cluster identification; K = 1–10, 10 replicates each, 10,000 burn-in + 100,000 MCMC iterations, no location prior | 611 individuals, 16,749 loci | not stated |
| Evanno ΔK method (rate of change in log-probability between successive K values) via Structure Harvester | Selection of the most likely number of genetic clusters K across STRUCTURE replicates | 10 replicates per K, K = 1–10 | not stated |
| Mean log-likelihood L(K) | Supplementary K-selection criterion alongside Evanno ΔK (authors note ΔK cannot detect K = 1) | 10 replicates per K | not stated |
| Principal component analysis (PCA) via adegenet v2.1.1 | Genotypic diversity summary across full, neutral-only, and outlier-only SNP datasets; also used to evaluate structure by sex and age | 611 individuals, 16,749 loci (full dataset); subsets also analysed | not stated |
| Discriminant analysis of principal components (DAPC) with BIC-based cluster selection (find.clusters) via adegenet v2.1.1 | Model-free genetic cluster identification as complement to STRUCTURE | 611 individuals, 16,749 loci | not stated |
| Pairwise FST (F-statistics) | Quantification of genetic differentiation among 42 sampling sites and between northern vs. southern cluster assignments (admixed individuals excluded); locus-specific F-statistics also calculated — details truncated in supplied text | n per site: 3–41; total 611 individuals | not stated |
-
The optimal number of clusters K was selected primarily via the Evanno ΔK method; the authors themselves note this method cannot detect K = 1 and used L(K) as a secondary check↳ Could also: Cross-validation (CV) error from ADMIXTURE, or a combination of ΔK, L(K), and CV error reviewed together, could also be used for K selection — ΔK tends to recover the uppermost level of hierarchical structure and can be sensitive to uneven sampling across sites; CV error is an independent, prediction-based criterion that can complement ΔK and L(K), reducing ambiguity when multiple values of K receive similar support
-
Bayesian clustering was performed with STRUCTURE using an MCMC admixture model with no location prior↳ Could also: ADMIXTURE (Alexander et al. 2009), fastSTRUCTURE, or TESS3 (which integrates geographic coordinates) could also be used for ancestry estimation — ADMIXTURE implements an equivalent admixture model but is substantially faster for large SNP panels, enabling more thorough exploration of K and longer MCMC chains; TESS3 additionally accounts for continuous isolation-by-distance, which is relevant given the latitudinal cline and admixed zone observed in this study
-
A single SNP per RADseq locus was retained using a fixed physical-distance threshold (--thin 5000) to limit linkage among markers↳ Could also: Retaining all SNPs followed by empirical LD pruning (e.g., PLINK --indep-pairwise with a sliding window) could also control for linkage disequilibrium — LD-based pruning uses the observed correlation structure among markers rather than a fixed distance, retaining more information in regions of low LD while more thoroughly filtering high-LD regions such as the putative chromosomal inversion identified in this study
-
Genetic differentiation was quantified with FST across 42 sites and between cluster pairs↳ Could also: Hedrick's G'ST or Jost's D could also be used as complementary differentiation measures, particularly for the outlier and fixed loci — FST is bounded by within-population heterozygosity and can underestimate differentiation at highly polymorphic loci; standardized measures such as G'ST or Jost's D are less sensitive to allele frequencies and may better reflect the magnitude of differentiation at the 182 outlier loci where differentiation is strongest
-
Admixed individuals were characterised by near-50:50 STRUCTURE Q-values and excluded from pairwise FST comparisons↳ Could also: Hybrid-class assignment using hybrid index combined with interspecific heterozygosity (HI–Hobs plots, e.g., via the R package introgress or hybriddetective) could also be applied to the contact zone individuals — HI–Hobs approaches distinguish among F1, F2, and backcross hybrid classes rather than treating admixed individuals as a single undifferentiated group, providing additional resolution on the reproductive dynamics at the Point Reyes contact zone that could inform management of the boundary region
-
Population structure analyses (STRUCTURE, PCA, DAPC) were run on the full SNP dataset and then repeated separately on neutral-only and outlier-only subsets↳ Could also: A redundancy analysis (RDA) or partial RDA controlling for geographic distance (isolation by distance) could also partition structure attributable to environment vs. neutral demographic history — Separating datasets post hoc into neutral and outlier loci does not quantify the relative contributions of selection and gene flow to observed structure; RDA-based landscape genomics approaches can simultaneously test environmental associations while controlling for spatial autocorrelation, complementing the outlier-scan approach used here
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Where directly checkable the deposit holds up — 16,749 loci and n=611 are exact and global FIS 0.0024 vs 0.0027 is negligible — so the broad picture of low overall differentiation is corroborated. The main deviations (global FST 0.0115 vs 0.0141, FIT offsets, and the non-recoverable per-locus range −0.0575…0.8896) are best explained as a Weir&Cockerham-vs-Nei estimator drift compounded by the authors depositing only Nei stats; magnitude and direction are preserved. This is partly our forced method choice (Nei was what shipped) and partly an authors-side documentation/deposit gap, not fabrication. The outlier-driven core claim (182 shared outliers, K=2, FST-excl-outliers 0.0009) remains pending regeneration, so confirmation is only partial.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.