Evolutionary Genomics of Sex-Related Chromosomes at the Base of the Green Lineage.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓No relevant deviation in data/preprocessing
- ✓Any deviation was negligible
- 🟡Reported values were only indirectly comparable
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL but clean reproduction. The paper's OWN analysis pipeline (pico-PLAZA orthogroups + phylogenies + Ka/Ks) is NOT shipped; the linked github.com/ropensci/onekp is an archived generic 1KP data-access R package, not the authors' code (P16). We reproduced the headline GENOMIC signature from public chromosome-level assemblies on «our HPC» (SLURM «job», node n094): the sharp GC-content decrease on the candidate MT chromosome 2. Measured low-GC core decreases of 10.82 pts (O. tauri RCC4221) and 13.66 pts (O. lucimarinus CCE9901) vs genome background -- BOTH inside the reported 9-17 pt band -> 1:1 on the GC-decrease signature; chr2 identity exact; genome stats exact vs NCBI. The 450-650 kb MT-locus span graded PARTIAL: it is the full allele (gene content + suppressed recombination), wider than our composition-only deepest-drop core (159/216 kb) -- a subset, not a contradiction. Fresh job reproduced the prior archived run's numbers IDENTICALLY (deterministic). Gene-family counts, MT phylogenies, Ka/Ks, GC3-of-MT-genes, divergence times NOT attempted (require the unshipped custom pipeline / undefined gene sets). METADATA FLAG: paper's Data Availability cites PRJNA337288 as 'O. lucimarinus', but NCBI = Ostreococcus tauri RCC1115 (MT+) -- a species/accession mislabel.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 90assessed: 2026-06-18 ⛓ 5219e85dc0c3
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper investigates whether the candidate MT+ and MT- mating-type alleles in Ostreococcus tauri and related Mamiellales diverged before (trans-specific evolution) or after speciation within the Ostreococcus genus, using gene family, synteny, and phylogenetic analyses.
- ★ The divergence of the MT+ and MT- alleles predates speciation events within the Ostreococcus genus finding
- ★ MT+ and MT- regions of O. tauri show no synteny and cannot be aligned at the nucleotide level, unlike syntenic regions outside the MT locus finding
- ★ A sharp (~9-17%) decrease in GC content on the outlier chromosome was used to define MT locus boundaries in eight Mamiellales genomes finding
- ★ Phylogenetic profiling of MT gametolog trans-specific polymorphisms disclosed candidate MT loci in two additional Mamiellales species, and possibly a third finding
- ★ There is no evidence of discrete evolutionary strata across the MT region of O. tauri finding
- ★ A past large-scale translocation of gene segments between MT+ and MT- significantly improves their colinearity compared with random gene order finding
- OrthoFinder was used to classify MT and non-MT genes into gene families across eight Mamiellales genomes method
- ★ The Mamiellales MT candidates are likely the oldest mating-type loci described to date finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Comparative genomics / GC content analysis | 8 Mamiellales genomes (incl. O. tauri RCC4221 and RCC1115) | none | GC content and chromosome/locus size to define MT boundaries | — |
| Orthology/gene family clustering | Ostreococcus spp., Bathycoccus prasinos, Micromonas commoda, Micromonas pusilla genomes | none | assignment of genes to gene families (GFs) | OrthoFinder |
| Synteny and gene-order (Sdist statistic) analysis | O. tauri RCC4221 (MT-) and RCC1115 (MT+) | none | relative colinearity of orthologous gene positions between MT+ and MT- | — |
| Ka/Ks (synonymous/nonsynonymous substitution rate) analysis | 69 shared MT gene families, O. tauri MT- and MT+ | none | Ks and Ka values per gametolog pair, search for evolutionary strata | — |
| Phylogenetic tree reconstruction of gene families | Mamiellales/Mamiellophyceae genomes (gametologs) | none | tree topology classified as ante- or post-speciation divergence of MT alleles | — |
| Functional/Gene Ontology annotation | MT-specific and core MT genes in O. tauri RCC4221 and RCC1115 | none | predicted gene function/GO terms | — |
- ▼ Sharp decrease in GC content on the outlier chromosome used to define MT boundaries 9-17%
- – 23 core MT gene families present in MT regions of all 8 Mamiellales genomes 23 genes in 23 GFs (both MT- and MT+)
- – MT-specific gene families are few and asymmetric between mating types 6 genes/6 GFs (MT-) vs 2 genes/2 GFs (MT+)
- ▼ Translocation of the [b,c] segment to the 5' start of MT- significantly improves colinearity with MT+ versus random gene order P=0.0054 (100,000 permutations)
- – Overall observed Sdist not significantly different from random gene placement P>0.10 (10,000 permutations)
- – Most Ka/Ks-computable gametolog pairs had Ks<1 and were not adjacent on both MT loci, arguing against discrete evolutionary strata 19/22 gene pairs with Ks<1
- – Total gene content of the two candidate MT loci is similar in size 244 genes (MT-, RCC4221) vs 240 genes (MT+, RCC1115)
- pvalue P=0.0054 (Significance of translocated segment improving MT+/MT- colinearity (permutation test))
- pvalue P>0.10 (Non-significance of overall observed Sdist vs random gene order (permutation test))
- count 244 genes (MT-, RCC4221); 240 genes (MT+, RCC1115) (Total gene count in each candidate MT locus)
- count 23 genes in 23 GFs (core MT GFs) in both MT- and MT+ (Core mating-type gene families shared across all Mamiellales genomes)
- count 75 genes in 69 GFs (MT-); 79 genes in 69 GFs (MT+) (Shared (noncore) MT gene families between the two Ostreococcus MT loci)
- fold_change 9-17% lower GC content (GC content decrease on outlier chromosome defining MT region boundaries)
- other 19/22 gametolog pairs with Ks<1 (Ka/Ks analysis of shared MT gene families used to search for evolutionary strata)
- other up to 640 Myr (Estimated divergence time span across the eight analyzed Mamiellales genomes)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This comparative genomics study used GC content profiling to define mating-type (MT) locus boundaries across eight Mamiellales genomes, OrthoFinder-based gene family clustering to classify gametologs into functional categories, and a custom colinearity statistic (Sdist) evaluated via permutation tests to assess gene-order relationships between MT+ and MT− regions. Synonymous (Ks) and non-synonymous (Ka) substitution rates were computed for shared gene pairs to investigate evolutionary strata, and phylogenetic tree topologies were used to classify gene families by whether allele divergence predated or postdated speciation events within Ostreococcus. Note: the provided text is truncated before the Materials and Methods section, so additional statistical procedures may have been described there.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Permutation test (random shuffling of gene order) on custom Sdist colinearity statistic — overall arrangement | Overall gene-order colinearity between MT+ and MT− regions (fig. 2A) | 10,000 permutations | not stated |
| Permutation test (random shuffling of gene order) on custom Sdist colinearity statistic — translocation-adjusted arrangement | Colinearity improvement after proposed transposition of the 5′ [b,c] segment of MT− (fig. 2A) | 100,000 permutations | not stated |
| Synonymous substitution rate (Ks) and non-synonymous substitution rate (Ka) computation (Tzeng et al. 2004 method) | 69 shared MT gene family pairs between MT+ and MT− in O. tauri; Ka computed for 22 pairs with non-saturated Ks | 22 gene pairs (19 of 22 with Ks < 1) | not stated |
| GC content profiling — threshold-based boundary detection (~9–17% drop relative to genome-wide average) | Defining MT locus boundaries in O. tauri RCC4221 (MT−), O. tauri RCC1115 (MT+), and six Mamiellales genomes | 8 genomes | not stated |
| Phylogenetic tree topology classification (ante- vs. post-speciation allele divergence) | Gene family trees for core, shared, and MT-specific gene families across Mamiellales | — | not stated |
-
MT locus boundaries were defined by visual inspection of a ~9–17% GC content drop on the outlier chromosome↳ Could also: Formal compositional segmentation algorithms (e.g., hidden Markov models for GC shifts, or circular binary segmentation as used in CNV analysis) could also be applied to detect and localize breakpoints — Algorithmic boundary detection provides statistical confidence intervals on breakpoint positions and reduces dependence on manual threshold selection, facilitating comparison across genomes
-
Gene-order colinearity between MT+ and MT− was assessed using a custom Sdist statistic with permutation-derived null distributions↳ Could also: Established synteny-scoring frameworks such as MCScanX, i-ADHoRe, or rank-correlation statistics (e.g., Spearman's ρ on gene positions) could also quantify colinearity — Well-documented tools provide standardized null models, enable multi-genome comparisons in a single framework, and allow independent reproduction of results without re-implementing the statistic
-
Evolutionary strata were investigated by examining the Ks values and adjacency patterns of 22 gene pairs with non-saturated substitution rates↳ Could also: Permutation-based tests of Ks clustering along the chromosome (as in Lahn and Page 1999 and subsequent stratum analyses) or Gaussian mixture models over Ks distributions could also be used to identify discrete strata — Formal stratum detection methods provide statistical support for the number and boundaries of strata rather than relying on patterns of gene adjacency in a small, saturation-filtered subset
-
Phylogenetic tree topologies classifying gene families as ante- or post-speciation divergence were used without reporting bootstrap or Bayesian posterior support thresholds↳ Could also: Reporting bootstrap support (e.g., ≥70%) or Bayesian posterior probabilities (e.g., ≥0.95) per node, or applying concordance factor analyses, could also accompany each topological classification — Explicit support values allow readers to gauge confidence in each classification and quantify how many gene trees are genuinely resolved versus ambiguous, which is important when drawing conclusions about trans-specific evolution
-
GC content differences between MT and non-MT regions were reported as descriptive ranges without a formal statistical comparison↳ Could also: A Wilcoxon rank-sum test or t-test comparing per-gene GC content inside versus outside MT boundaries, with a reported effect size, could also be applied (analogous to the approach in Hamaji et al. 2018 for volvocine algae) — A formal test with effect size and confidence interval would let readers assess the magnitude and uncertainty of the GC composition difference between recombining and non-recombining regions across genomes
-
Two permutation tests were performed on the same data (overall Sdist and translocation-adjusted Sdist) without multiple-testing correction↳ Could also: A Bonferroni correction or a sequential Holm procedure could also be applied to the family of permutation tests, or both tests could be framed as a single composite hypothesis — Correction for the number of tests performed on the same gene-order data controls the family-wise error rate and makes explicit whether the translocation finding would survive adjustment
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34599324
Paper: Benites et al. 2021, Genome Biol Evol 13(12):evab216. "Evolutionary Genomics of Sex-Related Chromosomes at the Base of the Green Lineage." Comparative genomics of candidate mating-type (MT) chromosomes in Mamiellophyceae (picoplanktonic green algae: Ostreococcus, Bathycoccus, Micromonas, ...).
Code artifact reality (P16)
- Linked "code":
https://github.com/ropensci/onekp→ ropensci-archive/onekp, an archived generic R package to access 1000-Plants (1KP) sequence data. It is NOT the authors' analysis pipeline. The paper's actual analysis used a custom, unshipped pipeline: pico-PLAZA framework + BLASTP/OrthoFinder/MAFFT/IQ-TREE/ TransDecoder/seqinr(Ka/Ks). No author code repository, no workflow, no parameter files are deposited. → Authors' end-to-end pipeline is not reproducible (no_code for the orthogroup/phylogeny machinery). - Per brief P16: applying standard third-party tools to the paper's public data is equally valid. So we reproduce the clearly-specified genomic measurements that are derivable from the deposited chromosome-level assemblies with standard methods.
Datasets the paper relies on
| accession | paper label | NCBI reality | role |
|---|---|---|---|
| CAID00000000.1 / GCA_000214015.2 | O. tauri (RCC4221, MT−) | O. tauri RCC4221, 20 chr, 12.9 Mb | reference genome, candidate MT chr2 |
| PRJNA337288 | "O. lucimarinus" (paper) | O. tauri RCC1115 (MT+) — MISLABEL in paper | second reference (MT+) |
| GCA_000092065.1 | O. lucimarinus CCE9901 | O. lucimarinus CCE9901, 21 chr, 13.2 Mb | cross-species MT |
| PRJNA15676 | Micromonas commoda | genome | gene-family input |
| PRJNA15678 | M. pusilla | genome | gene-family input |
| PRJNA394752 | Bathycoccus prasinos | genome | gene-family input |
| PRJNA248394 | MMETSP | 33 transcriptomes | transcriptome panel |
IN SCOPE (pipeline-derived, reproducible from public assemblies)
- R1. GC-content signature of the candidate MT chromosome. Paper: a "sharp (~9–17%) decrease in GC content on the big outlier chromosome was used to define MT boundaries" in O. tauri RCC4221 (chr2) and others (Fig. 1; suppl. Table S1). The MT locus spans 450–650 kb. → Reproduce: per-chromosome + sliding-window GC of chr2; detect the low-GC outlier region; measure its GC decrease vs genome background. Pipeline: sequence composition (standard; no author code needed).
- R2. Genome-level assembly statistics (chromosome count, genome size, genome-wide GC) vs values implied by paper / NCBI. Sanity QC.
OUT OF SCOPE (custom unshipped pipeline / wet-lab / not pinnable)
- Gene-family composition counts (Table 1: 244/240 genes, 23 core MT GFs, etc.) — require the custom pico-PLAZA orthogroup definitions (not shipped).
- Phylogenetic topology of MT GFs ("21 of 23 support ancient MT"), Micromonas MT subclustering — require the per-family alignments/trees (not shipped; gene set undefined without the pipeline output).
- Ka/Ks of 22 gene pairs — requires the specific paired gene set (not shipped).
- GC3 of "core MT GFs" (~20% lower than background) — requires the MT-gene set (defined only by the unshipped pipeline); the which-genes list is not deposited.
- Divergence-time estimates (330–640 Myr) — molecular-clock with manual calibration.
Reproduction strategy
R1 is the headline, clearly-specified, fully public-data signal → primary target. Run sequence-composition analysis on «our HPC» (SLURM), pull small CSV/JSON back, compare.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The headline genomic signature — the sharp ~9-17% GC decrease on the candidate MT chromosome (chr2) — reproduces cleanly from public, identical chromosome-level assemblies: 10.82 pts (O. tauri) and 13.66 pts (O. lucimarinus), both inside the reported band, with genome stats exact. The limiting problems are authors-side: the linked GitHub repo is a generic 1KP data-access package rather than their pico-PLAZA pipeline, so gene-family/phylogeny/Ka-Ks claims are not derivable, and PRJNA337288 is mislabeled (NCBI = O. tauri RCC1115, not O. lucimarinus). No fabrication concern — the tested numbers match within tolerance; the deviation is one of completeness/availability, not value discrepancy, so overall yellow with a confirmed-but-partial central claim.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.