Whole genome and transcriptome maps of the entirely black native Korean chicken breed Yeonsan Ogye.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡A deviation arose in the data or preprocessing
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the QC, mostly 1:1. The registry 'code' (github.com/sohnjangil/tsrator = TSRATOR/ClusToR) is a 27 KB C++ helper for ONE step of a large hybrid WGS assembly; no reported number ties 1:1 to its output and running it needs the whole assembly stack, so per 80/20 it was NOT run and the full de-novo re-assembly + gene annotation were left out of scope. Instead I reproduced the pipeline-derived QC of the deposited assembly (GCA_002798355.1 Ogye1.0 / WGS PDMY01) with standard third-party tools on «our HPC» («job», 14m42s). RESULTS: assembly contiguity recomputed EXACTLY against the deposited record (total length 1,021,022,236 bp exact; scaffold N50 90,110,113 bp exact; gap-aware contigs 7722 vs 7721 and contig N50 639 kbp exact; GC 41.56% vs 41.5%; 1822 vs 1821 scaffolds), and BUSCO genome completeness reproduced WITHIN TOLERANCE (Complete 97.0% vs reported 97.60%; Duplicated 0.5% EXACT; Fragmented 0.8% vs 0.90%; Missing 2.3% vs 1.00%). The completeness rerun used vertebrata_odb10 + miniprot because the paper's exact vertebrata_odb9 lineage is no longer obtainable (all archive URLs 404) and current BUSCO v6 refuses odb9 — so C1-C4 are an honest corroboration, not a byte-exact same-lineage match (not chased, 80/20). One apparent mismatch is fully explained, not fabrication: the paper's 16.8 Mbp scaffold N50 is the pre-chromosome-anchoring scaffold stage, while the deposited GCA was anchored one step further (90.1 Mbp). NOT attempted: full hybrid re-assembly, TSRATOR/ClusToR run, MAKER annotation (15,766 genes), transcriptome/lncRNA/pigmentation analyses. No fabrication concern observed; all checked figures are consistent with independent recomputation on the deposited data.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 91assessed: 2026-06-14 ⛓ 5c49da49fd86
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study aims to generate the first whole-genome assembly and comprehensive coding/noncoding transcriptome and DNA methylation maps for the entirely black indigenous Korean chicken breed Yeonsan Ogye (YO), to provide resources for understanding breed-specific phenotypes, particularly hyperpigmentation linked to the fibromelanosis (FM) locus.
- ★ A hybrid de novo assembly combining high-depth Illumina short reads (376.6X) and low-depth PacBio long reads (9.7X) produced the YO draft genome Ogye_1.1 with contig and scaffold NG50 of 362.3 Kbp and 16.8 Mbp. method
- ★ The Ogye_1.1 draft genome has 97.6% BUSCO completeness, comparable to galGal5 (97.4%) and superior to other avian genomes (92%-93%). finding
- ★ Compared to galGal4 and galGal5, the YO genome contains 551 structural variations, including a duplication of the fibromelanosis (FM) locus related to hyperpigmentation. finding
- ★ Transcriptome maps reconstructed from 20 tissues include 15,766 protein-coding and 6,900 long noncoding RNA genes, many tissue-specifically expressed with tissue-specific promoter DNA methylation. resource
- ★ The FM locus rearrangement in YO is best explained by a single-step inverted duplication (rearrangement 1), supported by scaffolds discontinuous on both sides of the duplicated regions. mechanism
- The genome and transcriptome maps serve as resources for studying domestic/black-skinned chicken breeds, breed genomic differences, and the evolution of hyperpigmentation. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-genome sequencing (short read) | Yeonsan Ogye chicken (Gallus gallus domesticus), blood-derived genomic DNA | none | genome sequence/coverage (376.6X) for de novo assembly | Illumina HiSeq 2000; Wizard DNA extraction kit |
| Whole-genome sequencing (long read) | Yeonsan Ogye chicken, genomic DNA | none | long reads (9.7X, avg 6 Kbp) for gap-filling/scaffolding | PacBio RS II with P6C4 chemistry |
| RNA sequencing (RNA-seq) | 20 YO tissues (breast, liver, bone marrow, fascia, cerebrum, gizzard, mature/immature eggs, comb, spleen, cerebellum, gallbladder, kidney, heart, uterus, pancreas, lung, skin, eye, shank) | none | transcript expression / gene counts (~1.5 billion reads) | Illumina TruSeq Stranded Total RNA Sample Prep Kit; Ribo-Zero rRNA depletion |
| Reduced representation bisulfite sequencing (RRBS) | 20 YO tissues, MspI-digested genomic DNA | none | DNA methylation levels (~123 million reads) | Illumina HiSeq 2500; TruSeq Nano DNA Library Prep Kit; EpiTect Bisulfite Kit |
| Structural variation detection / comparative genome alignment | Ogye_1.1 genome vs galGal4 and galGal5 | none | large SVs (>1 Kbp) counts and types | LASTZ; Delly, Lumpy, FermiKit, novoBreak |
| FM locus read-depth and discordant-read mapping analysis | YO FM locus on chromosome 20 vs galGal4 | none | read depth and paired-end/mate-pair mapping pattern to infer duplication/rearrangement | — |
| Genome completeness assessment | Ogye_1.1 vs galGal4, galGal5, turkey, duck, zebra finch | none | complete single-copy / duplicated / fragmented / missing ortholog percentages (2,586 conserved vertebrate genes) | BUSCO with OrthoDB v9 |
- – Hybrid assembly yielded contig and scaffold NG50 of 362.3 Kbp and 16.8 Mbp (final contig/scaffold N50 506.3 Kbp and 21.2 Mbp) 362.3 Kbp / 16.8 Mbp NG50
- ▲ Ogye_1.1 BUSCO complete single-copy genes 97.60%, exceeding galGal4 (96.90%) and galGal5 (97.40%) 97.60% vs 97.40%/96.90%
- – 551 structural variations detected including 185 deletions, 180 insertions, 158 duplications, 23 inversions, 5 translocations 551 SVs
- – 290 and 447 distinct SVs detected relative to galGal4 and galGal5, respectively 290 / 447
- – Transcriptome maps include 15,766 protein-coding and 6,900 lncRNA genes across 20 tissues 15,766 / 6,900 genes
- ▲ FM locus showed doubled read depth at two loci, indicating duplication; intervening region estimated at 412.6 Kbp; Gap_1 and Gap_2 estimated at 164.5 Kbp and 63.3 Kbp 412.6 Kbp intervening; 164.5/63.3 Kbp gaps
- – Gap percentage and contig N50 improved from 1.87%/53.6 Kbp (initial) to 0.85%/506.3 Kbp (final) 1.87%→0.85%; 53.6→506.3 Kbp
- other 376.6X raw Illumina short reads (100.2X small insert + 276.4X large insert) (whole-genome short-read coverage)
- other 9.7X PacBio long reads, average length 6 Kbp (long-read coverage for gap filling)
- other 97.60% complete single-copy BUSCO (Ogye_1.1 completeness over 2,586 conserved vertebrate genes)
- count 551 structural variations (>1 Kbp) (SVs vs galGal4/galGal5)
- count 15,766 protein-coding and 6,900 lncRNA genes (transcriptome maps from 20 tissues)
- count ~1.5 billion RNA-seq reads; ~123 million RRBS reads (sequencing depth across 20 tissues)
- other contig NG50 362.3 Kbp, scaffold NG50 16.8 Mbp (final Ogye_1.1 assembly using estimated genome size 1.25 Gbp)
- other FM intervening region 412.6 Kbp; Gap_1 164.5 Kbp, Gap_2 63.3 Kbp (FM locus duplication structure on chromosome 20)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a genome/transcriptome Data Note that primarily reports descriptive bioinformatic metrics rather than inferential hypothesis testing. The overall approach centers on de novo hybrid genome assembly from a single bird, with assembly quality summarized via N50/NG50 length statistics and BUSCO single-copy ortholog completeness percentages, and with structural variants called by multiple programs and retained when validated by at least one of them. Transcriptome and methylation data from 20 tissues are summarized as read counts and mapping rates, and copy-number evidence for the FM-locus duplication is presented as observed read-depth doubling rather than a formal statistical test.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Descriptive assembly metrics (contig/scaffold N50 and NG50, gap percentage) | Hybrid whole-genome assembly evaluation (Fig. 1B/1C, Supplementary Fig. S1) | single individual (1 YO chicken, object no. 02127) | na |
| BUSCO single-copy ortholog completeness assessment | Genome completeness comparison across assemblies (Table 4), 2,586 conserved vertebrate genes / OrthoDB v9 | 2,586 conserved vertebrate genes; one assembly per species | na |
| Structural-variant calling with multi-caller consensus (validated by ≥1 of Delly, Lumpy, FermiKit, novoBreak) | Large SV detection vs galGal4 and galGal5 (Supplementary Fig. S5, Table S2) | single genome compared to two reference assemblies | na |
| Read-depth and discordant paired-end/mate-pair mapping inspection (copy-number / rearrangement inference) | FM-locus duplication and rearrangement scenarios (Fig. 2A–2C) | single individual | na |
-
Assembly completeness was summarized with BUSCO single-copy ortholog percentages and N50/NG50 length statistics.↳ Could also: Complementary reference-free metrics such as k-mer-based completeness/quality (e.g., Merqury via k-mer spectra) or LAI/assembly-consistency scores could also be reported. — K-mer-based and consistency metrics provide an orthogonal, reference-independent view of base-level accuracy and completeness alongside ortholog-based scores.
-
Structural variants were retained when validated by at least one of four SV callers.↳ Could also: A consensus or majority-vote scheme (e.g., requiring support from two or more callers, or merging with a tool such as SURVIVOR) could also be used. — A higher consensus threshold trades sensitivity for specificity and lets one report how call confidence varies with the number of supporting callers.
-
The FM-locus duplication was inferred from observed doubled read depth and discordant read mapping.↳ Could also: A formal read-depth copy-number model (e.g., a statistical CNV caller with a baseline/expected-coverage distribution) could also quantify the duplication. — A modeled approach would attach an explicit confidence or copy-number estimate to the observed depth change rather than relying on a qualitative doubling.
-
Quantities such as mapping rates and assembly metrics are reported as single point values from one individual.↳ Could also: Where biological generality is of interest, additional individuals or technical replicates with reported dispersion (SD/CI) could also be incorporated. — Replication with a dispersion measure would let readers gauge variability and the breed-level generalizability of estimates beyond the single reference bird, consistent with the paper's stated aim as a single-individual reference resource.
-
Genome comparisons (Table 4) are presented as descriptive percentage differences between assemblies.↳ Could also: When comparing categorical completeness counts, a contingency-table summary (e.g., counts of complete/fragmented/missing genes) could also accompany the percentages. — Presenting underlying counts alongside percentages makes the basis of small inter-assembly differences fully transparent to the reader.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Ogye_1.1 assembly achieves 97.60% BUSCO complete single-copy orthologs, exceeding galGal4 (96.90%) and galGal5 (97.40%)WGS chicken yeonsan-ogye up 2018×1papers★ This paper is the founder (earliest)
-
FM locus on chromosome 20 shows doubled read depth at two sub-loci indicating tandem duplication, with intervening region estimated at 412.6 Kbp and flanking gaps of 164.5 Kbp and 63.3 KbpWGS chicken yeonsan-ogye up 2018×1papers★ This paper is the founder (earliest)
-
551 structural variations detected in Yeonsan Ogye genome: 185 deletions, 180 insertions, 158 duplications, 23 inversions, and 5 translocationsWGS chicken yeonsan-ogye 2018×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-30010758
Paper: Sohn et al. (2018) Whole genome and transcriptome maps of the entirely black native Korean chicken breed Yeonsan Ogye. GigaScience 7(7):giy086. DOI 10.1093/gigascience/giy086 · PMID 30010758 · PMCID PMC6065499.
Code link in registry: https://github.com/sohnjangil/tsrator → the repo is TSRATOR / ClusToR, a small (~27 KB) C++ utility: "a clustering tool for reference-assisted whole-genome de novo assembly" (clusters scaffolds + reads by reference chromosome to simplify the assembly graph). It is one helper step inside the paper's large hybrid-assembly pipeline (ALLPATHS-LG → SSPACE-LongRead → GapCloser → OPERA → PBJelly → pseudo-reference-assisted step using TSRATOR). Per the brief (P16) a third-party tool on the paper's own data counts as a valid reproduction — but TSRATOR emits intermediate split-FASTA files, and the paper reports no specific number tied directly to TSRATOR's output that can be compared 1:1. Re-running TSRATOR would also require the unsplit scaffolds + the full lastz/BWA mapping stack (the whole ~1 Gb hybrid assembly), i.e. the entire 100 %. That is out of scope (HARD RULE 3, 80/20).
Pipeline-derived results in the paper (what could be reproduced)
| Result | Pipeline | Reproducible from public data? | In scope |
|---|---|---|---|
| Hybrid genome assembly (the 1.02 Gb assembly itself) | ALLPATHS-LG + SSPACE + OPERA + PBJelly + TSRATOR | Needs raw PacBio+Illumina+Fosmid + days of compute = the full 100 % | No (80/20 drop of the hard 20→100 %) |
| Assembly completeness — BUSCO (Table 4): Ogye_1.1 = 97.60 % complete, 0.50 % duplicated, 0.90 % fragmented, 1.00 % missing, vertebrata_odb9 (2 586 genes) | BUSCO genome mode on the final assembly | Yes — deposited assembly FASTA is public (GCA_002798355 / WGS PDMY01); BUSCO is a standard third-party tool, gene-content QC is invariant to scaffold-vs-chromosome stage | YES — primary |
| Assembly contiguity (text/Fig 1C): scaffold N50 16.8 Mbp, contig NG50 362.3 Kbp, pseudo-contig N50 506.3 Kbp, gap 0.85 %; total ≈1.02 Gb | direct sequence statistics | Yes — recompute total length, #scaffolds/#contigs, N50, GC from the deposited FASTA | YES — secondary (with a stage caveat, see below) |
| GC content 41.5 % (NCBI metadata) | sequence stat | Yes — recompute | secondary |
| 15,766 protein-coding genes | MAKER/AUGUSTUS annotation pipeline | Needs full re-annotation (RNA-seq + ab initio) = hard 20 % | No (80/20) |
| Transcriptome / lncRNA / pigmentation (EDN3, fibromelanosis) findings | RNA-seq DE, manual curation | Mostly wet-lab-anchored / large RNA-seq pipeline | No |
Chosen in-scope targets (clear, low-hanging, deterministic)
- BUSCO completeness of the deposited Ogye assembly (primary, 1:1 vs Table 4). Uses the same lineage the paper used — vertebrata_odb9 (2 586 BUSCOs) so the denominator is identical and the % values are directly comparable. Gene predictor may differ from the paper's (modern BUSCO default vs the paper's 2017 BUSCO/Augustus) — noted as a methodological variation, not a different metric.
- Assembly contiguity statistics recomputed from the deposited FASTA (secondary): total length, #scaffolds, #contigs, scaffold N50, contig N50, GC.
Important auditing caveat (recorded up front)
The deposited assembly GCA_002798355 "Ogye1.0" is chromosome-level (30 chromosomes anchored): NCBI metadata gives scaffold N50 = 90.1 Mbp, contig N50 = 639.8 Kbp, 1,821 scaffolds, 7,721 contigs, 1,021,005,462 bp, 41.5 % GC. The paper instead quotes the pre-anchoring scaffold-stage figure (16.8 Mbp scaffold N50). So the contig-level numbers and total length should match the deposited file, but the scaffold N50 will legitimately differ because the deposited assembly was anchored to chromosomes one step further than the number printed in the text. This is a stage difference, not evidence of fabrication — flagged
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Pipeline-derived QC of the deposited Ogye1.0 assembly reproduces cleanly. All deterministic contiguity statistics are EXACT against the deposited record — total length 1,021,022,236 bp, scaffold N50 90,110,113 bp, contig N50 639,813 bp, GC 41.56%, with contigs/scaffolds off by +1 (definition rounding). BUSCO completeness corroborates within tolerance (Complete 97.0% vs 97.60%; Duplicated 0.5% exact; Fragmented 0.8% vs 0.9%), the larger Missing gap (2.3% vs 1.0%) being a forced lineage-version artifact (odb9 unobtainable -> odb10+miniprot). The lone apparent mismatch — scaffold N50 16.8 Mbp (paper) vs 90.1 Mbp (deposited) — is a legitimate assembly-stage difference (pre- vs post-chromosome-anchoring), not fabrication. Full de-novo re-assembly/annotation were out of scope.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.