Gap-free telomere-to-telomere haplotype assembly of the tomato hind (Cephalopholis sonnerati).
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
FINAL. Sci Data data-descriptor of two gap-free telomere-to-telomere haplotype assemblies of Cephalopholis sonnerati (HA GCA_043388425.1, HB GCA_043388385.1), both PUBLIC at NCBI GenBank. Reproduced via P16 (standard third-party QC tools run FRESH on the paper's own deposited FASTAs) on «our HPC» SLURM. RESULT: 8/10 claims EXACT to the base pair — per-haplotype total length, 24 chromosomes, contig N50, and gap-free (zero N) for both haplotypes (seqkit v2.13.0); #chromosomes independently re-confirmed. 2/10 PARTIAL — BUSCO completeness (actinopterygii_odb10): reproduced HA 99.3% (C:3616) / HB 99.4% (C:3618) with BUSCO 5.7.1/miniprot vs the paper's 97.9% / 97.8% with BUSCO 5.3.0; the ~+1.5pp difference is the well-understood drift from the BUSCO version / gene-predictor change, and both agree the assemblies are near-complete. No fabrication indicators: every contiguity value matches the deposit exactly and BUSCO corroborates high completeness. NOT attempted (out of scope, non-deterministic/raw-read-heavy): full de novo re-assembly, Merqury QV, mapping/coverage, repeat content, gene annotation. All compute ran as SLURM jobs on compute nodes (login-node compute forbidden); a transient account-wide «infra»+HOME quota outage was worked around using node-local /tmp scratch.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 90assessed: 2026-06-19 ⛓ a3a67cd7831a
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · probe · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study aims to determine whether integrating NGS, ONT ultra-long, PacBio HiFi, and Hi-C sequencing data can produce telomere-to-telomere (T2T), gap-free haplotype-resolved genome assemblies of the tomato hind (Cephalopholis sonnerati) that improve upon the existing reference genome in continuity and completeness.
- ★ Two T2T gap-free haplotype assemblies of C. sonnerati (YSFRI_Csonn_HA_1.0 and YSFRI_Csonn_HB_1.0) were successfully generated, each spanning 24 chromosomes with no gaps. finding
- ★ NGS, ONT ultra-long, and PacBio HiFi reads mapped to both assemblies at rates exceeding 99.8%, with ≥20X coverage over more than 99.74% of the genome. finding
- ★ Merqury quality value (QV) analysis indicates high assembly accuracy, with average QVs of 51.80 and 51.83 for the two haplotypes. finding
- ★ BUSCO analysis against the Actinopterygii database found 97.9% and 97.8% complete single-copy orthologs in the two assemblies, indicating high completeness. finding
- ★ 23,270 and 23,184 protein-coding genes were predicted and annotated in the HA and HB haplotype assemblies, respectively. finding
- ★ Telomere identification, centromere prediction, and repetitive sequence annotation were performed for both haplotype assemblies. method
- ★ Compared with the previously published reference genome (JJU_Cson_1.0), the two new assemblies show significantly improved completeness and continuity (e.g., contig N50 of 43.83–44.09 Mb vs 2.48 Mb). finding
- Sequencing data, genome assemblies, and annotation files were deposited in CNGB Sequence Archive, NCBI SRA, NCBI GenBank, and FigShare as public resources. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Next-generation sequencing (NGS, whole-genome shotgun) | Cephalopholis sonnerati, muscle tissue (juvenile, 3 months old) | none | Short-read genomic DNA sequence data for assembly polishing/validation | DNBSEQ-T7 (MGI Tech) |
| PacBio HiFi sequencing | Cephalopholis sonnerati, muscle tissue | none | Long high-fidelity reads for contig assembly | Revio system (PacBio) |
| ONT ultra-long sequencing | Cephalopholis sonnerati, muscle tissue | none | Ultra-long reads for scaffolding, telomere/gap patching | PromethION (Oxford Nanopore Technologies) |
| Hi-C sequencing | Cephalopholis sonnerati, liver tissue | none | Chromatin interaction data for chromosome clustering/ordering/orientation | DNBSEQ-T7 (MGI Tech) |
| RNA sequencing (NGS + ONT full-length transcriptome) | Cephalopholis sonnerati, gill and fin tissues | none | Transcript assemblies for gene structure prediction | DNBSEQ-T7 and PromethION |
| BUSCO completeness assessment | In silico analysis of assembled genome (Actinopterygii odb10 database) | none | Percentage of complete, fragmented, and missing conserved single-copy orthologs | BUSCO v5.3.0 |
| Merqury k-mer based quality value (QV) analysis | In silico analysis of assembled genome vs sequencing reads | none | Assembly quality value (accuracy/correctness) | Merqury v1.3 |
| Repeat annotation (de novo + homology-based) | In silico analysis of assembled genome | none | Proportion and length of interspersed and tandem repeats (DNA transposons, LINEs, SINEs, LTRs) | RepeatModeler v2.0.4, RepeatMasker v4.1.5 |
- – Assembled two T2T gap-free haplotypes: YSFRI_Csonn_HA_1.0 (1,039.53 Mb, contig N50 43.83 Mb) and YSFRI_Csonn_HB_1.0 (1,039.91 Mb, contig N50 44.09 Mb), each with 24 chromosomes
- ▲ Mapping rates of NGS, ONT ultra-long, and PacBio HiFi reads to both assemblies >99.8%
- ▲ Genome coverage at ≥20X depth across both assemblies >99.74%
- ▲ Average Merqury quality values for HA and HB assemblies QV 51.80 / 51.83
- ▲ Complete BUSCOs identified out of 3,640 searched genes (Actinopterygii odb10) 97.9% (3,564) / 97.8% (3,599)
- – Protein-coding genes predicted in HA and HB assemblies 23,270 / 23,184 genes
- – Interspersed repeat content identified in HA and HB assemblies 484.33 Mb (46.59%) / 488.17 Mb (46.94%)
- ▲ Contig N50 of new assemblies vastly exceeds that of prior reference genome JJU_Cson_1.0 (2.48 Mb) ~17.7-fold (HA) / ~17.8-fold (HB)
- count 1,039,525,268 bp (HA) / 1,039,913,711 bp (HB) (Total genome assembly length)
- fold_change Contig N50 43,828,159 bp (HA) / 44,092,975 bp (HB) vs 2,482,587 bp reference (Assembly contiguity comparison to JJU_Cson_1.0)
- mean QV 51.80 (HA) / 51.83 (HB) (Merqury assembly accuracy)
- count 3,564/3,640 (97.9%) and 3,599/3,640 (97.8%) complete BUSCOs (Genome completeness assessment)
- count 23,270 (HA) / 23,184 (HB) (Number of predicted protein-coding genes)
- other >99.8% mapping rate; >99.74% coverage at ≥20X (Read alignment quality to assemblies)
- other 407.54 Gb NGS, 98.84 Gb PacBio HiFi, 61.43 Gb ONT ultra-long, 311.60 Gb Hi-C (Total sequencing data generated for assembly)
- other 46.59% (484.33 Mb) HA and 46.94% (488.17 Mb) HB interspersed repeats; 5.54%/5.59% tandem repeats (Repeat sequence content of genome)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a data-descriptor paper presenting two telomere-to-telomere haplotype genome assemblies for a single tomato hind (Cephalopholis sonnerati) individual, generated from NGS, ONT ultra-long, PacBio HiFi, and Hi-C sequencing. The paper reports assembly quality using standard genomics benchmarking metrics (mapping/coverage rates, BUSCO completeness percentages, Merqury quality-value scores, N50) rather than inferential hypothesis-testing statistics. Results are presented as point-estimate summary tables and figures comparing the two haplotype assemblies to each other and to a previously published reference genome, without significance testing, dispersion measures, or p-values.
-
Assembly correctness was summarized as a single Merqury quality-value (QV) point estimate (e.g., 51.80 and 51.83) for each haplotype, without an accompanying confidence interval or variance measure.↳ Could also: Reporting a bootstrap or per-chromosome distribution of QV alongside the genome-wide average would also be a standard approach. — This would convey how consistent assembly accuracy is across regions of the genome, rather than only a single summary number.
-
Genome completeness was assessed primarily through BUSCO single-copy ortholog percentages against the Actinopterygii database.↳ Could also: Supplementing this with additional independent completeness metrics, such as the k-mer-based Merqury completeness score or the LTR Assembly Index (LAI), could also be reported. — Cross-validating completeness from multiple independent angles (ortholog-based and k-mer-based) can provide a fuller picture of assembly quality than a single metric.
-
The reference assembly was constructed from sequencing a single juvenile individual.↳ Could also: An alternative design could sequence and assemble genomes from multiple individuals, or use trio-binning with parental short-read data, to validate haplotype phasing. — This would help distinguish genome features that are representative of the species from those that may be specific to the single sampled individual.
-
Mapping rates and coverage were reported as single genome-wide percentages (e.g., >99.8% mapping, >99.74% coverage at 20X).↳ Could also: Reporting the distribution of these values per chromosome or in non-overlapping windows (e.g., as a range or interquartile range) could also be presented. — This would highlight whether coverage and mapping quality are uniform across the genome or vary in specific regions, such as repeat-rich areas.
-
Centromere locations were inferred using the top nine most abundant tandem repeat clusters (TRCs) combined with gene density, without a formal statistical threshold for cluster inclusion.↳ Could also: A formal enrichment or clustering statistic (e.g., a threshold based on repeat density distribution) could also be used to define candidate centromeric boundaries. — This would provide a quantifiable confidence measure for centromere calls, complementing the descriptive TRC-ranking approach, which the authors themselves note would benefit from further validation via ChIP-seq or FISH.
-
Comparisons between the two haplotype assemblies (HA vs. HB) and against the prior reference (JJU_Cson_1.0) were made via side-by-side summary tables of metrics like BUSCO percentages and N50.↳ Could also: Where appropriate, paired or distributional comparisons (e.g., per-chromosome length or synteny concordance statistics) could also be used to characterize similarity between the two haplotypes. — This could offer a more granular, quantitative view of haplotype concordance beyond aggregate summary statistics.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-39578472
Paper: Lu et al. 2024, Sci Data. "Gap-free telomere-to-telomere haplotype assembly of the tomato hind (Cephalopholis sonnerati)." DOI 10.1038/s41597-024-04093-3 · PMCID PMC11584678.
This is a Data Descriptor: two haplotype-resolved, gap-free T2T chromosome assemblies of a grouper fish, plus annotation. Two assemblies deposited at NCBI GenBank:
- HA =
GCA_043388425.1(ASM4338842v1, submitter name YSFRI_Csonn_HA_1.0) - HB =
GCA_043388385.1(ASM4338838v1, submitter name YSFRI_Csonn_HB_1.0)
Raw reads: NGS SRR30963279 · PacBio HiFi SRR30963278 · ONT ultra-long SRR30963277 · Hi-C SRR30963276 (the brief's nominal accession) · CNGB CNP0005738. Both assemblies+annotation also on figshare 10.6084/m9.figshare.27300720.
Reproduction strategy (P16 — third-party tools on the paper's own deposited data)
A full de novo T2T re-assembly from raw reads (98 Gb HiFi + 61 Gb ONT + 311 Gb
Hi-C, hifiasm/nextDenovo/ALLHiC/3D-DNA/medaka...) is enormous and non-deterministic;
not attempted. Instead we recompute the assembly-QC layer — the clearly-specified
contiguity + completeness statistics — directly on the deposited GenBank FASTAs,
using standard third-party tools (seqkit, BUSCO5). This is the same, equally
valid reproduction pattern used for prior genome RUs (oilpalm pmid-38918881, goat
pmid-34507524, chlorops pmid-36028584).
IN SCOPE (gradeable, cheap → run first)
Computed from each deposited FASTA:
- Total assembly length (bp) —
seqkit stats. Reported HA 1,039,525,268 / HB 1,039,913,711. Pipeline: hifiasm+nextDenovo+ALLHiC+3D-DNA+medaka (final FASTA). - Number of chromosomes (24) — count of chromosome-scale sequences.
- Contig N50 — N50 over sequences; for a gap-free assembly contig N50 == scaffold/sequence N50. Reported HA 43,828,159 / HB 44,092,975.
- Gap-free — total ambiguous (N) bases == 0, verifying the central "gap-free T2T" claim. (paper: gap-free)
- BUSCO completeness (Actinopterygii lineage) —
busco -m genome -l actinopterygii_odb10. Reported Complete HA 97.9% (3,564) / HB 97.8%; single-copy 96.8% / 96.7%. Tool: BUSCO v5.3.0 (paper) ≈ v5.x (ours).
OUT OF SCOPE (not attempted — heavy / needs raw reads / non-deterministic)
- Merqury QV (HA 51.80 / HB 51.83) — k-mer based, needs the raw HiFi reads (98 Gb download + meryl). Optional last-20%; skipped for cost.
- Mapping rate >99.88% / ≥20X coverage >99.74% — needs full read mapping (bwa/ minimap2) of NGS+HiFi to the assembly. Skipped (raw-read heavy).
- Repeat content (interspersed 46.59%/46.94%, tandem 5.54%/5.59%) — full RepeatModeler/RepeatMasker de-novo pipeline; heavy + library-dependent. Skipped.
- Gene annotation (23,270 / 23,184 protein-coding genes; exon/intron stats; functional-annotation %) — full RNA-seq + MAKER/AUGUSTUS pipeline. Skipped.
- Telomere / centromere counts — Winnowmap/TRF pipeline; not gradeable to a single robust printed number we can cheaply check. Skipped (gap-free N-count above is the cheap proxy for the T2T claim).
Notes
code_urlin registry = nanoporetech/medaka = the ONT polishing tool used in gap-filling; it is one third-party tool in the pipeline, not an authors' repo. We reproduce the outputs on deposited data rather than re-running medaka.- All grades are PROVISIONAL; a human reviewer decides match. BUSCO % can drift across BUSCO/lineage-dataset versions (paper used v5.3.0; we pin a v5.x).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.