Chromosome-level genome assembly of agar-producing red seaweed Gracilaria vermiculophylla.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡A deviation arose in the data or preprocessing
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough and reproduced ~1:1. Approach: ran third-party tools (seqkit, BUSCO v5.3.1) on the paper's OWN deposited final assembly + annotation (figshare 10.6084/m9.figshare.28667702.v3, CC BY 4.0) on «our HPC»/«infra». 8/9 reported metrics reproduce exactly or within tolerance: genome size 77,513,337 bp, scaffold N50 3,164,947 bp, largest contig 4,532,714 bp, anchored 95.19%, 22 chromosomes, and 10,689 genes are all bit-exact; GC 50.45% (=50.5%); genome BUSCO within-tol (D 1.6% / M 16.5% / n=255 identical, Complete +0.8pp = 193 vs 191). Only the gene-set (protein) BUSCO is partial: C 74.9% vs reported 77.3% (~2.4pp lower; duplicated 0.4% matches), likely a BUSCO mode/version difference. NOT attempted (out of scope, 80/20): full de-novo assembly (NextDenovo/Racon/Pilon/Juicer/3D-DNA), repeat/TE annotation (59.26%), functional-annotation DB percentages, and read-level QV/Merqury/coverage metrics (would require the ~55 Gb raw reads). No fabrication concern — every reproduced value derives cleanly from the shipped figshare data. Note: the manifest's data accession SRR23609112 was a text-mining mis-resolution; the paper's real data is SRA SRP564194 (SRR32361123/124/125) + the figshare deposit.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 91assessed: 2026-06-16 ⛓ 903c995a2f0d
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study addresses the lack of a high-quality chromosome-level genome for the agar-producing red seaweed Gracilaria vermiculophylla, aiming to generate one as a reference resource for agar biosynthesis, evolutionary, comparative genomic, and ecological research.
- ★ A chromosome-level genome assembly of G. vermiculophylla was generated by combining DNBSeq short reads, Nanopore long reads, and Hi-C data. resource
- ★ The assembled genome is 77.5 Mb with contig N50 of 2.61 Mb and scaffold N50 of 3.16 Mb, comprising 22 pseudochromosomes. finding
- ★ Transposable elements constitute 45.93 Mb (59.26%) of the genome, with LTRs the predominant retrotransposons (55.03%). finding
- ★ The genome contains 10,689 protein-coding genes, of which 86.14% were functionally annotated. finding
- ★ Hi-C interaction data were used to decontaminate the assembly by distinguishing host chromosomal contigs from bacterial sequences. method
- ★ The new assembly markedly improves contiguity over the previous G. vermiculophylla draft (size 45.9→77.5 Mb; contig N50 30.15 Kb→2.61 Mb, an 86.5-fold increase). finding
- ★ BUSCO, GC content, and sequencing depth assessments demonstrate the high quality of the assembly and successful decontamination. finding
- Chr07 aligns to the sex-determination region; 21 of 26 published sex-linked markers mapped, with seven aligning specifically to Chr07. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-genome short-read sequencing (DNBSEQ WGS) | Gracilaria vermiculophylla thallus | none | short-read genomic sequence (PE150) | MGI-T7 / DNBSEQ platform; Covaris E220 ultrasonicator |
| Nanopore long-read sequencing | Gracilaria vermiculophylla thallus | none | long-read genomic sequence | Nanopore CycloneSEQ; motor protein BCH-X |
| Hi-C (chromatin conformation capture) sequencing | Gracilaria vermiculophylla fresh thalli | none | chromatin contact reads for chromosome anchoring and decontamination | DNBSEQ-T7 (PE150); MboI/DpnII (NEB), T4 DNA Ligase |
| HMW DNA extraction | Gracilaria vermiculophylla (~3 g fresh sample) | none | DNA quality (OD260/280=1.93, OD260/230=1.87) | phenol-chloroform/CTAB protocol |
| Transcriptome-based gene annotation (RNA-seq mapping) | Gracilaria vermiculophylla genome | none | aligned/assembled transcripts for gene model prediction | public SRA RNA-seq data; Hisat2, StringTie |
| BUSCO completeness evaluation | Gracilaria vermiculophylla genome and gene set | none | % complete/single-copy/duplicated/fragmented/missing BUSCOs | BUSCO v5.3.1, eukaryota_odb10 (255 genes) |
| Repeat/TE annotation | Gracilaria vermiculophylla genome | none | TE content and classification | TRF, RepeatMasker, RepeatModeler, LTR_FINDER, RepBase |
| Functional annotation | Gracilaria vermiculophylla protein-coding genes | none | % genes annotated across databases | NR, Swiss-Prot, KEGG, KOG, TrEMBL, InterPro, GO |
- – Final nuclear genome assembled at 77.5 Mb with contig N50 2.61 Mb and scaffold N50 3.16 Mb; 22 pseudochromosomes totaling 73.78 Mb (95.2% anchored) 77.5 Mb; N50 2.61/3.16 Mb
- – Transposable elements account for 45.93 Mb (59.26%); LTRs 55.03% with Gypsy 36.8% and Copia 12.8% 45.93 Mb / 59.26%
- – 10,689 protein-coding genes predicted; 9,207 (86.14%) functionally annotated 10,689 genes; 86.14%
- ▲ Contig N50 improved from 30.15 Kb to 2.61 Mb vs prior draft; contig count dropped from 4,039 to 80 86.5-fold (contig N50); ~50.5-fold (contig count)
- – Genome BUSCO completeness 74.9% complete; gene set 77.3% complete BUSCOs C:74.9% (genome), C:77.3% (gene)
- – Short-read QV of 34.45 (>99.96% accuracy) and 95.65% completeness via Merqury; average short-read depth 62.6 QV 34.45; depth 62.6
- – Nanopore reads mapped at 70.93% with 99.88% coverage and 66.92 average depth; GC clustered tightly around 50.5% indicating successful decontamination mapping 70.93%; depth 66.92; GC 50.5%
- ▲ Scaffold N50 increased >16.7-fold vs G. domingensis (189.32 Kb → 3.16 Mb) 16.7-fold
- count 10,689 protein-coding genes (total genes predicted by EVM)
- fold_change 86.5-fold increase in contig N50 (30.15 Kb → 2.61 Mb) (vs previous G. vermiculophylla draft)
- other QV 34.45 (>99.96% accuracy), completeness 95.65% (Merqury assessment from short reads)
- count 22,323,461 Hi-C contact reads (30.64% of total) (after deduplication)
- mean average short-read sequencing depth 62.6; Nanopore depth 66.92 (BWA/Minimap2 mapping to assembly)
- other GC content 50.5% (genome GC rate, vs 70.06% in Neoporphyra haitanensis)
- count 18.16 Gb clean short reads (Q20 97.13%, Q30 92.21%); 14.9 Gb Nanopore (avg 4,799 bp) (sequencing output)
- other BUSCO genome C:74.9% [S:73.3%,D:1.6%], F:8.6%, M:16.5%, n:255 (eukaryota_odb10 evaluation)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a data descriptor reporting a chromosome-level genome assembly for Gracilaria vermiculophylla, generated by combining DNBSeq short-read, Nanopore long-read, and Hi-C sequencing. No inferential statistical hypothesis testing was performed; quality validation relied entirely on standard bioinformatics metrics including BUSCO completeness scores, Merqury QV, sequencing depth, mapping rates, and GC-content distribution. Gene and repeat content were reported as percentages of the assembled genome or gene set annotated across multiple databases. All results are point estimates without measures of dispersion, consistent with convention for single-specimen genome data descriptors.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| BUSCO completeness evaluation (v5.3.1, eukaryota_odb10 database) | Genome assembly and gene set quality assessment (Table 4); also used to compare gene set completeness against previously published Gracilaria assemblies | 255 conserved eukaryotic single-copy orthologs | not stated |
| Merqury k-mer-based Quality Value (QV) scoring and completeness estimation | Assembly accuracy validation; QV reported as 34.45 (>99.96% accuracy), completeness 95.65% | — | not stated |
| Short-read mapping rate and mean sequencing depth (BWA v0.7.17) | Technical validation of assembly; average depth 62.6× | 18.16 Gb clean paired-end 150 bp reads | not stated |
| Nanopore read mapping (Minimap2) and GC-content distribution inspection | Decontamination validation and coverage assessment (Fig. 6); mapping rate 70.93%, coverage 99.88%, mean depth 66.92× | 14.9 Gb Nanopore long reads | not stated |
| Descriptive comparison of assembly contiguity metrics (genome size, N50, contig count) against previously published Gracilaria and Florideophyceae genomes | Technical Validation section; fold-change ratios reported (e.g., 86.5-fold contig N50 increase over prior G. vermiculophylla assembly) | — | na |
-
Assembly completeness was benchmarked with the generic eukaryota_odb10 BUSCO database (255 genes), which is the broadest eukaryote lineage available↳ Could also: A more lineage-specific BUSCO database (e.g., a Rhodophyta- or algae-specific set, if one becomes available) or the LTR Assembly Index (LAI) metric could also supplement evaluation — Lineage-specific databases capture a higher proportion of expected orthologs for the focal clade and yield more sensitive completeness estimates; LAI additionally quantifies the assembly quality of repeat-rich regions, which is particularly relevant given that TEs constitute 59.26% of this genome
-
Decontamination relied on Hi-C intra-scaffold interaction strength combined with BLAST-based taxonomic assignment to separate host, bacterial, and organellar sequences↳ Could also: Blobtools (BlobDB) integrating sequencing depth, GC content, and taxonomic classification could also be used for contamination screening — Blobtools produces blobplots that simultaneously visualize all three axes across every contig, enabling a complementary or standalone decontamination approach that does not depend on Hi-C signal and can detect contaminants that share similar interaction patterns with the host
-
Nanopore long-read polishing used three rounds of Racon followed by two rounds of Pilon with short reads↳ Could also: Medaka (Oxford Nanopore Technologies' neural-network-based polisher) could also be applied for the Nanopore-specific polishing step, before or instead of Racon — Medaka is trained on Nanopore signal characteristics and can resolve systematic basecalling errors that consensus-based approaches may handle less precisely, potentially increasing the QV prior to short-read Pilon correction
-
Repeat annotation combined a de novo library from RepeatModeler with homology-based masking via RepeatMasker against RepBase↳ Could also: EDTA (Extensive de-novo TE Annotator) could also be used for de novo transposable element annotation — EDTA integrates multiple TE-discovery tools under a standardized, end-to-end filtering workflow and has shown improved sensitivity and classification accuracy for LTRs — the dominant repeat class here (55.03% of the genome) — particularly in taxa with limited RepBase representation
-
Assembly contiguity improvements over prior assemblies were described as simple fold-change ratios in N50 and contig count (e.g., 86.5-fold contig N50 increase)↳ Could also: Reference-based evaluation with QUAST or Assembly Likelihood Evaluation (ALE) against a related reference genome could also be applied — Fold-change in N50 captures contiguity gains but not structural correctness; QUAST can additionally quantify putative misassemblies, indel rates, and substitution rates, providing a more complete picture of assembly quality beyond contiguity metrics alone
-
Gene annotation integrated three evidence streams (de novo Augustus, homology via GeMoMa, transcriptome via StringTie/TransDecoder) combined with EVidenceModeler↳ Could also: BRAKER2 or MAKER2 could also serve as alternative integrative gene annotation frameworks — BRAKER2 couples Augustus with RNA-seq evidence via GeneMark-ET and automates species-specific parameter training, while MAKER2 provides a weighted-evidence framework with standardized repeat-masking integration; outputs from these pipelines could be compared against EVM-based results to assess annotation sensitivity and specificity
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41629338
Paper: Chromosome-level genome assembly of agar-producing red seaweed Gracilaria vermiculophylla. Jian et al., Sci Data 2026. DOI 10.1038/s41597-026-06635-3 · PMCID PMC12966425.
This is a Data Descriptor (genome assembly). Reported numeric results come from a multi-tool bioinformatic pipeline. We reproduce the clearly-specified, deterministic outputs derivable from the deposited final assembly + annotation, using standard third-party tools on the paper's own deposited data (per BRIEF rule P16: third-party tool on the paper's data is equally valid).
Deposited data actually used (NOTE: differs from scaffold accession)
The scaffold/manifest lists SRR23609112, but the paper's own accessions are:
- SRA project SRP564194: SRR32361123 (Nanopore), SRR32361124 (short-read DNBSEQ), SRR32361125 (Hi-C).
- GenBank assembly JBPJGC000000000.
- figshare 10.6084/m9.figshare.28667702.v3 (CC BY 4.0) — final assembly +
annotation:
Gv_unknow_new.fa(genome),Gv_unknow_new.gff(annotation),Gv_unknow_new.cds.fa(CDS),Gv_unknow_pep.fa(peptides).SRR23609112appears to be a text-mining mis-resolution; we use the figshare deposit + SRP564194 as the true paper data.
IN SCOPE (attempted) — derivable from the deposited assembly/annotation
| # | Reported result | Paper value | How we reproduce |
|---|---|---|---|
| ASM-1 | Genome assembly size | 77,513,337 bp (77.5 Mb) | seqkit stats -a on Gv_unknow_new.fa |
| ASM-2 | # pseudochromosomes/sequences | 22 (figshare page says 23) | count records in genome FASTA |
| ASM-3 | Scaffold N50 | 3,164,947 bp (3.16 Mb) | seqkit N50 on genome FASTA |
| ASM-4 | Largest scaffold | 4,532,714 bp (4.53 Mb) | seqkit max-len |
| ASM-5 | GC content | 50.5% | seqkit GC% |
| ANN-1 | # protein-coding genes | 10,689 | count gene/mRNA in GFF; count records in pep.fa |
| VAL-1 | BUSCO completeness (genome) | 74.9% complete | BUSCO v5.3.1 genome mode, eukaryota_odb10 |
| VAL-2 | BUSCO completeness (gene set) | 77.3% complete | BUSCO v5.3.1 protein mode on pep.fa |
OUT OF SCOPE (not attempted) — and why
- Full de-novo assembly (NextDenovo → Racon → Pilon → Juicer/3D-DNA Hi-C scaffolding): this is the hard ~80%; requires the full raw read sets (≈55 Gb), days of compute, and many non-deterministic steps. We instead verify the deposited assembly's reported metrics. Skipped per 80/20 rule.
- Repeat/TE content (59.26%, RepeatMasker/RepeatModeler/LTR_FINDER): a full repeat-annotation pipeline; not low-hanging. Skipped.
- Functional annotation counts (NR/KEGG/Swiss-Prot %): depend on external DB versions; not cleanly reproducible. Skipped.
- Read-level QV / Merqury / mapping depth (62.6×, QV 34.45, 95.65%): require downloading the full raw reads. Skipped (would chase the last 20%).
- Wet-lab (DNA extraction, OD ratios): out of scope (non-computational).
Contig N50 caveat
Contig N50 (2,613,191 bp) and contig count (80) require splitting scaffolds at
assembly gaps (N-runs). The deposited Gv_unknow_new.fa is the scaffold/pseudo-
chromosome level FASTA, so scaffold-level metrics (ASM-1..5) are the direct,
honest comparison. Contig-level metrics are reported as secondary if gaps exist.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Reproduction ran third-party tools (seqkit, BUSCO 5.3.1) on the authors' own deposited final assembly, and 8/9 reported metrics reproduce bit-exactly or within tolerance — genome size, N50, largest contig, chromosome count, 95.19% anchored, and 10,689 genes are all exact. The lone deviation is the gene-set BUSCO Complete (~2.4pp lower), an explainable BUSCO mode/version artifact on our side, not an authors' or data-availability defect. No fabrication concern: every value derives cleanly from the shared figshare deposit, and the central chromosome-level-assembly claim holds fully.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.