Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Chromosome-level genome assembly of agar-producing red seaweed Gracilaria vermiculophylla.

Sci Data · 2026
L1 91/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
How its reproducibility compares
91/100
Reproducibility score
1.0 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 82% of all assessed papers rank 197 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough and reproduced ~1:1. Approach: ran third-party tools (seqkit, BUSCO v5.3.1) on the paper's OWN deposited final assembly + annotation (figshare 10.6084/m9.figshare.28667702.v3, CC BY 4.0) on «our HPC»/«infra». 8/9 reported metrics reproduce exactly or within tolerance: genome size 77,513,337 bp, scaffold N50 3,164,947 bp, largest contig 4,532,714 bp, anchored 95.19%, 22 chromosomes, and 10,689 genes are all bit-exact; GC 50.45% (=50.5%); genome BUSCO within-tol (D 1.6% / M 16.5% / n=255 identical, Complete +0.8pp = 193 vs 191). Only the gene-set (protein) BUSCO is partial: C 74.9% vs reported 77.3% (~2.4pp lower; duplicated 0.4% matches), likely a BUSCO mode/version difference. NOT attempted (out of scope, 80/20): full de-novo assembly (NextDenovo/Racon/Pilon/Juicer/3D-DNA), repeat/TE annotation (59.26%), functional-annotation DB percentages, and read-level QV/Merqury/coverage metrics (would require the ~55 Gb raw reads). No fabrication concern — every reproduced value derives cleanly from the shipped figshare data. Note: the manifest's data accession SRR23609112 was a text-mining mis-resolution; the paper's real data is SRA SRP564194 (SRR32361123/124/125) + the figshare deposit.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 91
    assessed: 2026-06-16 ⛓ 903c995a2f0d
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study addresses the lack of a high-quality chromosome-level genome for the agar-producing red seaweed Gracilaria vermiculophylla, aiming to generate one as a reference resource for agar biosynthesis, evolutionary, comparative genomic, and ecological research.

Core claims
  • A chromosome-level genome assembly of G. vermiculophylla was generated by combining DNBSeq short reads, Nanopore long reads, and Hi-C data. resource
  • The assembled genome is 77.5 Mb with contig N50 of 2.61 Mb and scaffold N50 of 3.16 Mb, comprising 22 pseudochromosomes. finding
  • Transposable elements constitute 45.93 Mb (59.26%) of the genome, with LTRs the predominant retrotransposons (55.03%). finding
  • The genome contains 10,689 protein-coding genes, of which 86.14% were functionally annotated. finding
  • Hi-C interaction data were used to decontaminate the assembly by distinguishing host chromosomal contigs from bacterial sequences. method
  • The new assembly markedly improves contiguity over the previous G. vermiculophylla draft (size 45.9→77.5 Mb; contig N50 30.15 Kb→2.61 Mb, an 86.5-fold increase). finding
  • BUSCO, GC content, and sequencing depth assessments demonstrate the high quality of the assembly and successful decontamination. finding
  • Chr07 aligns to the sex-determination region; 21 of 26 published sex-linked markers mapped, with seven aligning specifically to Chr07. finding
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome short-read sequencing (DNBSEQ WGS) Gracilaria vermiculophylla thallus none short-read genomic sequence (PE150) MGI-T7 / DNBSEQ platform; Covaris E220 ultrasonicator
Nanopore long-read sequencing Gracilaria vermiculophylla thallus none long-read genomic sequence Nanopore CycloneSEQ; motor protein BCH-X
Hi-C (chromatin conformation capture) sequencing Gracilaria vermiculophylla fresh thalli none chromatin contact reads for chromosome anchoring and decontamination DNBSEQ-T7 (PE150); MboI/DpnII (NEB), T4 DNA Ligase
HMW DNA extraction Gracilaria vermiculophylla (~3 g fresh sample) none DNA quality (OD260/280=1.93, OD260/230=1.87) phenol-chloroform/CTAB protocol
Transcriptome-based gene annotation (RNA-seq mapping) Gracilaria vermiculophylla genome none aligned/assembled transcripts for gene model prediction public SRA RNA-seq data; Hisat2, StringTie
BUSCO completeness evaluation Gracilaria vermiculophylla genome and gene set none % complete/single-copy/duplicated/fragmented/missing BUSCOs BUSCO v5.3.1, eukaryota_odb10 (255 genes)
Repeat/TE annotation Gracilaria vermiculophylla genome none TE content and classification TRF, RepeatMasker, RepeatModeler, LTR_FINDER, RepBase
Functional annotation Gracilaria vermiculophylla protein-coding genes none % genes annotated across databases NR, Swiss-Prot, KEGG, KOG, TrEMBL, InterPro, GO
Key results
  • Final nuclear genome assembled at 77.5 Mb with contig N50 2.61 Mb and scaffold N50 3.16 Mb; 22 pseudochromosomes totaling 73.78 Mb (95.2% anchored) 77.5 Mb; N50 2.61/3.16 Mb
  • Transposable elements account for 45.93 Mb (59.26%); LTRs 55.03% with Gypsy 36.8% and Copia 12.8% 45.93 Mb / 59.26%
  • 10,689 protein-coding genes predicted; 9,207 (86.14%) functionally annotated 10,689 genes; 86.14%
  • Contig N50 improved from 30.15 Kb to 2.61 Mb vs prior draft; contig count dropped from 4,039 to 80 86.5-fold (contig N50); ~50.5-fold (contig count)
  • Genome BUSCO completeness 74.9% complete; gene set 77.3% complete BUSCOs C:74.9% (genome), C:77.3% (gene)
  • Short-read QV of 34.45 (>99.96% accuracy) and 95.65% completeness via Merqury; average short-read depth 62.6 QV 34.45; depth 62.6
  • Nanopore reads mapped at 70.93% with 99.88% coverage and 66.92 average depth; GC clustered tightly around 50.5% indicating successful decontamination mapping 70.93%; depth 66.92; GC 50.5%
  • Scaffold N50 increased >16.7-fold vs G. domingensis (189.32 Kb → 3.16 Mb) 16.7-fold
Key statistics
  • count 10,689 protein-coding genes (total genes predicted by EVM)
  • fold_change 86.5-fold increase in contig N50 (30.15 Kb → 2.61 Mb) (vs previous G. vermiculophylla draft)
  • other QV 34.45 (>99.96% accuracy), completeness 95.65% (Merqury assessment from short reads)
  • count 22,323,461 Hi-C contact reads (30.64% of total) (after deduplication)
  • mean average short-read sequencing depth 62.6; Nanopore depth 66.92 (BWA/Minimap2 mapping to assembly)
  • other GC content 50.5% (genome GC rate, vs 70.06% in Neoporphyra haitanensis)
  • count 18.16 Gb clean short reads (Q20 97.13%, Q30 92.21%); 14.9 Gb Nanopore (avg 4,799 bp) (sequencing output)
  • other BUSCO genome C:74.9% [S:73.3%,D:1.6%], F:8.6%, M:16.5%, n:255 (eukaryota_odb10 evaluation)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a data descriptor reporting a chromosome-level genome assembly for Gracilaria vermiculophylla, generated by combining DNBSeq short-read, Nanopore long-read, and Hi-C sequencing. No inferential statistical hypothesis testing was performed; quality validation relied entirely on standard bioinformatics metrics including BUSCO completeness scores, Merqury QV, sequencing depth, mapping rates, and GC-content distribution. Gene and repeat content were reported as percentages of the assembled genome or gene set annotated across multiple databases. All results are point estimates without measures of dispersion, consistent with convention for single-specimen genome data descriptors.

Replicationunclear Sample sizeSingle individual collected; no formal sample size justification or power calculation stated; one genome assembled from one specimen GroupsNo experimental treatment groups; assembly quality metrics compared descriptively against previously published Gracilaria and Florideophyceae reference genomes Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
BUSCO completeness evaluation (v5.3.1, eukaryota_odb10 database) Genome assembly and gene set quality assessment (Table 4); also used to compare gene set completeness against previously published Gracilaria assemblies 255 conserved eukaryotic single-copy orthologs not stated
Merqury k-mer-based Quality Value (QV) scoring and completeness estimation Assembly accuracy validation; QV reported as 34.45 (>99.96% accuracy), completeness 95.65% not stated
Short-read mapping rate and mean sequencing depth (BWA v0.7.17) Technical validation of assembly; average depth 62.6× 18.16 Gb clean paired-end 150 bp reads not stated
Nanopore read mapping (Minimap2) and GC-content distribution inspection Decontamination validation and coverage assessment (Fig. 6); mapping rate 70.93%, coverage 99.88%, mean depth 66.92× 14.9 Gb Nanopore long reads not stated
Descriptive comparison of assembly contiguity metrics (genome size, N50, contig count) against previously published Gracilaria and Florideophyceae genomes Technical Validation section; fold-change ratios reported (e.g., 86.5-fold contig N50 increase over prior G. vermiculophylla assembly) na
Approaches that could also have been used
  • Assembly completeness was benchmarked with the generic eukaryota_odb10 BUSCO database (255 genes), which is the broadest eukaryote lineage available
    Could also: A more lineage-specific BUSCO database (e.g., a Rhodophyta- or algae-specific set, if one becomes available) or the LTR Assembly Index (LAI) metric could also supplement evaluation — Lineage-specific databases capture a higher proportion of expected orthologs for the focal clade and yield more sensitive completeness estimates; LAI additionally quantifies the assembly quality of repeat-rich regions, which is particularly relevant given that TEs constitute 59.26% of this genome
  • Decontamination relied on Hi-C intra-scaffold interaction strength combined with BLAST-based taxonomic assignment to separate host, bacterial, and organellar sequences
    Could also: Blobtools (BlobDB) integrating sequencing depth, GC content, and taxonomic classification could also be used for contamination screening — Blobtools produces blobplots that simultaneously visualize all three axes across every contig, enabling a complementary or standalone decontamination approach that does not depend on Hi-C signal and can detect contaminants that share similar interaction patterns with the host
  • Nanopore long-read polishing used three rounds of Racon followed by two rounds of Pilon with short reads
    Could also: Medaka (Oxford Nanopore Technologies' neural-network-based polisher) could also be applied for the Nanopore-specific polishing step, before or instead of Racon — Medaka is trained on Nanopore signal characteristics and can resolve systematic basecalling errors that consensus-based approaches may handle less precisely, potentially increasing the QV prior to short-read Pilon correction
  • Repeat annotation combined a de novo library from RepeatModeler with homology-based masking via RepeatMasker against RepBase
    Could also: EDTA (Extensive de-novo TE Annotator) could also be used for de novo transposable element annotation — EDTA integrates multiple TE-discovery tools under a standardized, end-to-end filtering workflow and has shown improved sensitivity and classification accuracy for LTRs — the dominant repeat class here (55.03% of the genome) — particularly in taxa with limited RepBase representation
  • Assembly contiguity improvements over prior assemblies were described as simple fold-change ratios in N50 and contig count (e.g., 86.5-fold contig N50 increase)
    Could also: Reference-based evaluation with QUAST or Assembly Likelihood Evaluation (ALE) against a related reference genome could also be applied — Fold-change in N50 captures contiguity gains but not structural correctness; QUAST can additionally quantify putative misassemblies, indel rates, and substitution rates, providing a more complete picture of assembly quality beyond contiguity metrics alone
  • Gene annotation integrated three evidence streams (de novo Augustus, homology via GeMoMa, transcriptome via StringTie/TransDecoder) combined with EVidenceModeler
    Could also: BRAKER2 or MAKER2 could also serve as alternative integrative gene annotation frameworks — BRAKER2 couples Augustus with RNA-seq evidence via GeneMark-ET and automates species-specific parameter training, while MAKER2 provides a weighted-evidence framework with standardized repeat-masking integration; outputs from these pipelines could be compared against EVM-based results to assess annotation sensitivity and specificity
Software: SOAPnuke v2.0 · NextDenovo v2.5.2 · racon 1.3.3 · Pilon 1.24 · BWA (Burrows-Wheeler Aligner) 0.7.12 / 0.7.17 · Juicer pipeline v1.5 · 3D-DNA pipeline v180922 · TRF (Tandem Repeats Finder) 4.9 · RepeatMasker open-4.0.9 · RepeatProteinMask v4.0.7 · RepeatModeler open-1.0.11 · LTR_FINDER_parallel 1.0.7 · Augustus 3.2.1 · Hisat2 v2.1.0 · StringTie 2.2.1 · GeMoMa 1.9 · TransDecoder v5.5.0 · EVidenceModeler (EVM) v2.1.0 · Minimap2 · Merqury v1.3 · BUSCO v5.3.1

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41629338

Paper: Chromosome-level genome assembly of agar-producing red seaweed Gracilaria vermiculophylla. Jian et al., Sci Data 2026. DOI 10.1038/s41597-026-06635-3 · PMCID PMC12966425.

This is a Data Descriptor (genome assembly). Reported numeric results come from a multi-tool bioinformatic pipeline. We reproduce the clearly-specified, deterministic outputs derivable from the deposited final assembly + annotation, using standard third-party tools on the paper's own deposited data (per BRIEF rule P16: third-party tool on the paper's data is equally valid).

Deposited data actually used (NOTE: differs from scaffold accession)

The scaffold/manifest lists SRR23609112, but the paper's own accessions are:

  • SRA project SRP564194: SRR32361123 (Nanopore), SRR32361124 (short-read DNBSEQ), SRR32361125 (Hi-C).
  • GenBank assembly JBPJGC000000000.
  • figshare 10.6084/m9.figshare.28667702.v3 (CC BY 4.0) — final assembly + annotation: Gv_unknow_new.fa (genome), Gv_unknow_new.gff (annotation), Gv_unknow_new.cds.fa (CDS), Gv_unknow_pep.fa (peptides). SRR23609112 appears to be a text-mining mis-resolution; we use the figshare deposit + SRP564194 as the true paper data.

IN SCOPE (attempted) — derivable from the deposited assembly/annotation

# Reported result Paper value How we reproduce
ASM-1 Genome assembly size 77,513,337 bp (77.5 Mb) seqkit stats -a on Gv_unknow_new.fa
ASM-2 # pseudochromosomes/sequences 22 (figshare page says 23) count records in genome FASTA
ASM-3 Scaffold N50 3,164,947 bp (3.16 Mb) seqkit N50 on genome FASTA
ASM-4 Largest scaffold 4,532,714 bp (4.53 Mb) seqkit max-len
ASM-5 GC content 50.5% seqkit GC%
ANN-1 # protein-coding genes 10,689 count gene/mRNA in GFF; count records in pep.fa
VAL-1 BUSCO completeness (genome) 74.9% complete BUSCO v5.3.1 genome mode, eukaryota_odb10
VAL-2 BUSCO completeness (gene set) 77.3% complete BUSCO v5.3.1 protein mode on pep.fa

OUT OF SCOPE (not attempted) — and why

  • Full de-novo assembly (NextDenovo → Racon → Pilon → Juicer/3D-DNA Hi-C scaffolding): this is the hard ~80%; requires the full raw read sets (≈55 Gb), days of compute, and many non-deterministic steps. We instead verify the deposited assembly's reported metrics. Skipped per 80/20 rule.
  • Repeat/TE content (59.26%, RepeatMasker/RepeatModeler/LTR_FINDER): a full repeat-annotation pipeline; not low-hanging. Skipped.
  • Functional annotation counts (NR/KEGG/Swiss-Prot %): depend on external DB versions; not cleanly reproducible. Skipped.
  • Read-level QV / Merqury / mapping depth (62.6×, QV 34.45, 95.65%): require downloading the full raw reads. Skipped (would chase the last 20%).
  • Wet-lab (DNA extraction, OD ratios): out of scope (non-computational).

Contig N50 caveat

Contig N50 (2,613,191 bp) and contig count (80) require splitting scaffolds at assembly gaps (N-runs). The deposited Gv_unknow_new.fa is the scaffold/pseudo- chromosome level FASTA, so scaffold-level metrics (ASM-1..5) are the direct, honest comparison. Contig-level metrics are reported as secondary if gaps exist.

ASM-1
Reported
77,513,337 bp (77.5 Mb) genome size
Reproduced
77,513,337 bp
exact
ASM-2
Reported
22 pseudochromosomes
Reproduced
22 chromosomes (Chr01-22) + 16 unplaced contigs = 38 seqs
exact
ASM-3
Reported
scaffold N50 3,164,947 bp
Reproduced
3,164,947 bp
exact
ASM-4
Reported
largest contig 4,532,714 bp
Reproduced
4,532,714 bp (Chr04)
exact
ASM-5
Reported
GC 50.5%
Reproduced
50.45%
within tolerance
ANCHOR
Reported
73.78 Mb (95.2%) anchored to chromosomes
Reproduced
73,785,327 bp (95.19%)
exact
ANN-1
Reported
10,689 protein-coding genes
Reproduced
10,689 (gene & mRNA in GFF; 10,689 cds & pep records)
exact
VAL-1
Reported
BUSCO genome C:74.9%[S:73.3%,D:1.6%],F:8.6%,M:16.5%,n:255 eukaryota_odb10
Reproduced
C:75.7%[S:74.1%,D:1.6%],F:7.8%,M:16.5%,n:255
within tolerance
VAL-2
Reported
BUSCO gene set C:77.3%[S:76.9%,D:0.4%],F:5.5%,M:17.2%,n:255
Reproduced
C:74.9%[S:74.5%,D:0.4%],F:4.7%,M:20.4%,n:255 (proteins mode)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 91/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5

Reproduction ran third-party tools (seqkit, BUSCO 5.3.1) on the authors' own deposited final assembly, and 8/9 reported metrics reproduce bit-exactly or within tolerance — genome size, N50, largest contig, chromosome count, 95.19% anchored, and 10,689 genes are all exact. The lone deviation is the gene-set BUSCO Complete (~2.4pp lower), an explainable BUSCO mode/version artifact on our side, not an authors' or data-availability defect. No fabrication concern: every value derives cleanly from the shared figshare deposit, and the central chromosome-level-assembly claim holds fully.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

153.6 k
tokens (I/O) · 10.4 M incl. cache
29 min
runtime · 1.92 CPU-h
9.5 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine