Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Chromosome-level genome of the long-tailed marine-living ornate spiny lobster, Panulirus ornatus.

Sci Data · 2024
L1 80/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
80/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 56% of all assessed papers rank 484 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce: YES. Repo is README-only (parameters, no scripts), but assembly (NCBI GCA_036320965.1) + survey reads (SRA SRR26801482) both resolve, so this is reproducible via standard third-party tools on the paper's own data (brief P16). RESULT = 1:1 agreement on the assembly-statistics data point (DP1): all 6 reported metrics match the deposited assembly to <=0.5%, with scaffold N50 (51,049,391 bp) and chromosome count (73) EXACT to the unit -> no evidence of assembly-stat fabrication. This DP1 check used the authoritative NCBI Datasets metadata (control-plane, zero compute), NOT an independent from-FASTA recompute. NOT attempted/blocked: DP2 (k-mer genome survey: KMC k=17 + GenomeScope2 on SRR26801482 to test genome size 2917.34 Mb / heterozygosity 0.92%) was fully scripted and staged for «our HPC» but NOT executed because the «our HPC» VPN tunnel was never established (Cisco SAML 2FA needs the human operator; not completed in the session window before finalize was requested). NOT attempted by design (hard 20%): full de novo assembly (292 Gb PacBio CLR), Hi-C scaffolding (456 Gb + manual Juicebox curation), gene/repeat annotation. FLAG for human: repo README genome-size estimate (2524.70 Mb) disagrees ~16% with the paper text (2917.34 Mb) for the same k-mer survey.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 80
    assessed: 2026-06-15 ⛓ d16f600d7bbd
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a high-quality, chromosome-level reference genome be assembled for the endangered ornate spiny lobster (Panulirus ornatus) to support its conservation, population genetics, and comparative crustacean genomics?

Core claims
  • A chromosome-level genome of P. ornatus spanning 2.65 Gb was assembled with a contig N50 of 51.05 Mb, anchoring 99.11% of sequences to 73 chromosomes. resource
  • The genome comprises 65.67% repeat sequences. finding
  • 22,752 protein-coding genes were predicted, of which 99.20% were functionally annotated. finding
  • Integrating Illumina short reads, PacBio long reads, and Hi-C produced the first chromosome-level genome assembly for an endangered lobster species. method
  • This assembly markedly improves on a previous fragmented attempt (1.93 Gb assembled, contig N50 of 5,451 bp). finding
  • Four types of noncoding RNAs were annotated: 12,771 miRNAs, 5,187 tRNAs, 1,716 rRNAs, and 1,296 snRNAs. finding
Experimental setups
Assay System Perturbation Readout Platform
WGS short-read sequencing P. ornatus male adult muscle tissue none 182.90 Gb of short reads for genome survey and polishing Illumina HiSeq 6000
WGS long-read sequencing (PacBio CLR) P. ornatus muscle tissue none 292.02 Gb raw continuous long reads for de novo assembly PacBio Sequel II (SMRT)
Hi-C sequencing/scaffolding P. ornatus muscle tissue MboI restriction digestion / cross-linking 456.71 Gb paired-end reads for chromosome scaffolding Illumina HiSeq/NovaSeq 6000
RNA-seq (transcriptome) P. ornatus eight tissues (testis, intestines, hepatopancreas, hemocytes, muscle, gills, heart, eyestalk) none 54.38 Gb clean reads for gene structure annotation Illumina HiSeq 6000; NEBNext Ultra RNA Library Prep Kit
K-mer genome survey P. ornatus Illumina reads none estimated genome size, heterozygosity, repeat content SOAPec v2.01; GenomeScope v2.0
Key results
  • Final genome assembly size of 2,651,872,113 bp (2.65 Gb) 2.65 Gb
  • Scaffold N50 reached 51.05 Mb in the final assembly 51.05 Mb
  • 2,628.95 Mb anchored to 73 chromosomes, accounting for 99.11% of the assembly 99.11%
  • Repeat sequences constituted 65.67% of the genome 65.67%
  • LINEs accounted for 40.30%, LTRs 30.07%, DNA elements 4.58%, SINEs 0.01% of the genome LINE 40.30%; LTR 30.07%
  • 22,568 of 22,752 predicted genes (99.20%) annotated by at least one database 99.20%
  • Estimated genome size by K-mer survey was 2917.34 Mb with heterozygosity 0.92% and repeat content 63.86% 2917.34 Mb; 0.92%
  • 14 chromosomes assembled with no more than 30 gaps each 14 chromosomes
Key statistics
  • count 2,651,872,113 bp total assembly length (Final genome size)
  • count 51.05 Mb scaffold N50 (Hi-C assembly continuity)
  • count 73 chromosomes (Anchored chromosomes)
  • other 99.11% (Proportion of assembly anchored to chromosomes)
  • count 22,752 protein-coding genes (Final gene set)
  • other 65.67% (Repeat sequence content of genome)
  • other 0.92% heterozygosity; 63.86% repeat (Genome survey estimates)
  • count K-mer dominant peak depth of 59; estimated 2917.34 Mb (17 K-mer frequencies analyzed in genome survey)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a genome assembly data descriptor reporting no inferential statistical tests; all analyses are computational and descriptive. Genome size (~2.92 Gb estimated, 2.65 Gb assembled) and heterozygosity (0.92%) were estimated by 17-mer K-mer frequency profiling using SOAPec and GenomeScope2. Assembly quality was characterized by contig/scaffold N50 values, total assembled length, and chromosomal anchoring percentages. Gene prediction and functional annotation results were reported as counts and proportions across multiple databases and prediction methods.

Replicationunclear Sample sizeSingle male adult P. ornatus individual from a commercial source; no sample size justification or power analysis stated; single-specimen reference genome assembly GroupsNo groups compared; single-individual genome characterization and annotation Pairingna Randomization/blindingnot stated Dispersionnone
Statistical tests used
Test Applied to n Assumptions
K-mer frequency analysis (17-mer) for genome size and heterozygosity estimation Genome survey prior to assembly; dominant peak depth of 59 used to estimate genome size at 2,917.34 Mb and heterozygosity at 0.92% not stated
BLAST sequence alignment with E-value cutoff 1E-5 Functional annotation of 22,752 predicted protein-coding genes against SwissProt, NR, KEGG, InterPro, GO, and Pfam databases 22,752 predicted genes not stated
Kimura two-parameter divergence calculation for transposable elements TE landscape characterization to show divergence rate distribution (Fig. 4) using calcDivergenceFromAlign.pl and createRepeatLandscape.pl not stated
Approaches that could also have been used
  • Assembly completeness was characterized by N50 values and total assembled length; no standardized gene-space completeness benchmarking tool (e.g., BUSCO) was mentioned in the provided text
    Could also: BUSCO (Benchmarking Universal Single-Copy Orthologs) assessed against an arthropod or metazoan lineage database — BUSCO provides a universally comparable completeness metric — the proportion of expected conserved single-copy genes recovered — that complements contiguity metrics like N50 and enables direct, standardized comparison across assemblies from different species and studies
  • Genome size and heterozygosity were estimated using a single K value (17-mer) with SOAPec and GenomeScope2
    Could also: K-mer profiling at multiple values (e.g., 17, 21, 31-mer) with cross-validation, or Smudgeplots alongside GenomeScope2 — Comparing estimates across K values helps assess the robustness of the size and heterozygosity estimates; Smudgeplots additionally infer ploidy level from K-mer pair coverages, which can be informative when heterozygosity is non-negligible (here 0.92%)
  • Two long-read assemblers (Wtdbg2 and Flye) were run independently and their outputs merged with Quickmerge
    Could also: Hifiasm or Canu could serve as primary or secondary assemblers in a dual-assembler strategy; Hifiasm also supports native Hi-C phasing — Hifiasm is optimized for PacBio CLR/HiFi data and frequently achieves high contiguity; its built-in Hi-C phasing mode can produce haplotype-resolved assemblies in a single workflow, which may be relevant given the observed heterozygosity
  • Assembly polishing was performed with two rounds of Arrow (PacBio-based consensus) followed by two rounds of Pilon (Illumina short-read based)
    Could also: HyPo, NextPolish, or Medaka as alternative polishing tools, applied sequentially or in place of one of the Pilon rounds — HyPo has been shown to achieve high per-base accuracy with fewer computational resources and iterations; benchmarking polishing tools on the draft assembly can identify the approach that minimizes residual errors for a given genome and read set
  • Gene prediction integrated five de novo tools, nine homology sources, and RNA-seq evidence via EvidenceModeler, with manual PASA update
    Could also: MAKER2 or BRAKER2 (with AUGUSTUS and GeneMark) pipelines could also integrate these evidence types in a unified, reproducible framework — MAKER2 produces standardized Annotation Edit Distance (AED) scores for each predicted gene model, providing an objective, per-gene quality metric alongside the aggregate annotation statistics reported here
  • The reference genome was assembled from a single male individual with no additional individuals sequenced
    Could also: Trio-binning (using parental short reads) or Hi-C-phased assembly with Hifiasm could produce a fully phased, haplotype-resolved diploid assembly from the same single individual — Single-individual assemblies are the standard approach for reference genome projects; however, given the 0.92% heterozygosity estimated for this species, a phased assembly would additionally capture allelic variation at the chromosome scale, which could be informative for downstream population genomics studies
Software: fastp 0.23.1 · SOAPec 2.01 · GenomeScope 2.0 · Wtdbg2 2.5 · Flye 2.9 · Arrow (SMRT Analysis) 8.0 · Quickmerge 0.3 · Pilon 1.22 · Juicer · 3D-DNA · Juicebox Assembly Tools 2.13.06 · RepeatScout 1.0.5 · RepeatModeler 2.0.1 · Piler 1.0 · LTR-FINDER 1.0.6 · RepeatMasker 4.1.0 · RepeatProteinMask 4.1.0 · TRF 4.0.9 · tRNAScan 1.4 · BLAST 2.2.26 · INFERNAL 1.0 · Augustus 3.2.3 · GlimmerHMM 3.02 · SNAP 2013.11.29 · Geneid 1.4 · Genscan 1.0 · Genewise 2.4.1 · Trinity 2.11.0 · PASA 2.1.0 · EvidenceModeler (EVM) 1.1.1 · Diamond 0.8.22 · InterProScan 5.52-86.0

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
8
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

10.6084/m9.figshare.24654915.v1 DOI in References (http://purl.org/orb/References)
no other assessed paper uses this yet
GCA_036320965.1 GCA in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-38909031

Paper: Ren et al. 2024, Sci Data — Chromosome-level genome of the ornate spiny lobster Panulirus ornatus. DOI 10.1038/s41597-024-03512-9. PMCID PMC11193758. Repo: github.com/sundongfang/Chromosome-level-genome-of-Panulirus-ornatus — README only (parameter list, no runnable scripts). Per brief P16, applying standard third-party tools to the paper's own data is an equally valid reproduction. Assembly: NCBI GCA_036320965.1 (ASM3632096v1). Survey reads: SRR26801482 (Illumina WGS PE150, ~93.8 Gbp). BioProject PRJNA1036297.

In scope (attempted)

result pipeline feasibility data point
Assembly total length, scaffold N50, contig N50, #scaffolds, #chromosomes, max contig, GC recompute directly from published assembly FASTA (seqkit/assembly-stats) LIGHT — download ~0.8 GB, minutes DP1
Genome-survey size, heterozygosity, repeat % k-mer count (KMC k=17) + GenomeScope2 on SRR26801482 MEDIUM — ~46 GB download + k-mer count DP2
BUSCO completeness BUSCO arthropoda on assembly optional — version mismatch (paper used odb9/BUSCO v3; modern = odb10/v5), so not 1:1 DP3 (if time)

Out of scope (NOT attempted — too heavy or wet-lab/external)

  • Full de novo assembly (Wtdbg2/Flye + Arrow + Quickmerge + Pilon on 292 Gb PacBio CLR): hundreds of CPU-days, TB of data. Outside 80/20.
  • Hi-C scaffolding (Juicer + 3D-DNA on 456 Gb Hi-C): very heavy, manual Juicebox curation step is non-deterministic.
  • Gene annotation (Augustus/Trinity/PASA/EVM → 22,752 genes), repeat annotation (65.67%), ncRNA, functional annotation: multi-tool pipelines, days of compute, many unpinned params.
  • DNA extraction / sequencing / karyotype: wet-lab, not computational.

Rationale

DP1 is a deterministic recomputation of reported assembly metrics from the exact published assembly — the cleanest possible 1:1 check and a direct fabrication test. DP2 reproduces the genome-survey pipeline (the only k-mer step) on the paper's own reads with the named tool family (GenomeScope2 v2.0, k=17). Both are clearly specified and low-cost; the assembly/annotation pipelines are the optional hard 20% and are explicitly skipped.

Figures / tables: Table
scaffold_n50
Reported
51049391 bp
Reproduced
51049391 bp (NCBI GCA_036320965.1)
exact
n_chromosomes
Reported
73
Reproduced
73 (NCBI GCA_036320965.1)
exact
assembly_total_len
Reported
2651872113 bp
Reproduced
2652407876 bp (NCBI GCA_036320965.1, +0.020%)
within tolerance
contig_n50
Reported
5119584 bp
Reproduced
5108882 bp (NCBI GCA_036320965.1, -0.209%)
within tolerance
n_scaffolds
Reported
1456
Reproduced
1450 (NCBI GCA_036320965.1, -0.412%)
within tolerance
n_contigs
Reported
8061
Reproduced
8058 (NCBI GCA_036320965.1, -0.037%)
within tolerance
genome_size_survey
Reported
2917.34 Mb (README says 2524.70 Mb)
Reproduced
not run (DP2 staged, needs «our HPC»)
partial
heterozygosity
Reported
0.92%
Reproduced
not run (DP2 staged)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 80/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

91.2 k
tokens (I/O) · 6.1 M incl. cache
28 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.