A chromosome-level genome assembly of Plantago ovata.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED. The deposited chromosome-level assembly GCA_028274465.1 (UofA_Burton_1) was independently recomputed on a «our HPC» compute node (SLURM «job») directly from the NCBI FASTA (md5-verified d9a1cc9f...). All 13 Table 1 / Results structural metrics reproduce 1:1: total 500.94 Mb, 876 scaffolds, scaffold N50 128.87 Mb, contig N50 249.86 Kb, GC 38.4%, 4 pseudochromosomes of 137.73/128.87/114.44/106.35 Mb summing to 487.38 Mb (97.29%); contig count 4280 vs paper 4301 (within-tol, gap-definition). BUSCO v5.7.1 (viridiplantae_odb10, genome mode) reproduces the headline completeness EXACTLY at C:99.3% (422/425 Complete, 0 Missing); the S/D/F/M sub-split differs <=1.1pp, explained by gene-predictor difference (ours miniprot, paper metaeuk/augustus). OUT OF SCOPE (not attempted, deliberately): de novo assembly from raw PacBio CLR reads via Canu + 3D-DNA Hi-C scaffolding (multi-day, non-deterministic, deposited assembly is authoritative); MAKER gene annotation (41,820 genes); RepeatModeler/RepeatMasker repeat content (61.9%); wet-lab steps. The named repo (phasegenomics/matlock) is a single Hi-C BAM-filter step of the scaffolding pipeline, not a full assembly pipeline. Dataset profiling completed for all 4 accessions.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 94assessed: 2026-06-20 ⛓ fc8397395d56
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study aims to construct the first chromosome-scale reference genome assembly of Plantago ovata to overcome the lack of a genomic resource that has limited breeding-based improvement of psyllium husk quantity and quality.
- ★ A chromosome-level reference genome assembly of P. ovata was constructed using PacBio long reads and Hi-C scaffolding. resource
- ★ The final assembly covers ~500.94 Mb with 99.3% BUSCO gene set completeness. finding
- ★ 97.29% of the assembled sequence is anchored to four chromosomes with a scaffold N50 of ~128.87 Mb. finding
- ★ The P. ovata genome contains 61.90% repetitive content, with LTR retrotransposons (40.04%) as the dominant class. finding
- ★ 41,820 protein-coding genes, 411 non-coding RNAs, 108 rRNAs, and 1,295 tRNAs were annotated in the genome. finding
- Chromosome identity was assigned using 5S and 45S rDNA cluster distribution patterns. method
- Three short genomic regions show evidence of nuclear mitochondrial DNA (NUMT) insertions. finding
- Comparative orthology analysis places P. ovata among Laminales species using OrthoFinder-derived orthogroups. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| PacBio long-read (CLR) whole-genome sequencing | Plantago ovata (whole plant/genomic DNA) | none | contig assembly | Pacific Biosciences (PacBio) CLR |
| Hi-C chromosome conformation capture sequencing | Plantago ovata genomic DNA | none | chromosome-scale scaffolding | Hi-C |
| k-mer based genome size estimation | P. ovata corrected PacBio reads | none | estimated haploid genome size | findGSE v0.1.0; genomescope2 v2.0 (21-mer) |
| RNA-seq transcript aggregation and gene model prediction | P. ovata (multiple tissues, public and in-house RNA-seq) | none | gene models / annotation | MAKER v2.31.11 |
| Genome completeness assessment | P. ovata assembled genome and predicted proteome | none | % complete/fragmented/missing conserved orthologs | BUSCO v5.4.3 (viridiplantae_odb10) |
| Repeat element annotation and LTR Assembly Index scoring | P. ovata genome assembly | none | repeat content %, LAI score | LAI (LTR_retriever-based) |
| Comparative orthology analysis | P. ovata and 9 other plant species (Laminales, Brassicales, Solanales) | none | orthogroup assignment, gene duplication events | OrthoFinder v2.5.4 |
| Short/long-read remapping validation | P. ovata genomic Illumina and PacBio reads (SRR10076762, SRR14643405) | none | % reads mapped to assembly | Illumina; PacBio |
- – Final assembly size 500.94 Mb with scaffold N50 of 128.87 Mb across 876 sequences 500.94 Mb / N50 128.87 Mb
- ▲ BUSCO assembly completeness of 99.3% (only 1 of 425 genes missing) 99.3%
- – 97.29% (487.38 Mb) of genome anchored to four chromosomes; unplaced scaffolds only 2.71% 97.29%
- – Repeat content estimated at 61.90% (310.10 Mb), with LTRs at 40.04% (200.59 Mb), split between Ty1/Copia (19.69%) and Gypsy (20.29%) 61.90% / 40.04%
- – 41,820 protein-coding genes identified, of which only 56% (23,638) have AED < 0.5 41,820 genes; 56%
- – k-mer based genome size estimates (551.02 Mb via findGSE; 415.78 Mb via genomescope2) bracket the assembled genome size 551.02 Mb / 415.78 Mb
- ▲ Genomic GC content is 38.4%, while CDS GC content is 44.3%, ~6% higher 38.4% vs 44.3% (+6%)
- ▲ High read mapping rates confirm assembly quality: 95.81% Illumina, 92.25% PacBio, up to 96.10% RNA-seq 95.81% / 92.25% / 96.10%
- count 500.94 Mb (total assembly size)
- other N50 = 128.87 Mb (scaffold N50)
- other C:99.3% [S:94.1%, D:5.2%], F:0.5%, M:0.2%, n:425 (BUSCO assembly completeness (viridiplantae_odb10))
- count 41,820 (total protein-coding genes annotated)
- other 61.90% (total repeat content of genome)
- other 40.04% (LTR retrotransposon proportion of genome)
- other 38.40% (overall genomic GC content)
- other LAI = 10.27 (LTR Assembly Index score for assembly continuity)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper reports a chromosome-level genome assembly and annotation of Plantago ovata, combining PacBio long-read sequencing, Hi-C scaffolding, and RNA-seq-based gene prediction. The statistical/quantitative content consists mainly of descriptive bioinformatics metrics (assembly contiguity statistics, BUSCO completeness scores, LTR Assembly Index, k-mer-based genome size estimation, and GC-content distribution) rather than classical inferential hypothesis testing, and results are reported as summary values and comparisons to prior single-value estimates from the literature rather than as replicated experimental measurements with formal statistical tests.
-
Genome size was estimated using two k-mer-based methods (findGSE and GenomeScope2), and the resulting single estimates were compared descriptively to the final assembly size and to prior flow-cytometry-based values from other studies.↳ Could also: A formal statistical comparison (e.g., reporting confidence intervals around k-mer-based estimates, or a meta-analytic synthesis of the multiple published genome-size estimates) could also be used — This would let readers gauge the precision of each estimate and quantify how consistent the various methods (k-mer, flow cytometry, assembly length) are with one another, beyond a narrative comparison of point values.
-
Genome assembly quality was assessed using single-value benchmarking metrics (BUSCO completeness, LAI score, mapping rate percentages) computed once on the final assembly.↳ Could also: Bootstrapping or resampling-based approaches (e.g., subsampling reads or BUSCO gene sets to generate a distribution of completeness scores) could also be used — This would provide an estimate of variability/uncertainty around these single-assembly quality metrics, which are otherwise reported as fixed point values.
-
Comparative GC-content distributions across species were described qualitatively as unimodal versus bimodal patterns, based on visual comparison to figures in other publications (e.g., Kotwal et al., Singh et al.).↳ Could also: A quantitative test of distribution shape, such as Hartigan's dip test for multimodality, could also be used — This would allow the unimodal versus bimodal classification of GC-content distributions to be supported by a formal statistical criterion rather than visual inspection of published figures.
-
Orthogroup and gene-family comparisons across ten plant species were performed using OrthoFinder, with results shown as counts and proportions in bar charts and Venn diagrams (Fig. 3).↳ Could also: Statistical enrichment testing (e.g., Fisher's exact test or a hypergeometric test) for species-specific orthogroup expansions could also be used — This would let readers assess whether differences in orthogroup counts or gene duplication events between species/nodes are greater than expected by chance, complementing the descriptive counts shown.
-
Repeat content and LTR retrotransposon proportions were reported as single percentages for the P. ovata genome and compared narratively to ranges reported across 103 other genomes (Ou et al.).↳ Could also: Reporting the P. ovata values alongside the distribution (e.g., percentile rank or z-score relative to the reference set of genomes) could also be used — This would situate the P. ovata repeat content and LAI score more precisely within the comparator distribution, rather than relying on stating that the value falls within a previously reported range.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36707685 (A chromosome-level genome assembly of Plantago ovata)
Sci Rep 2023, DOI 10.1038/s41598-022-25078-5. Assembly = GCA_028274465.1 (UofA_Burton_1), BioProject PRJNA732452. Named code: phasegenomics/matlock (a Hi-C BAM filter used as one step of the scaffolding pipeline).
Pipeline (from Methods)
PacBio Sequel CLR (40.21 Gb, ~76x) -> Canu v2.1 -> pbgcpp polish -> purge_haplotigs -> Hi-C (bwa + samblaster + matlock filter) -> 3D-DNA scaffolding + Juicebox -> 4 pseudochromosomes. Annotation: MAKER-style gene models (41,820 genes), RepeatModeler/RepeatMasker (61.9% repeats), BUSCO (Viridiplantae_odb10).
IN SCOPE (pipeline-derived, reproducible from the deposited assembly)
- Assembly structural metrics (Table 1): total size, #scaffolds, scaffold N50, #contigs, contig N50, GC%, #pseudochromosomes, chromosome sizes, longest scaffold. -> Reproduced 1:1 by recomputing directly from the deposited GCA_028274465.1 FASTA (an independent third-party recomputation of the authors' reported numbers).
- BUSCO genome completeness (Viridiplantae_odb10): re-run BUSCO v5.7.1 on the assembly.
OUT OF SCOPE / not attempted (and why)
- De novo assembly from raw reads (Canu v2.1 on 40 Gb PacBio CLR + 3D-DNA Hi-C scaffolding): multi-day/multi-week compute, non-deterministic, and the authoritative output (the assembly) is already deposited. We instead verify the deposited assembly's reported metrics 1:1. Attempting the full assembly would not add reproducibility evidence beyond the deposited FASTA.
- Gene annotation (41,820 genes): MAKER pipeline requires curated evidence sets, training, and many tools; not reproducible at low cost; out of scope for this pass.
- Repeat content (61.9%) / LAI: RepeatModeler de-novo modelling is heavy and library-dependent; not attempted in this pass (candidate for a later extension).
- Wet-lab (DNA extraction, Hi-C library prep, flow cytometry genome size): out of scope.
Datasets profiled (PRJNA732452)
SRR14643405 (PacBio CLR WGS), SRR14643406 (Hi-C), + 76 RNA-seq runs. Profiled in data/dataset_profile.json.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.