Prospects of telomere-to-telomere assembly in barley: Analysis of sequence gaps in the MorexV3 reference genome.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED. Ran the paper's exact seqtk->TRF read-level repeat pipeline (seqtk 1.5-r133, TRF 4.10.0rc2, params '2 5 7 80 10 50 2000 -l 1 -h') on the paper's OWN long-read data. PRIMARY HiFi claims EXACT: C1 = 44 TTTAGGG arrays >1 kb (44/44), C2 = 317,475 bp ~ 317 kb («job», HiFi N=6,617,223 = Table 1). ONT stretch (PRJEB40588, all 22 runs, observed N=30,517,372 ~ 94% of Table 1's 32.5M; «job»): C7 the single most specific number in the paper reproduced EXACTLY -> longest trinucleotide array 50,473 AAG copies spanning 152,808 bp (153 kb), read ERR4659224.385782. C3=1646/1848 (89%), C4=408/452 (90%), C5=11.7/13 Mb (90%), C6=6272/7655 (82%, lower bound) -- all within ~10-18% of reported, the residual explained by observed-vs-reported read count and the cancelled trinucleotide-TRF tail. NO fabrication indicators: every value is derivable from the public deposit + named tools, and the two precise claims (C1/C2, C7) land on the nose. NOT attempted (out of scope, not a seqtk/TRF read pipeline): Bionano optical mapping, CENH3 ChIP-seq, cytogenetics/FISH, k-mer genome-size, MorexV3 assembly statistics. Infra notes: shared-account «infra»+/home quota exhaustion forced a checkpoint+resume of the ONT stream; data all on «infra», results only on «host».
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 70assessed: 2026-06-20 ⛓ 580e9171d2bf
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper investigates whether and how a telomere-to-telomere (T2T) assembly of the barley genome could be achieved by identifying what sequences are still missing from the current MorexV3 reference assembly and why gaps remain.
- ★ Almost all centromeric sequences and 45S ribosomal DNA repeat arrays are absent from the MorexV3 pseudomolecules finding
- ★ The majority of sequence gaps in MorexV3 can be attributed to assembly breakdown in long stretches of satellite repeats finding
- ★ Missing sequences cannot fully account for the difference between MorexV3 assembly size and flow cytometric genome size estimates finding
- ★ Telomeric satellite arrays (TTTAGGG) are not captured in their entirety by long reads (HiFi or ONT) finding
- ★ Subtelomeric repeat arrays (HvT01, pAS1) are disrupted by sequence gaps in the assembly finding
- ONT sequencing data used for MorexV3 were contaminated with chloroplast DNA, inflating raw coverage-based genome size estimates finding
- ★ Ultra-long sequence reads (>100 kb to 1 Mb) will likely be required to assemble centromeric and rDNA loci mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Optical genome mapping (Bionano) | barley cv. Morex | none | missing sequence at chromosome termini relative to pseudomolecules | Bionano optical map, DLE-1 labelling |
| ChIP-seq | barley cv. Morex, centromeric chromatin | none | centromeric histone variant CENH3 binding / centromeric repeat abundance | ChIP-seq |
| k-mer spectrum genome size estimation | barley cv. Morex genomic DNA | none | genome size estimate (GSE) at k=21,51,101 | findGSE on PacBio HiFi, ONT, and PE450 Illumina reads |
| Read-depth based genome size estimation | barley cv. Morex | none | genome size estimate from mean/median coverage | alignment of HiFi/ONT/PE450 reads to MorexV3 pseudomolecules |
| Fluorescence in situ hybridization (FISH) | barley cv. Morex metaphase chromosomes | none | localization of chloroplast DNA insertions, subtelomeric satellite HvT01, and 45S rDNA | FISH with cpDNA, HvT01, and 45S rDNA probes |
| Tandem Repeat Finder (TRF) annotation | barley cv. Morex ONT long reads | none | length and location of TTTAGGG telomeric repeat arrays | Tandem Repeat Finder |
| BLAST sequence alignment | barley cv. Morex HiFi long reads | none | abundance/length of subtelomeric repeats HvT01 and pAS1 | BLAST |
| Flow cytometry (literature values used for comparison) | barley cv. Morex nuclei | none | haploid (1C) genome size | flow cytometry |
- ▼ MorexV3 pseudomolecule assembly totals 4.196 Gb (plus 29.1 Mb unplaced contigs), below flow cytometric estimates of 4.88–5.04 Gb ~0.7-0.85 Gb difference
- ▼ Missing sequence at chromosome termini detected via optical map overhangs 17-220 kb (short arms), 10-80 kb (long arms)
- ▼ ONT reads show TTTAGGG telomeric arrays >1kb totaling 13 Mb cumulative, averaging ~11 kb per chromosome arm; HiFi captured far less HiFi: 741 bp per telomere vs ONT ~11 kb per arm
- – HvT01 subtelomeric repeat abundance estimated from HiFi reads 3.1 Mb total, 225 kb per chromosome arm
- – pAS1 subtelomeric repeat abundance estimated from HiFi reads 25.9 Mb total, 1.85 Mb per chromosome arm
- – 9.3% of ONT reads aligned to chloroplast genome, indicating cpDNA contamination rather than true nuclear insertions 29.5 Gb of ONT reads matched cpDNA
- ▼ After correcting for cpDNA contamination, ONT coverage-based genome size estimate revised 5 Gb to 4.7 Gb
- ▲ HvT01 sequence disproportionately assigned to unanchored contigs compared to pAS1 19.3% (707 kb) of HvT01 vs 0.5% (134 kb) of pAS1
- count 4.196 Gb pseudomolecules + 29.1 Mb unplaced contigs (MorexV3 assembly size)
- mean 4.88 and 5.04 Gb (1C genome size) (flow cytometric estimate by Doležel et al. 2018)
- other 3.559–5.349 Gb (range of published flow cytometric 1C genome size estimates for H. vulgare)
- fold_change GSE at k=51: HiFi 4.2 Gb, PE450 4.3 Gb (k-mer based genome size estimates)
- count 1848 ONT reads with TTTAGGG arrays >1 kb (452 with >10 kb); average array size 6.9 kb; longest 37.2 kb; cumulative 13 Mb (telomeric repeat quantification in ONT reads)
- count 44 HiFi reads with TTTAGGG arrays >1 kb, cumulative 317 kb (741 bp per telomere at 31x coverage) (telomeric repeat quantification in HiFi reads)
- count 3803 reads with HvT01 alignments ≥1 kb, cumulative 97.6 Mb, estimated 3.1 Mb total array size (HvT01 subtelomeric repeat abundance from HiFi reads)
- count 158,607 reads with pAS1 alignments >1 kb, cumulative 804 Mb, estimated 25.9 Mb non-redundant sequence (pAS1 subtelomeric repeat abundance from HiFi reads)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This bioinformatics paper characterizes sequence gaps in the MorexV3 barley reference assembly through descriptive comparisons of genome size estimates derived from three methods — k-mer frequency spectra (findGSE at k = 21, 51, 101), coverage-based read-depth arithmetic, and published flow cytometric measurements — applied to three sequencing datasets (PacBio HiFi, Oxford Nanopore, Illumina PE450). Repeat-class content (telomeric, subtelomeric, centromeric, ribosomal DNA) was quantified by aligning consensus sequences or annotating tandem repeats in raw reads via BLAST and Tandem Repeat Finder, then comparing cumulative alignment lengths against assembly content. No formal hypothesis tests were employed; all results are reported as descriptive totals (Mb, Gb) and proportions.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| k-mer frequency spectrum analysis (findGSE) at k = 21, 51, 101 | Genome size estimation from HiFi, ONT, and PE450 read sets (Table 1) | 6.6 M HiFi reads (132.7 Gb); 32.5 M ONT reads (426.9 Gb); 684.6 M PE450 reads (364.2 Gb) | not stated |
| Coverage-based genome size estimation (total base pairs sequenced / median read depth) | Genome size estimation from HiFi, ONT, and PE450 read alignments to MorexV3 pseudomolecules (Table 1) | 31x HiFi median depth; 85.2x ONT median depth; 76.2x PE450 median depth | not stated |
| BLAST alignment with threshold-based filtering (≥110 bp for HvT01; ≥330 bp for pAS1; ≥90% alignment length for chloroplast reads) | Quantification of subtelomeric repeat (HvT01, pAS1) and chloroplast DNA content in HiFi reads | 3803 HvT01-positive HiFi reads; 158,607 pAS1-positive HiFi reads; 9.3% ONT reads matching chloroplast | not stated |
| Tandem Repeat Finder (TRF) annotation of tandem and satellite arrays | Detection and sizing of telomeric TTTAGGG arrays in ONT and HiFi reads; trinucleotide microsatellites in ONT reads | 1848 ONT reads with TTTAGGG arrays >1 kb; 452 with arrays >10 kb; 44 HiFi reads with arrays >1 kb | not stated |
| Optical map–to–pseudomolecule alignment (terminal overhang measurement) | Detection of terminal sequence gaps across all seven barley chromosome arms (Table 2, Figure 2) | 7 chromosomes (14 arms) | not stated |
-
Genome size was estimated from k-mer spectra using findGSE; only point estimates were reported for each k-mer size and dataset combination↳ Could also: GenomeScope2 or a bootstrap/parametric confidence interval around the k-mer heterozygosity peak could also quantify uncertainty in each genome size estimate — Reporting uncertainty bounds would make it possible to judge whether apparent discrepancies between estimation methods (e.g., HiFi k=21 giving 2.9 Gb vs. k=101 giving 4.3 Gb) exceed sampling variability or reflect a genuine methodological difference
-
Repeat abundance in raw reads was estimated as cumulative alignment length divided by an assumed uniform genome-wide coverage depth (e.g., 31× for HiFi)↳ Could also: A simulation-based or bootstrap resampling approach that propagates uncertainty in per-locus coverage could also produce interval estimates for total repeat content — Read depth is not perfectly uniform across the genome; accounting for coverage variance would clarify the precision of repeat-content inferences and whether differences between assembly and raw-read estimates are statistically meaningful
-
Genome size estimates from five methods (k-mer at three k values, coverage-based, flow cytometry) were compared by tabulating point estimates side by side↳ Could also: A meta-analytic or Bayesian hierarchical model combining the heterogeneous estimates could also synthesize them into a single posterior estimate with credible intervals — The paper's central goal is to reconcile discrepant size estimates; a formal synthesis would make the degree of concordance explicit and quantify how much unassembled sequence remains, rather than leaving it to qualitative comparison
-
The threshold for classifying a read as repeat-positive was fixed at >1 kb cumulative alignment length↳ Could also: A sensitivity analysis varying the threshold (e.g., 0.5 kb, 2 kb, 5 kb) could also be reported alongside the main estimates — Threshold choice directly affects estimated repeat abundance; showing how totals change across thresholds would demonstrate robustness (or flag sensitivity) of the key quantitative conclusions
-
Completeness of subtelomeric and telomeric sequences was assessed separately per repeat class and per chromosome arm without a unified statistical test of under-representation↳ Could also: A binomial or Poisson test comparing observed vs. expected copy numbers per chromosome arm (derived from coverage-normalised read counts) could also formally test whether each arm's assembly is significantly depleted for a given repeat class — Formal testing of under-representation per locus would complement the descriptive comparisons and identify which specific chromosomal regions are most incomplete, supporting prioritisation for future gap-closing efforts
-
The ONT-based genome size estimate was corrected for chloroplast contamination by a single correction factor derived from the fraction of reads aligning to the chloroplast genome↳ Could also: Uncertainty in the contamination fraction could also be propagated through the genome size calculation using standard error propagation or bootstrapping over the ONT read set — The correction factor itself (9.3% of reads, 29.5 Gb of sequence) is an estimate with sampling variance; propagating that uncertainty into the corrected genome size estimate (4.7 Gb) would indicate how much the contamination correction affects the final conclusion about assembly completeness
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35338551
Paper: Navrátilová et al. 2022, Prospects of telomere-to-telomere assembly in barley: Analysis of sequence gaps in the MorexV3 reference genome. Plant Biotechnol J. PMID 35338551 · PMCID PMC9241371 · DOI 10.1111/pbi.13816.
Code pointer (BRIEF): https://github.com/lh3/seqtk — a third-party tool
(Heng Li's seqtk). Per P16 this is fully valid: we reproduce by running the
described tool/pipeline on the paper's own data. seqtk's role in the paper is the
FASTQ→FASTA conversion that feeds Tandem Repeat Finder (TRF).
Data: raw long reads on ENA:
- HiFi (PacBio)
PRJEB40587— 5 runs, 6.62 M reads, 132.7 Gb (Table 1) ✔ verified via ENA filereport - ONT
PRJEB40588— 22 runs, 32.5 M reads, 426.9 Gb (Table 1) ✔ verified via ENA filereport - PE450 (Illumina)
PRJEB31444; MorexV3 assemblyPRJEB40589
The pipeline we reproduce (in scope)
Methods, verbatim: "Read files in FASTQ format were converted to FASTA format
with seqtk. Tandem repeats were identified with Tandem Repeat Finder (TRF, Benson
(1999)) using the parameter setting 2 5 7 80 10 50 2000 -l 1 -h." Telomeric
arrays are then identified as TTTAGGG (plant telomere) repeat tracts and filtered
by length.
| In-scope result | Reported value | Location | Pipeline |
|---|---|---|---|
| C1 (PRIMARY) HiFi reads: TTTAGGG arrays > 1 kb | 44 arrays | Results / Telomeres | seqtk → TRF → filter motif=TTTAGGG, span>1kb |
| C2 (PRIMARY) HiFi: cumulative length of those arrays | 317 kb | Results / Telomeres | same |
| C3 (stretch) ONT reads with TTTAGGG arrays > 1 kb | 1848 reads | Results / Telomeres | same on ONT |
| C4 (stretch) ONT reads with TTTAGGG arrays > 10 kb | 452 reads | Results / Telomeres | same |
| C5 (stretch) ONT cumulative TTTAGGG (>1 kb) | 13 Mb | Results / Telomeres | same |
| C6 (stretch) ONT reads with trinucleotide arrays > 20 kb | 7655 reads | Results / microsatellites | seqtk → TRF → period-3 filter |
| C7 (stretch) longest AAG array | 50 473 copies / 153 kb | Results / microsatellites | same |
Primary target = C1+C2 (HiFi telomere arrays). Crisp, small, fully specified, and HiFi is the smaller dataset (~126 GB vs ONT ~400 GB). C3–C7 (ONT) are the heavy "last 20%": ~3× the data + TRF time; attempted only if C1/C2 land cleanly with budget to spare.
Out of scope (not a seqtk/TRF pipeline output — wet-lab / optical-mapping /
manual / external-tool, not attempted)
- Bionano optical-map gap spanning (51 of 147 gaps), rDNA unit counts from Bionano (2435 / 829 units; 21.71 / 8.12 Mb) — optical mapping, not in the read→TRF path.
- MorexV3 assembly statistics (4.196 Gb, 147 gaps) — produced by the prior assembly paper (Mascher et al. 2021), not this analysis.
- CENH3 ChIP-seq peak widths, cereba integrase copy estimates, centromere size (3.2 Mb) — separate mapping/annotation + manual interpretation.
- FISH / cytogenetics / flow-cytometry genome-size estimates — wet-lab.
- k-mer genome-size estimates (4.2/4.3 Gb) — different tool (not seqtk/TRF).
Equivalence-preserving optimisation (documented, not a deviation in result)
Running TRF on all 132.7 Gb HiFi for a single motif is wasteful. A genuine >1 kb TTTAGGG tract (≥143 perfect copies) in HiFi (Q≥30) necessarily contains many exact consecutive copies, so we pre-select reads containing ≥3 consecutive canonical telomere units (any rotation, both strands), then run the exact paper TRF parameter string on those reads and filter TRF's own output. This cannot drop any read that holds a >1 kb array; it only removes reads that provably can't. We report the candidate-read count alongside the final array count for transparency.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.