Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Prospects of telomere-to-telomere assembly in barley: Analysis of sequence gaps in the MorexV3 reference genome.

Plant Biotechnol J · 2022
71/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
71/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 38% of all assessed papers rank 694 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED. Ran the paper's exact seqtk->TRF read-level repeat pipeline (seqtk 1.5-r133, TRF 4.10.0rc2, params '2 5 7 80 10 50 2000 -l 1 -h') on the paper's OWN long-read data. PRIMARY HiFi claims EXACT: C1 = 44 TTTAGGG arrays >1 kb (44/44), C2 = 317,475 bp ~ 317 kb («job», HiFi N=6,617,223 = Table 1). ONT stretch (PRJEB40588, all 22 runs, observed N=30,517,372 ~ 94% of Table 1's 32.5M; «job»): C7 the single most specific number in the paper reproduced EXACTLY -> longest trinucleotide array 50,473 AAG copies spanning 152,808 bp (153 kb), read ERR4659224.385782. C3=1646/1848 (89%), C4=408/452 (90%), C5=11.7/13 Mb (90%), C6=6272/7655 (82%, lower bound) -- all within ~10-18% of reported, the residual explained by observed-vs-reported read count and the cancelled trinucleotide-TRF tail. NO fabrication indicators: every value is derivable from the public deposit + named tools, and the two precise claims (C1/C2, C7) land on the nose. NOT attempted (out of scope, not a seqtk/TRF read pipeline): Bionano optical mapping, CENH3 ChIP-seq, cytogenetics/FISH, k-mer genome-size, MorexV3 assembly statistics. Infra notes: shared-account «infra»+/home quota exhaustion forced a checkpoint+resume of the ONT stream; data all on «infra», results only on «host».

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 70
    assessed: 2026-06-20 ⛓ 580e9171d2bf
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper investigates whether and how a telomere-to-telomere (T2T) assembly of the barley genome could be achieved by identifying what sequences are still missing from the current MorexV3 reference assembly and why gaps remain.

Core claims
  • Almost all centromeric sequences and 45S ribosomal DNA repeat arrays are absent from the MorexV3 pseudomolecules finding
  • The majority of sequence gaps in MorexV3 can be attributed to assembly breakdown in long stretches of satellite repeats finding
  • Missing sequences cannot fully account for the difference between MorexV3 assembly size and flow cytometric genome size estimates finding
  • Telomeric satellite arrays (TTTAGGG) are not captured in their entirety by long reads (HiFi or ONT) finding
  • Subtelomeric repeat arrays (HvT01, pAS1) are disrupted by sequence gaps in the assembly finding
  • ONT sequencing data used for MorexV3 were contaminated with chloroplast DNA, inflating raw coverage-based genome size estimates finding
  • Ultra-long sequence reads (>100 kb to 1 Mb) will likely be required to assemble centromeric and rDNA loci mechanism
Experimental setups
Assay System Perturbation Readout Platform
Optical genome mapping (Bionano) barley cv. Morex none missing sequence at chromosome termini relative to pseudomolecules Bionano optical map, DLE-1 labelling
ChIP-seq barley cv. Morex, centromeric chromatin none centromeric histone variant CENH3 binding / centromeric repeat abundance ChIP-seq
k-mer spectrum genome size estimation barley cv. Morex genomic DNA none genome size estimate (GSE) at k=21,51,101 findGSE on PacBio HiFi, ONT, and PE450 Illumina reads
Read-depth based genome size estimation barley cv. Morex none genome size estimate from mean/median coverage alignment of HiFi/ONT/PE450 reads to MorexV3 pseudomolecules
Fluorescence in situ hybridization (FISH) barley cv. Morex metaphase chromosomes none localization of chloroplast DNA insertions, subtelomeric satellite HvT01, and 45S rDNA FISH with cpDNA, HvT01, and 45S rDNA probes
Tandem Repeat Finder (TRF) annotation barley cv. Morex ONT long reads none length and location of TTTAGGG telomeric repeat arrays Tandem Repeat Finder
BLAST sequence alignment barley cv. Morex HiFi long reads none abundance/length of subtelomeric repeats HvT01 and pAS1 BLAST
Flow cytometry (literature values used for comparison) barley cv. Morex nuclei none haploid (1C) genome size flow cytometry
Key results
  • MorexV3 pseudomolecule assembly totals 4.196 Gb (plus 29.1 Mb unplaced contigs), below flow cytometric estimates of 4.88–5.04 Gb ~0.7-0.85 Gb difference
  • Missing sequence at chromosome termini detected via optical map overhangs 17-220 kb (short arms), 10-80 kb (long arms)
  • ONT reads show TTTAGGG telomeric arrays >1kb totaling 13 Mb cumulative, averaging ~11 kb per chromosome arm; HiFi captured far less HiFi: 741 bp per telomere vs ONT ~11 kb per arm
  • HvT01 subtelomeric repeat abundance estimated from HiFi reads 3.1 Mb total, 225 kb per chromosome arm
  • pAS1 subtelomeric repeat abundance estimated from HiFi reads 25.9 Mb total, 1.85 Mb per chromosome arm
  • 9.3% of ONT reads aligned to chloroplast genome, indicating cpDNA contamination rather than true nuclear insertions 29.5 Gb of ONT reads matched cpDNA
  • After correcting for cpDNA contamination, ONT coverage-based genome size estimate revised 5 Gb to 4.7 Gb
  • HvT01 sequence disproportionately assigned to unanchored contigs compared to pAS1 19.3% (707 kb) of HvT01 vs 0.5% (134 kb) of pAS1
Key statistics
  • count 4.196 Gb pseudomolecules + 29.1 Mb unplaced contigs (MorexV3 assembly size)
  • mean 4.88 and 5.04 Gb (1C genome size) (flow cytometric estimate by Doležel et al. 2018)
  • other 3.559–5.349 Gb (range of published flow cytometric 1C genome size estimates for H. vulgare)
  • fold_change GSE at k=51: HiFi 4.2 Gb, PE450 4.3 Gb (k-mer based genome size estimates)
  • count 1848 ONT reads with TTTAGGG arrays >1 kb (452 with >10 kb); average array size 6.9 kb; longest 37.2 kb; cumulative 13 Mb (telomeric repeat quantification in ONT reads)
  • count 44 HiFi reads with TTTAGGG arrays >1 kb, cumulative 317 kb (741 bp per telomere at 31x coverage) (telomeric repeat quantification in HiFi reads)
  • count 3803 reads with HvT01 alignments ≥1 kb, cumulative 97.6 Mb, estimated 3.1 Mb total array size (HvT01 subtelomeric repeat abundance from HiFi reads)
  • count 158,607 reads with pAS1 alignments >1 kb, cumulative 804 Mb, estimated 25.9 Mb non-redundant sequence (pAS1 subtelomeric repeat abundance from HiFi reads)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This bioinformatics paper characterizes sequence gaps in the MorexV3 barley reference assembly through descriptive comparisons of genome size estimates derived from three methods — k-mer frequency spectra (findGSE at k = 21, 51, 101), coverage-based read-depth arithmetic, and published flow cytometric measurements — applied to three sequencing datasets (PacBio HiFi, Oxford Nanopore, Illumina PE450). Repeat-class content (telomeric, subtelomeric, centromeric, ribosomal DNA) was quantified by aligning consensus sequences or annotating tandem repeats in raw reads via BLAST and Tandem Repeat Finder, then comparing cumulative alignment lengths against assembly content. No formal hypothesis tests were employed; all results are reported as descriptive totals (Mb, Gb) and proportions.

Replicationunclear Sample sizeDatasets described by total read count, total base pairs, and mean/median/mode coverage depth (Table 1); no formal power analysis or sample size justification reported GroupsThree sequencing datasets (PacBio HiFi, ONT, Illumina PE450) compared against each other and against MorexV3 assembly size and published flow cytometric genome size estimates for a single barley cultivar (Morex) Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
k-mer frequency spectrum analysis (findGSE) at k = 21, 51, 101 Genome size estimation from HiFi, ONT, and PE450 read sets (Table 1) 6.6 M HiFi reads (132.7 Gb); 32.5 M ONT reads (426.9 Gb); 684.6 M PE450 reads (364.2 Gb) not stated
Coverage-based genome size estimation (total base pairs sequenced / median read depth) Genome size estimation from HiFi, ONT, and PE450 read alignments to MorexV3 pseudomolecules (Table 1) 31x HiFi median depth; 85.2x ONT median depth; 76.2x PE450 median depth not stated
BLAST alignment with threshold-based filtering (≥110 bp for HvT01; ≥330 bp for pAS1; ≥90% alignment length for chloroplast reads) Quantification of subtelomeric repeat (HvT01, pAS1) and chloroplast DNA content in HiFi reads 3803 HvT01-positive HiFi reads; 158,607 pAS1-positive HiFi reads; 9.3% ONT reads matching chloroplast not stated
Tandem Repeat Finder (TRF) annotation of tandem and satellite arrays Detection and sizing of telomeric TTTAGGG arrays in ONT and HiFi reads; trinucleotide microsatellites in ONT reads 1848 ONT reads with TTTAGGG arrays >1 kb; 452 with arrays >10 kb; 44 HiFi reads with arrays >1 kb not stated
Optical map–to–pseudomolecule alignment (terminal overhang measurement) Detection of terminal sequence gaps across all seven barley chromosome arms (Table 2, Figure 2) 7 chromosomes (14 arms) not stated
Approaches that could also have been used
  • Genome size was estimated from k-mer spectra using findGSE; only point estimates were reported for each k-mer size and dataset combination
    Could also: GenomeScope2 or a bootstrap/parametric confidence interval around the k-mer heterozygosity peak could also quantify uncertainty in each genome size estimate — Reporting uncertainty bounds would make it possible to judge whether apparent discrepancies between estimation methods (e.g., HiFi k=21 giving 2.9 Gb vs. k=101 giving 4.3 Gb) exceed sampling variability or reflect a genuine methodological difference
  • Repeat abundance in raw reads was estimated as cumulative alignment length divided by an assumed uniform genome-wide coverage depth (e.g., 31× for HiFi)
    Could also: A simulation-based or bootstrap resampling approach that propagates uncertainty in per-locus coverage could also produce interval estimates for total repeat content — Read depth is not perfectly uniform across the genome; accounting for coverage variance would clarify the precision of repeat-content inferences and whether differences between assembly and raw-read estimates are statistically meaningful
  • Genome size estimates from five methods (k-mer at three k values, coverage-based, flow cytometry) were compared by tabulating point estimates side by side
    Could also: A meta-analytic or Bayesian hierarchical model combining the heterogeneous estimates could also synthesize them into a single posterior estimate with credible intervals — The paper's central goal is to reconcile discrepant size estimates; a formal synthesis would make the degree of concordance explicit and quantify how much unassembled sequence remains, rather than leaving it to qualitative comparison
  • The threshold for classifying a read as repeat-positive was fixed at >1 kb cumulative alignment length
    Could also: A sensitivity analysis varying the threshold (e.g., 0.5 kb, 2 kb, 5 kb) could also be reported alongside the main estimates — Threshold choice directly affects estimated repeat abundance; showing how totals change across thresholds would demonstrate robustness (or flag sensitivity) of the key quantitative conclusions
  • Completeness of subtelomeric and telomeric sequences was assessed separately per repeat class and per chromosome arm without a unified statistical test of under-representation
    Could also: A binomial or Poisson test comparing observed vs. expected copy numbers per chromosome arm (derived from coverage-normalised read counts) could also formally test whether each arm's assembly is significantly depleted for a given repeat class — Formal testing of under-representation per locus would complement the descriptive comparisons and identify which specific chromosomal regions are most incomplete, supporting prioritisation for future gap-closing efforts
  • The ONT-based genome size estimate was corrected for chloroplast contamination by a single correction factor derived from the fraction of reads aligning to the chloroplast genome
    Could also: Uncertainty in the contamination fraction could also be propagated through the genome size calculation using standard error propagation or bootstrapping over the ONT read set — The correction factor itself (9.3% of reads, 29.5 Gb of sequence) is an estimate with sampling variance; propagating that uncertainty into the corrected genome size estimate (4.7 Gb) would indicate how much the contamination correction affects the final conclusion about assembly completeness
Software: findGSE · Tandem Repeat Finder (TRF) · BLAST · GenomeScope

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35338551

Paper: Navrátilová et al. 2022, Prospects of telomere-to-telomere assembly in barley: Analysis of sequence gaps in the MorexV3 reference genome. Plant Biotechnol J. PMID 35338551 · PMCID PMC9241371 · DOI 10.1111/pbi.13816.

Code pointer (BRIEF): https://github.com/lh3/seqtk — a third-party tool (Heng Li's seqtk). Per P16 this is fully valid: we reproduce by running the described tool/pipeline on the paper's own data. seqtk's role in the paper is the FASTQ→FASTA conversion that feeds Tandem Repeat Finder (TRF).

Data: raw long reads on ENA:

  • HiFi (PacBio) PRJEB40587 — 5 runs, 6.62 M reads, 132.7 Gb (Table 1) ✔ verified via ENA filereport
  • ONT PRJEB40588 — 22 runs, 32.5 M reads, 426.9 Gb (Table 1) ✔ verified via ENA filereport
  • PE450 (Illumina) PRJEB31444; MorexV3 assembly PRJEB40589

The pipeline we reproduce (in scope)

Methods, verbatim: "Read files in FASTQ format were converted to FASTA format with seqtk. Tandem repeats were identified with Tandem Repeat Finder (TRF, Benson (1999)) using the parameter setting 2 5 7 80 10 50 2000 -l 1 -h." Telomeric arrays are then identified as TTTAGGG (plant telomere) repeat tracts and filtered by length.

In-scope result Reported value Location Pipeline
C1 (PRIMARY) HiFi reads: TTTAGGG arrays > 1 kb 44 arrays Results / Telomeres seqtk → TRF → filter motif=TTTAGGG, span>1kb
C2 (PRIMARY) HiFi: cumulative length of those arrays 317 kb Results / Telomeres same
C3 (stretch) ONT reads with TTTAGGG arrays > 1 kb 1848 reads Results / Telomeres same on ONT
C4 (stretch) ONT reads with TTTAGGG arrays > 10 kb 452 reads Results / Telomeres same
C5 (stretch) ONT cumulative TTTAGGG (>1 kb) 13 Mb Results / Telomeres same
C6 (stretch) ONT reads with trinucleotide arrays > 20 kb 7655 reads Results / microsatellites seqtk → TRF → period-3 filter
C7 (stretch) longest AAG array 50 473 copies / 153 kb Results / microsatellites same

Primary target = C1+C2 (HiFi telomere arrays). Crisp, small, fully specified, and HiFi is the smaller dataset (~126 GB vs ONT ~400 GB). C3–C7 (ONT) are the heavy "last 20%": ~3× the data + TRF time; attempted only if C1/C2 land cleanly with budget to spare.

Out of scope (not a seqtk/TRF pipeline output — wet-lab / optical-mapping /

manual / external-tool, not attempted)

  • Bionano optical-map gap spanning (51 of 147 gaps), rDNA unit counts from Bionano (2435 / 829 units; 21.71 / 8.12 Mb) — optical mapping, not in the read→TRF path.
  • MorexV3 assembly statistics (4.196 Gb, 147 gaps) — produced by the prior assembly paper (Mascher et al. 2021), not this analysis.
  • CENH3 ChIP-seq peak widths, cereba integrase copy estimates, centromere size (3.2 Mb) — separate mapping/annotation + manual interpretation.
  • FISH / cytogenetics / flow-cytometry genome-size estimates — wet-lab.
  • k-mer genome-size estimates (4.2/4.3 Gb) — different tool (not seqtk/TRF).

Equivalence-preserving optimisation (documented, not a deviation in result)

Running TRF on all 132.7 Gb HiFi for a single motif is wasteful. A genuine >1 kb TTTAGGG tract (≥143 perfect copies) in HiFi (Q≥30) necessarily contains many exact consecutive copies, so we pre-select reads containing ≥3 consecutive canonical telomere units (any rotation, both strands), then run the exact paper TRF parameter string on those reads and filter TRF's own output. This cannot drop any read that holds a >1 kb array; it only removes reads that provably can't. We report the candidate-read count alongside the final array count for transparency.

C1
Reported
44 HiFi TTTAGGG telomeric arrays >1 kb
Reproduced
44
exact
C2
Reported
317 kb cumulative length of HiFi TTTAGGG arrays >1 kb
Reproduced
317475 bp (317 kb)
exact
C3
Reported
1848 ONT reads with TTTAGGG arrays >1 kb
Reproduced
1646
partial
C4
Reported
452 ONT reads with TTTAGGG arrays >10 kb
Reproduced
408
partial
C5
Reported
13 Mb cumulative TTTAGGG arrays >1 kb (ONT)
Reproduced
11696261 bp (11.7 Mb)
partial
C6
Reported
7655 ONT reads with trinucleotide arrays >20 kb
Reproduced
6272 (lower bound)
partial
C7
Reported
longest trinucleotide array 50473 AAG copies / 153 kb
Reproduced
50473 copies / 152808 bp (153 kb), AAG class
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

853.5 k
tokens (I/O) · 59.5 M incl. cache
546 min
runtime · 574.47 CPU-h
13.7 GB
peak RAM
4 (2 failed)
HPC jobs
hummel
machine