Chromosome-Scale Assembly of the Complete Genome Sequence of Leishmania (Mundinia) enriettii, Isolate CUR178, Strain LV763.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Genome announcement of L. (Mundinia) enriettii CUR178/LV763 via the LGAAP pipeline. All Table 1 assembly + annotation metrics were RECOMPUTED on «our HPC» from the deposited assembly GCA_017916305.1 + MAKER GFF and match EXACTLY (genome 33,318,864 bp; N50 1,075,649; 54 scaffolds=36 chr+18 unplaced; 8,353 genes; 8,584 exons; CDS 15.46 Mb/46.40%; density 250.7; mean gene 1,897.5); GC 59.57% vs 59.60% within tol. Illumina read counts EXACT (MiSeq 5,060,124; HiSeq 20,936,270). THREE auditable discrepancies in the read/base bookkeeping: MinION reads 793,030 reported vs 719,639 deposited; total reads off by that 73,391; and total bases 19.41 Gb reported vs 8.99 Gb deposited — the latter inconsistent with the paper's OWN 271.8x coverage (=>~9.06 Gb), a likely reporting error. BUSCO re-run 99.2% (129/130) vs reported 94.6% — same near-complete conclusion, numeric diff from BUSCO version/lineage. Flye 2.8.2 (LGAAP step 1) crashed in consensus (known 2.8.x EOFError); the Flye 2.9.5 re-run COMPLETED and recovers 32.81 Mb in 154 contigs (N50 664 kb) = 98.5% of genome size de-novo from the deposited Nanopore reads alone, corroborating the assembly (graded within-tol; final 36-chr contiguity comes from downstream RaGOO+Pilon, not run end-to-end). NOT attempted: full 314-step LGAAP Snakemake (Pilon/RaGOO/RepeatMasker/MAKER) end-to-end — verification was against the deposited authoritative artifacts plus the de-novo Flye contiguity check. FINAL: 13 exact + 2 within-tol + 2 partial + 3 mismatch; the assembly+annotation (the paper's substance) reproduce 1:1, discrepancies confined to the Table 1 read/base bookkeeping.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 92assessed: 2026-06-18 ⛓ 8d9891e20640
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnet- ★ Complete chromosome-scale genome sequence of Leishmania (Mundinia) enriettii isolate CUR178, strain LV763 was assembled using combined short-read and long-read sequencing resource
- ★ The new assembly improves on the previous L. (M.) enriettii genome (isolate LEM3045) which had many unplaced contigs, higher gap content, and lower N50 finding
- ★ All 36 chromosomes were aligned using L. major Friedlin genome as reference guide for scaffolding, with chromosome ends determined complete except for 18 unplaced contigs finding
- ★ Genome assembly achieved 94.6% BUSCO completeness using Euglenozoa lineage dataset finding
- Functional annotation was performed using MAKER2 combined with AUGUSTUS trained on Leishmania tarentolae method
- An in vitro culture system originally developed for Leishmania orientalis axenic amastigotes was used to grow L. enriettii parasites method
- A reproducible Snakemake-based workflow (LGAAP) was used for assembly, repeat masking, and annotation resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-genome short-read sequencing (DNBSEQ) | Leishmania (Mundinia) enriettii, isolate CUR178, strain LV763 | none | paired-end sequence reads (270bp, 500bp) | Illumina HiSeq (BGI) |
| Whole-genome short-read sequencing (TruSeq Nano) | Leishmania (Mundinia) enriettii, isolate CUR178, strain LV763 | none | paired-end sequence reads (300bp) | Illumina MiSeq (Aberystwyth University) |
| Long-read whole-genome sequencing | Leishmania (Mundinia) enriettii, isolate CUR178, strain LV763 | none | long reads for chromosome-scale scaffold assembly | Oxford Nanopore, SQK-LSK109 protocol, R9 flow cells (FLO-MIN106) |
| Genome assembly and polishing | Leishmania (Mundinia) enriettii, isolate CUR178, strain LV763 | none | chromosome-scale genome assembly | Flye, Minimap2, SAMtools, Pilon, Funannotate, RaGOO |
| Genome completeness assessment | Leishmania (Mundinia) enriettii, isolate CUR178, strain LV763 | none | single-copy ortholog completeness (%) | BUSCO (Euglenozoa lineage dataset, 130 orthologs, 31 species) |
| Functional gene annotation | Leishmania (Mundinia) enriettii, isolate CUR178, strain LV763 | none | predicted gene models, gene counts, exon counts | MAKER2 with AUGUSTUS (trained on Leishmania tarentolae) |
| Read quality assessment | Leishmania (Mundinia) enriettii, isolate CUR178, strain LV763 | none | sequencing read quality metrics | MultiQC |
- – BUSCO completeness of 123/130 single-copy orthologs identified 94.6%
- – Total genome size assembled 33,318,864 bp
- – Assembly N50 value 1,075,649 bp
- – Total number of reads generated across all platforms 26,789,424 reads
- – Total scaffolds in final assembly with 18 unplaced contigs 54 scaffolds; 76,607 bp unplaced
- – Number of predicted genes annotated 8,353 genes
- – Genome coverage achieved 271.8x
- – GC content of genome 59.60%
- count 26,789,424 total reads (combined MiSeq, HiSeq, and MinION reads)
- count 5,060,124 MiSeq reads (short-read sequencing)
- count 20,936,270 HiSeq reads (short-read sequencing)
- count 793,030 MinION reads (N50 12,070 bp) (long-read sequencing)
- other 19.41 Gb total bases, 271.8x coverage (sequencing depth)
- count 8,353 genes; 8,584 exons; mean gene length 1,897 bp (annotation summary)
- other 94.6% BUSCO completeness (123/130 orthologs) (Euglenozoa lineage dataset assembly completeness)
- other CDS total length 15.46 Mb (46.40% of genome) (coding sequence content)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a genome resource announcement reporting the chromosome-scale assembly and annotation of Leishmania (Mundinia) enriettii isolate CUR178 using hybrid short-read (Illumina MiSeq and HiSeq) and long-read (Oxford Nanopore MinION) sequencing. No inferential statistics were performed; results are presented entirely as descriptive bioinformatic metrics (N50, genome size, GC content, gene counts, BUSCO completeness percentage). An assembly comparison figure presents metrics for three genomes side by side without formal statistical testing.
-
Long reads were assembled with Flye using default parameters↳ Could also: Canu or wtdbg2 (Redbean) could also have been used for de novo long-read assembly from Nanopore data — Different assemblers use distinct graph algorithms (string-overlap graph in Canu vs. repeat graph in Flye); benchmarking two or more assemblers and selecting the best by N50 and BUSCO is a common approach that increases confidence in the chosen assembly
-
Polishing used Pilon with short-read alignments produced by Minimap2↳ Could also: Medaka (a Nanopore-native neural-network polisher) could also have been applied before Pilon as a first polishing pass — A two-stage strategy—Nanopore-aware polishing first (Medaka), then Illumina-based correction (Pilon)—is increasingly used to address platform-specific error profiles sequentially before hybrid correction, which can further reduce residual indels
-
Assembly completeness was assessed with BUSCO using the Euglenozoa lineage dataset (130 single-copy orthologs from 31 species)↳ Could also: Merqury (k-mer-based quality value and completeness scoring derived directly from the short reads) could also have been used alongside BUSCO — BUSCO measures conserved gene presence while Merqury provides a reference-free, k-mer-based quality value (QV) estimating base-level accuracy; the two metrics are complementary and their combined use is increasingly standard for hybrid assemblies
-
Reference-guided scaffolding was performed with RaGOO using the L. major Friedlin genome as a guide↳ Could also: ALLMAPS or, if Hi-C data were available, 3D-DNA or Salsa2 could also have been used for scaffolding — Reference-guided scaffolding relies on conserved synteny with a related organism; chromosome conformation capture (Hi-C)-based approaches generate scaffolding evidence from the target organism's own chromatin structure without assuming synteny conservation, which may be incomplete across the Leishmania subgenus boundary
-
Gene prediction used AUGUSTUS trained on L. tarentolae as part of the MAKER2 pipeline↳ Could also: BRAKER2 could also have been used, incorporating RNA-seq alignments to supplement or replace the cross-species ab initio predictor — Ab initio prediction trained on a related but distinct species may miss organism-specific gene models; integrating transcriptomic evidence from the target organism generally improves both gene model sensitivity and specificity
-
Cross-assembly comparison (Figure 1) was presented as a visual display of metrics without a standardized reporting framework↳ Could also: QUAST could also have been used to generate a systematic, tool-standardized multi-metric comparison report across the three assemblies — QUAST produces a reproducible, normalized set of assembly quality statistics (NGA50, misassembly counts, indel rates relative to a reference, etc.) in a single report, making cross-assembly comparisons more consistent and easier to interpret alongside metrics like BUSCO
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Every core assembly and annotation claim in Table 1 reproduces exactly against the authoritative deposited GenBank assembly GCA_017916305.1 and the authors' MAKER GFF (genome size, N50, 36 chromosomes, 8,353 genes, CDS 46.40%, gene density 250.7), with GC content within tolerance (59.60% vs 59.57%), so the central chromosome-scale-assembly claim fully holds. The one notable defect sits on the authors' side: the reported sequencing total of 19.41 Gb is not derivable from the deposited reads (~8.99 Gb) and is internally inconsistent with the paper's own 271.8x coverage (~9.06 Gb), and the MinION read count (793,030 vs deposited 719,639) also mismatches. This is a peripheral input-description anomaly, not a failure of the genome itself, so severity is moderate and overall quality is solid-with-explainable-deviation rather than 1:1.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.