Chromosome-Scale Assembly of the Complete Genome Sequence of Leishmania (Mundinia) orientalis, Isolate LSCM4, Strain LV768.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
1:1 REPRODUCED the central computational claim of this genome-announcement MRA. Ran the exact LGAAP Flye 2.8.2 de novo assembly step (flye --nano-raw SRR13558782 --genome-size 35m) on «our HPC» against the paper's own PRJNA691532 reads. Reproduced raw-assembly QUAST matches the repo's shipped Flye QUAST within 0.15% total length (34,428,603 vs 34,478,885 bp), 1 contig (169 vs 170), identical GC (59.73 vs 59.72) and L50 (12), N50 within 1.45%. Genome size matches the deposited paper genome GCA_017916335.1 (34,194,276 bp) within 0.69%; seqkit on the deposit independently confirms the paper headline (98 scaffolds, N50 1,120,138, GC 59.72, 36 chr). Read accounting: Illumina HiSeq 76,124,074 and MiSeq 3,831,060 match the paper EXACTLY; Nanopore deposited fastq is 543,781 vs reported 585,770 (~7% short -> mismatch/audit flag). NOT attempted (honest scope limits): Pilon Illumina polishing (repo shows negligible change), RaGOO chromosome-scaffolding (ragoo=1.1 deprecated; repo's shipped final 43-seq/32.47Mb QUAST does not even equal the deposited 98-scaffold/34.19Mb assembly -> audit flag), MAKER/funannotate annotation + BUSCO (~30h, out of scope). Verdict: de novo assembly reproduces faithfully; downstream scaffolding/annotation not reproduced; two audit flags (Nanopore read deficit; repo-shipped-final != deposited-final).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 70assessed: 2026-06-18 ⛓ 955f6e47ee74
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnet- ★ The complete genome sequence of Leishmania (Mundinia) orientalis, isolate LSCM4, strain LV768, was determined using combined short-read and long-read sequencing. finding
- ★ The assembly achieved chromosome-scale resolution, aligning all 36 chromosomes using the L. major Friedlin genome as a reference guide. finding
- ★ The genome assembly is 98.5% complete based on BUSCO analysis of Euglenozoa single-copy orthologs. finding
- ★ A reproducible Snakemake-based workflow (LGAAP) was used for assembly, repeat masking, and annotation, and is publicly available on GitHub. resource
- Functional annotation was performed using MAKER2 combined with AUGUSTUS, trained on Leishmania tarentolae. method
- ★ This genome sequence will facilitate greater understanding of L. (M.) orientalis and its relationship to other members of the subgenus Mundinia. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| in vitro culture | Leishmania (Mundinia) orientalis promastigotes/axenic amastigotes, isolate LSCM4 strain LV768 | none | sustained parasite growth and viability | — |
| DNA extraction/purification | L. (M.) orientalis cultured parasites | none | extracted DNA concentration and quality | Qiagen DNeasy blood and tissue kit; Qubit fluorometer, microplate reader, agarose gel electrophoresis |
| short-read whole-genome sequencing (DNBSEQ) | L. (M.) orientalis genomic DNA | none | paired-end sequence reads (170/270/500 bp) | Illumina HiSeq (BGI) |
| short-read whole-genome sequencing (TruSeq Nano) | L. (M.) orientalis genomic DNA | none | paired-end sequence reads (300 bp) | Illumina MiSeq (Aberystwyth University) |
| long-read whole-genome sequencing | L. (M.) orientalis genomic DNA | none | long sequence reads for scaffold assembly | Oxford Nanopore, SQK-LSK109 kit, R9 flow cells (FLO-MIN106), MinION |
| genome assembly and polishing (bioinformatics) | L. (M.) orientalis sequence reads | none | chromosome-scale genome assembly | Flye, Minimap2, SAMtools, Pilon, Funannotate, RaGOO, Snakemake |
| assembly completeness assessment (BUSCO) | L. (M.) orientalis genome assembly | none | proportion of conserved single-copy orthologs present | BUSCO, Euglenozoa lineage dataset (130 orthologs, 31 species) |
| genome annotation | L. (M.) orientalis genome assembly | none | predicted gene models, exons, CDS content | MAKER2 with AUGUSTUS (trained on L. tarentolae) |
- – Total sequencing reads generated 80,540,904 reads
- – Genome coverage achieved 390.7x
- – Final assembled genome size 34,194,276 bp
- – Assembly scaffold N50 1,120,138 bp
- – All 36 chromosomes aligned to L. major reference; 62 small unplaced contigs remained 257,579 bp in 62 contigs
- – BUSCO orthologs present 128/130 (98.5%)
- – Predicted gene count and density 8,158 genes; 238.6 genes/Mb
- – Coding sequence content of genome 15.40 Mb (45.05% of genome)
- count 80,540,904 total reads (3,831,060 MiSeq; 76,124,074 HiSeq; 585,770 MinION, read N50 11,497 bp) (sequencing read totals by platform)
- other 29.20 Gb total bases; 390.7x genome coverage (sequencing depth)
- other genome size 34,194,276 bp; 98 total scaffolds; N50 1,120,138 bp (assembly metrics)
- other GC content 59.70%; 1,707 Ns (0.005% of genome) (genome composition)
- other BUSCO completeness 128/130 orthologs (98.5%) (assembly completeness assessment (Euglenozoa lineage))
- count 8,158 genes; 8,488 exons; mean gene length 1,938 bp (genome annotation summary)
- other gene density 238.6 genes/Mb; total CDS length 15.40 Mb (45.05% of genome) (annotation density and coding content)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a genome resource announcement describing the sequencing, assembly, and annotation of Leishmania (Mundinia) orientalis isolate LSCM4/LV768. The study employed a hybrid sequencing strategy combining Illumina short reads (DNBSEQ and TruSeq Nano libraries) with Oxford Nanopore long reads. Assembly completeness was assessed using BUSCO against the Euglenozoa lineage dataset (98.5% of 130 single-copy orthologs recovered). No inferential statistical tests were performed; all reporting consists of descriptive assembly metrics.
-
Assembly completeness was assessed using BUSCO with the Euglenozoa lineage dataset (130 single-copy orthologs from 31 species)↳ Could also: A more specific lineage dataset (e.g., Kinetoplastida or Trypanosomatida if available in the BUSCO database version used) could also be applied, or CEGMA (Core Eukaryotic Genes Mapping Approach) could serve as a complementary completeness benchmark — A lineage dataset closer to the target organism captures genes more representative of the clade, potentially giving a more sensitive or specific completeness estimate; reporting both gives a fuller picture
-
Long reads were assembled with Flye using default parameters, then polished with short reads via Pilon in two rounds↳ Could also: Alternative hybrid assemblers such as MaSuRCA or Canu (for the long-read primary assembly) combined with multiple rounds of Medaka (for nanopore-specific error correction before Pilon) could also be used — Different assemblers optimize differently for repeat resolution and contig contiguity; benchmarking two assemblers and selecting the better N50/BUSCO result is common practice in genome announcement papers
-
Reference-guided scaffolding was performed using RaGOO with L. major Friedlin as the guide genome↳ Could also: Tools such as Chromosomer, ALLMAPS, or 3D-DNA (if Hi-C data were available) could also be used for scaffolding; or a more closely related Mundinia genome (e.g., L. enriettii) could serve as the guide if available at sufficient quality — Using a more phylogenetically proximate reference may reduce scaffolding artifacts arising from synteny breaks; Hi-C-based scaffolding would be fully reference-independent and is increasingly standard for chromosome-scale assemblies
-
Gene prediction was performed with MAKER2 in combination with AUGUSTUS trained on L. tarentolae↳ Could also: Training AUGUSTUS directly on L. orientalis (using a subset of high-confidence gene models) or using additional ab initio predictors (e.g., GeneMark-ES, GlimmerHMM) within a MAKER2 or BRAKER2 pipeline could also be applied — Using a trainer species from a different Leishmania subgenus introduces some phylogenetic distance; self-training or using a closer relative may improve sensitivity and specificity of gene models, and multi-predictor consensus is a common approach to reduce false positives
-
Assembly quality beyond BUSCO was summarized with descriptive metrics only (N50, genome size, GC content, number of Ns)↳ Could also: Genome assembly quality assessors such as QUAST (for contig-level statistics and misassembly detection relative to the reference) or Merqury (k-mer-based quality value and completeness estimation from the raw reads) could also be reported — BUSCO measures gene-space completeness but does not detect misassemblies or structural errors; QUAST and Merqury provide complementary perspectives on base-level accuracy and structural correctness that are increasingly expected in genome announcements
-
The genome assembly of a single isolate was reported without replicate sequencing runs or assessment of within-run variability↳ Could also: Reporting read-depth uniformity across chromosomes (e.g., coverage plots), or comparing the assembly to an independent short-read-only assembly as a sanity check, could also be included — Leishmania exhibits aneuploidy and copy-number variation among chromosomes; coverage uniformity plots are commonly included in Leishmania genome papers to document the somy status of each chromosome and validate the assembly
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Illumina read-count accounting reproduces exactly (HiSeq 76,124,074; MiSeq 3,831,060), but the Nanopore deposit is short one empty run (SRR13558783) and total deposited bases (~12.83 Gb) are ~2.3x below the paper's claimed 29.20 Gb. More seriously, the authors' own LGAAP repo ships a final QUAST of 32,471,852 bp / 43 contigs that contradicts the paper's 34,194,276 bp / 98 scaffolds — an authors-side inconsistency, not a method choice on our part. Genome-size magnitude is preserved (~32-34 Mb), so severity is moderate, but the exact published assembly metrics are not derivable from the shared artifacts. Reproduction is still preliminary (Flye assembly pending), so judgements are documented rather than final.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.