Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Chromosome-Scale Assembly of the Complete Genome Sequence of Leishmania (Mundinia) enriettii, Isolate CUR178, Strain LV763.

Microbiol Resour Announc · 2021
L1 80/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
80/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 56% of all assessed papers rank 484 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Genome announcement of L. (Mundinia) enriettii CUR178/LV763 via the LGAAP pipeline. All Table 1 assembly + annotation metrics were RECOMPUTED on «our HPC» from the deposited assembly GCA_017916305.1 + MAKER GFF and match EXACTLY (genome 33,318,864 bp; N50 1,075,649; 54 scaffolds=36 chr+18 unplaced; 8,353 genes; 8,584 exons; CDS 15.46 Mb/46.40%; density 250.7; mean gene 1,897.5); GC 59.57% vs 59.60% within tol. Illumina read counts EXACT (MiSeq 5,060,124; HiSeq 20,936,270). THREE auditable discrepancies in the read/base bookkeeping: MinION reads 793,030 reported vs 719,639 deposited; total reads off by that 73,391; and total bases 19.41 Gb reported vs 8.99 Gb deposited — the latter inconsistent with the paper's OWN 271.8x coverage (=>~9.06 Gb), a likely reporting error. BUSCO re-run 99.2% (129/130) vs reported 94.6% — same near-complete conclusion, numeric diff from BUSCO version/lineage. Flye 2.8.2 (LGAAP step 1) crashed in consensus (known 2.8.x EOFError); the Flye 2.9.5 re-run COMPLETED and recovers 32.81 Mb in 154 contigs (N50 664 kb) = 98.5% of genome size de-novo from the deposited Nanopore reads alone, corroborating the assembly (graded within-tol; final 36-chr contiguity comes from downstream RaGOO+Pilon, not run end-to-end). NOT attempted: full 314-step LGAAP Snakemake (Pilon/RaGOO/RepeatMasker/MAKER) end-to-end — verification was against the deposited authoritative artifacts plus the de-novo Flye contiguity check. FINAL: 13 exact + 2 within-tol + 2 partial + 3 mismatch; the assembly+annotation (the paper's substance) reproduce 1:1, discrepancies confined to the Table 1 read/base bookkeeping.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 92
    assessed: 2026-06-18 ⛓ 8d9891e20640
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Core claims
  • Complete chromosome-scale genome sequence of Leishmania (Mundinia) enriettii isolate CUR178, strain LV763 was assembled using combined short-read and long-read sequencing resource
  • The new assembly improves on the previous L. (M.) enriettii genome (isolate LEM3045) which had many unplaced contigs, higher gap content, and lower N50 finding
  • All 36 chromosomes were aligned using L. major Friedlin genome as reference guide for scaffolding, with chromosome ends determined complete except for 18 unplaced contigs finding
  • Genome assembly achieved 94.6% BUSCO completeness using Euglenozoa lineage dataset finding
  • Functional annotation was performed using MAKER2 combined with AUGUSTUS trained on Leishmania tarentolae method
  • An in vitro culture system originally developed for Leishmania orientalis axenic amastigotes was used to grow L. enriettii parasites method
  • A reproducible Snakemake-based workflow (LGAAP) was used for assembly, repeat masking, and annotation resource
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome short-read sequencing (DNBSEQ) Leishmania (Mundinia) enriettii, isolate CUR178, strain LV763 none paired-end sequence reads (270bp, 500bp) Illumina HiSeq (BGI)
Whole-genome short-read sequencing (TruSeq Nano) Leishmania (Mundinia) enriettii, isolate CUR178, strain LV763 none paired-end sequence reads (300bp) Illumina MiSeq (Aberystwyth University)
Long-read whole-genome sequencing Leishmania (Mundinia) enriettii, isolate CUR178, strain LV763 none long reads for chromosome-scale scaffold assembly Oxford Nanopore, SQK-LSK109 protocol, R9 flow cells (FLO-MIN106)
Genome assembly and polishing Leishmania (Mundinia) enriettii, isolate CUR178, strain LV763 none chromosome-scale genome assembly Flye, Minimap2, SAMtools, Pilon, Funannotate, RaGOO
Genome completeness assessment Leishmania (Mundinia) enriettii, isolate CUR178, strain LV763 none single-copy ortholog completeness (%) BUSCO (Euglenozoa lineage dataset, 130 orthologs, 31 species)
Functional gene annotation Leishmania (Mundinia) enriettii, isolate CUR178, strain LV763 none predicted gene models, gene counts, exon counts MAKER2 with AUGUSTUS (trained on Leishmania tarentolae)
Read quality assessment Leishmania (Mundinia) enriettii, isolate CUR178, strain LV763 none sequencing read quality metrics MultiQC
Key results
  • BUSCO completeness of 123/130 single-copy orthologs identified 94.6%
  • Total genome size assembled 33,318,864 bp
  • Assembly N50 value 1,075,649 bp
  • Total number of reads generated across all platforms 26,789,424 reads
  • Total scaffolds in final assembly with 18 unplaced contigs 54 scaffolds; 76,607 bp unplaced
  • Number of predicted genes annotated 8,353 genes
  • Genome coverage achieved 271.8x
  • GC content of genome 59.60%
Key statistics
  • count 26,789,424 total reads (combined MiSeq, HiSeq, and MinION reads)
  • count 5,060,124 MiSeq reads (short-read sequencing)
  • count 20,936,270 HiSeq reads (short-read sequencing)
  • count 793,030 MinION reads (N50 12,070 bp) (long-read sequencing)
  • other 19.41 Gb total bases, 271.8x coverage (sequencing depth)
  • count 8,353 genes; 8,584 exons; mean gene length 1,897 bp (annotation summary)
  • other 94.6% BUSCO completeness (123/130 orthologs) (Euglenozoa lineage dataset assembly completeness)
  • other CDS total length 15.46 Mb (46.40% of genome) (coding sequence content)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a genome resource announcement reporting the chromosome-scale assembly and annotation of Leishmania (Mundinia) enriettii isolate CUR178 using hybrid short-read (Illumina MiSeq and HiSeq) and long-read (Oxford Nanopore MinION) sequencing. No inferential statistics were performed; results are presented entirely as descriptive bioinformatic metrics (N50, genome size, GC content, gene counts, BUSCO completeness percentage). An assembly comparison figure presents metrics for three genomes side by side without formal statistical testing.

Replicationunclear Sample sizeSix naturally infected guinea pigs were the biological source of parasites; all sequencing libraries were explicitly stated to derive from a single extracted DNA sample to avoid inconsistency GroupsNo experimental groups compared; assembly metrics described descriptively alongside two reference genomes (L. (M.) enriettii LEM3045 and L. major Friedlin) in a figure Pairingna Randomization/blindingnot stated Dispersionnone
Approaches that could also have been used
  • Long reads were assembled with Flye using default parameters
    Could also: Canu or wtdbg2 (Redbean) could also have been used for de novo long-read assembly from Nanopore data — Different assemblers use distinct graph algorithms (string-overlap graph in Canu vs. repeat graph in Flye); benchmarking two or more assemblers and selecting the best by N50 and BUSCO is a common approach that increases confidence in the chosen assembly
  • Polishing used Pilon with short-read alignments produced by Minimap2
    Could also: Medaka (a Nanopore-native neural-network polisher) could also have been applied before Pilon as a first polishing pass — A two-stage strategy—Nanopore-aware polishing first (Medaka), then Illumina-based correction (Pilon)—is increasingly used to address platform-specific error profiles sequentially before hybrid correction, which can further reduce residual indels
  • Assembly completeness was assessed with BUSCO using the Euglenozoa lineage dataset (130 single-copy orthologs from 31 species)
    Could also: Merqury (k-mer-based quality value and completeness scoring derived directly from the short reads) could also have been used alongside BUSCO — BUSCO measures conserved gene presence while Merqury provides a reference-free, k-mer-based quality value (QV) estimating base-level accuracy; the two metrics are complementary and their combined use is increasingly standard for hybrid assemblies
  • Reference-guided scaffolding was performed with RaGOO using the L. major Friedlin genome as a guide
    Could also: ALLMAPS or, if Hi-C data were available, 3D-DNA or Salsa2 could also have been used for scaffolding — Reference-guided scaffolding relies on conserved synteny with a related organism; chromosome conformation capture (Hi-C)-based approaches generate scaffolding evidence from the target organism's own chromatin structure without assuming synteny conservation, which may be incomplete across the Leishmania subgenus boundary
  • Gene prediction used AUGUSTUS trained on L. tarentolae as part of the MAKER2 pipeline
    Could also: BRAKER2 could also have been used, incorporating RNA-seq alignments to supplement or replace the cross-species ab initio predictor — Ab initio prediction trained on a related but distinct species may miss organism-specific gene models; integrating transcriptomic evidence from the target organism generally improves both gene model sensitivity and specificity
  • Cross-assembly comparison (Figure 1) was presented as a visual display of metrics without a standardized reporting framework
    Could also: QUAST could also have been used to generate a systematic, tool-standardized multi-metric comparison report across the three assemblies — QUAST produces a reproducible, normalized set of assembly quality statistics (NGA50, misassembly counts, indel rates relative to a reference, etc.) in a single report, making cross-assembly comparisons more consistent and easier to interpret alongside metrics like BUSCO
Software: Flye · Minimap2 · SAMtools · Pilon · Funannotate v1.8.1 · RaGOO · BUSCO · MAKER2 · AUGUSTUS · MultiQC · Snakemake

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Table
genome_size
Reported
33,318,864 bp
Reproduced
33,318,864 bp
exact
n50
Reported
1,075,649 bp
Reproduced
1,075,649 bp
exact
gc_content
Reported
59.60%
Reproduced
59.57%
within tolerance
total_scaffolds
Reported
54
Reproduced
54
exact
chromosomes
Reported
36
Reproduced
36
exact
unplaced_contigs
Reported
18 (76,607 bp)
Reproduced
18 (76,607 bp)
exact
n_genes
Reported
8,353
Reproduced
8,353
exact
n_exons
Reported
8,584
Reproduced
8,584
exact
cds_total
Reported
15.46 Mb (46.40%)
Reproduced
15.46 Mb (46.40%)
exact
gene_density
Reported
250.7 genes/Mb
Reproduced
250.7 genes/Mb
exact
mean_gene_len
Reported
1,897 bp
Reproduced
1,897.5 bp
exact
coverage
Reported
271.8x
Reproduced
271.8x (corroborated by 8.99 Gb/33.32 Mb=270x)
exact
miseq_reads
Reported
5,060,124
Reproduced
5,060,124
exact
hiseq_reads
Reported
20,936,270
Reproduced
20,936,270
exact
minion_reads
Reported
793,030
Reproduced
719,639
did not match
total_reads
Reported
26,789,424
Reproduced
26,716,033
did not match
total_bases
Reported
19.41 Gb
Reproduced
8.99 Gb (8,994,020,225 bp)
did not match
minion_read_n50
Reported
12,070 bp
Reproduced
11,852 bp
partial
busco
Reported
94.6% (123/130)
Reproduced
99.2% (129/130)
partial
flye_denovo
Reported
chromosome-scale ~33 Mb (LGAAP step 1)
Reproduced
32.81 Mb / 154 contigs / N50 664 kb (98.5% of genome size)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 80/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

Every core assembly and annotation claim in Table 1 reproduces exactly against the authoritative deposited GenBank assembly GCA_017916305.1 and the authors' MAKER GFF (genome size, N50, 36 chromosomes, 8,353 genes, CDS 46.40%, gene density 250.7), with GC content within tolerance (59.60% vs 59.57%), so the central chromosome-scale-assembly claim fully holds. The one notable defect sits on the authors' side: the reported sequencing total of 19.41 Gb is not derivable from the deposited reads (~8.99 Gb) and is internally inconsistent with the paper's own 271.8x coverage (~9.06 Gb), and the MinION read count (793,030 vs deposited 719,639) also mismatches. This is a peripheral input-description anomaly, not a failure of the genome itself, so severity is moderate and overall quality is solid-with-explainable-deviation rather than 1:1.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

82.1 k
tokens (I/O) · 3.8 M incl. cache
8 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.