MOSAIK: a hash-based algorithm for accurate next-generation sequencing short-read mapping.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Built MOSAIK 2.2.3 from source and ran its full alignment pipeline (MosaikBuild -> MosaikAligner with the paper's non-default PacBio params -hs 10 -mmp 0.5 -act 15) against a V. cholerae N16961 reference and all 47 runs currently served under the paper's cited accession SRX032454. The pipeline itself works correctly (independently-written BAM FLAG parser matches MOSAIK's own internal U+M .stat accounting exactly), but the two headline PacBio numbers from Table 1 do not reproduce: reported avg read length 698.61bp [48;6084] vs our 1935.02bp (pooled all runs) / 555.23bp (consensus-reads-only subset); reported 85.79% aligned vs our 14.14% (pooled) / 55.11% (consensus-only). Root cause identified: the paper's inline citation points to a specific run SRR075103 under a 2014 'provisional' SRA path that no longer resolves in ENA (empty filereport, HTTP 404); today's SRX032454 accession instead serves a reorganized 47-run replacement from the same study/sample/library. This is a dataset-provenance/archival-drift finding, not a pipeline or methodology failure -- MOSAIK ran correctly against the best currently-available proxy for the paper's data. The reference genome size is close but not exact (4,091,269bp reproduced vs 4,033,464bp reported, +1.4%), consistent with a different assembly patch of the same N16961 strain. Did not attempt: other Table 1 platform rows (Illumina/SOLiD/454, out of scope per the RU's listed sra:SRX032454 accession); a per-run sweep of all 47 individual runs to find a closer match (not pursued once the provenance root cause was established). No completeness claim is made beyond the PacBio row of Table 1.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-08-06
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-06
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan a single hash-based, reference-guided short-read aligner map reads from all major second- and third-generation sequencing technologies while delivering high alignment accuracy, INDEL sensitivity, and well-calibrated (retrainable) mapping quality scores? The paper presents and benchmarks MOSAIK against existing aligners to test this.
- ★ MOSAIK is the only aligner that consistently aligns reads from all major sequencing platforms (Illumina, AB SOLiD, Roche 454, Ion Torrent, Pacific Biosciences SMRT) using the same algorithmic approach. resource
- ★ A hash-clustering strategy coupled with a Smith-Waterman polishing step yields highly accurate alignments that capture mismatches and short insertions/deletions. method
- ★ A neural-network-based training scheme (FANN library) produces well-calibrated Phred mapping quality scores from features such as best/second-best Smith-Waterman scores, read entropy, number of candidate mapping locations, hashes, and fragment length for paired ends. mechanism
- ★ MOSAIK achieves PPV of 99.5% for all alignments and 100.0% for high-confidence alignments (MQ > 20) on simulated data. finding
- ★ A retrainable mapping-quality pipeline lets users generate genome- or technology-specific neural networks, improving calibration for non-human genomes such as E. coli. method
- ★ MOSAIK's ROC/mapping-quality curve is smooth, so downstream tools using MQ cutoffs do not experience abrupt changes in read counts, unlike other aligners. finding
- ★ MOSAIK is the only mapper tested that is highly sensitive to both insertion and deletion polymorphisms (1-14 bp), and the most sensitive for deletions. finding
- MOSAIK provides explicit support for known-sequence structural variants (e.g. mobile element insertions) via IUPAC-aware alignment, optional output of all mapping locations, and is open-source, multi-threaded C++ integrated into the GKNO pipeline launcher. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Short-read alignment rate/throughput benchmarking | Simulated Illumina PE and SE reads against human genome hg19 | none (default MOSAIK parameters) | Percent reads aligned; alignment speed in reads/second | MASON read simulator; MOSAIK |
| Short-read alignment accuracy benchmarking (PPV vs mapping-quality cutoff) and ROC analysis | 12 million simulated Illumina paired-end reads (76 bp and 100 bp) from chromosome 20 of human hg19, aligned to the whole human genome | Simulated haplotype SNP rate of 0.1% | Positive predictive value (correctly placed reads / total mapped reads, 20 bp tolerance window); number of incorrectly mapped reads vs total mapped | MASON simulator; BWA-0.5.9, BOWTIE-2.0-beta5, STAMPY-1.0.13, MOSAIK-2.1.78 (default parameters) |
| Mapping quality calibration analysis (assigned vs actual Phred score) | Simulated Illumina paired-end reads (100 bp and 76 bp), human genome | none | Pearson correlation coefficient between assigned and actual mapping qualities | MOSAIK (FANN neural network), BWA, BOWTIE, STAMPY |
| Neural-network retraining and mapping-quality calibration | E. coli genome; 6 million simulated paired-end reads for training plus an independent set of 6 million simulated paired-end reads for testing | Retraining of the mapping-quality neural network on E. coli vs default human-trained network | Pearson correlation coefficient between assigned and actual mapping qualities | MOSAIK retrainable mapping-quality pipeline (FANN) |
| INDEL sensitivity benchmarking | Simulated Illumina paired-end reads containing 1-14 bp INDEL events | Introduced INDELs, average of ~100 events per INDEL length with ~800 spanning reads | Sensitivity (correctly mapped reads containing the simulated variant / total simulated reads) as a function of INDEL length | MUTATRIX genome simulator; MOSAIK, BWA, BOWTIE, STAMPY |
| Alignment of real second-generation sequencing data | Illumina (PE/SE) and AB SOLiD reads from the Han Chinese in Beijing (CHB) population, 1000 Genomes Project; human hg19 reference | none (default parameters) | Percent reads aligned; alignment speed | MOSAIK; SOLiD reads aligned in colorspace with post-alignment conversion to basespace |
| Alignment of third-generation sequencing data (Ion Torrent) | E. coli strain 536 reads released by Ion Torrent | none (default parameters) | Percent reads aligned; alignment speed; read length distribution | Ion Torrent PGM data; MOSAIK |
| Alignment of third-generation single-molecule sequencing data (PacBio SMRT) | V. cholerae (4,033,464 bp genome) reads released by Pacific Biosciences (SRR075103) | Non-default parameters: -hs 10 -mmp 0.5 -act 15 (vs default -hs 15 -mmp 0.15 -act 55) | Percent reads aligned; alignment speed | Pacific Biosciences SMRT data; MOSAIK |
- ▲ MOSAIK reaches 100.00% PPV at a mapping-quality cutoff of 20, versus 99.99% for BWA, 99.79% for BOWTIE and 99.63% for STAMPY 100.00% vs 99.99%/99.79%/99.63%
- ▲ MOSAIK has the highest average Pearson correlation between assigned and actual mapping qualities across read lengths 0.9698 vs 0.9061 (BWA), 0.9207 (BOWTIE), 0.8652 (STAMPY)
- ▲ Retraining the mapping-quality neural network on E. coli improves calibration over the default human-trained network r = 0.9749 (retrained) vs 0.8881 (default)
- – Lowering the mapping-quality cutoff from 30 to 29 causes a far smaller jump in incorrectly mapped reads for MOSAIK than for BWA, indicating a smoother ROC/dynamic range 6.25% increase (MOSAIK) vs 308.56% (BWA)
- – Alignment rates vary by technology, being highest for low-error simulated Illumina data and lowest for SOLiD colorspace data 99.98% (Illumina PE simulated), 99.75% (Illumina SE simulated), 99.42% (454 SE), 91.48% (1000G Illumina PE/SE), 85.79% (PacBio), 77.02% (Ion Torrent), 55.64% (SOLiD)
- – MOSAIK is the most sensitive aligner tested for deletions and is comparable to but slightly worse than STAMPY and BOWTIE for insertions, making it the only mapper highly sensitive to both
- – Alignment speed differs markedly across technologies, with longer and paired-end reads slower to align 153.98 reads/s (Illumina SE) to 0.69 reads/s (PacBio)
- – Run times of the most recent BWA, BOWTIE and MOSAIK versions are comparable, while STAMPY is approximately six times slower ~6x slower for STAMPY
- correlation greater than 0.98 (abstract); 0.9698 average across read lengths in Results (Pearson correlation between MOSAIK-assigned and actual mapping qualities)
- correlation MOSAIK 0.9609 (100 bp), 0.9497 (76 bp), 0.8881 (E. coli); MOSAIK-retrained 0.9749 (E. coli) (Table 3, Pearson's correlation coefficients of mapping qualities)
- correlation BWA 0.8987/0.8625/0.8936; BOWTIE 0.9027/0.9449/0.6989; STAMPY 0.8317/0.8818/0.5262 (Table 3 comparator aligners at 100 bp, 76 bp, E. coli)
- other PPV 99.5% for all alignments; 100.0% for alignments with mapping quality larger than 20 (MOSAIK accuracy on simulated data (Introduction))
- other MOSAIK PPV 1 / 1 (MQ30), 1 / 1 (MQ20), 0.9999 / 0.9999 (MQ10), 0.9962 / 0.9947 (MQ0) for 100 bp / 76 bp (Table 2 PPV by mapping quality cutoff)
- count 12 million simulated Illumina paired-end reads (76 bp and 100 bp) from chromosome 20 of hg19; haplotype SNP rate 0.1%; 20 bp tolerance window (Accuracy benchmarking dataset)
- count 6 million simulated E. coli paired-end reads for training plus 6 million independent reads for testing; INDEL benchmark used ~100 events per length (1-14 bp) with ~800 spanning reads (Neural-network retraining and INDEL sensitivity datasets)
- fold_change 308.56% increase (BWA) vs 6.25% increase (MOSAIK) in incorrectly mapped reads when lowering MQ cutoff from 30 to 29; STAMPY ~6x slower run time; low-memory mode ~8 Gb footprint for human genome (ROC dynamic range, runtime and memory figures)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper is a computational benchmarking study of the MOSAIK sequencing read aligner, comparing it against BWA, BOWTIE, and STAMPY using simulated (MASON, MUTATRIX) and real (1000 Genomes, Ion Torrent, Pacific Biosciences) sequencing datasets. Performance was assessed and reported through descriptive accuracy metrics such as positive predictive value (PPV), sensitivity, ROC curves, and Pearson correlation coefficients between assigned and actual mapping-quality scores, rather than through formal inferential hypothesis testing. Results are presented as point estimates, tables, and curves across read lengths, technologies, and mapping-quality cutoffs.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson correlation coefficient | Correlation between MOSAIK/BWA/BOWTIE/STAMPY assigned mapping-quality scores and actual (empirical) mapping-quality scores (Figure 3, Table 3) | — | not stated |
| Positive predictive value (proportion of correctly mapped reads among mapped reads) | Aligner accuracy comparison across mapping-quality cutoffs (Figure 1, Table 2) | 12 million simulated Illumina paired-end reads (76 bp and 100 bp) from chromosome 20 | na |
| Receiver operating characteristic (ROC) curve | Comparison of mapped-read counts vs. incorrectly-mapped-read counts across mapping-quality thresholds (Figure 2) | — | na |
| Sensitivity (proportion of correctly mapped reads among simulated reads) | INDEL detection accuracy across 1-14 bp insertion/deletion events (Figure 4) | ~100 INDEL events per length, ~800 spanning reads per event, generated with MUTATRIX | na |
-
Aligner accuracy metrics (PPV, sensitivity) are reported as single point-estimate percentages without an accompanying measure of uncertainty.↳ Could also: Report binomial (or bootstrap-derived) confidence intervals around these proportions, since they are calculated from counts of correctly/incorrectly mapped reads. — A confidence interval would convey how precisely each accuracy estimate is known, which can be especially informative when comparing aligners on smaller subsets (e.g., longer INDELs with fewer spanning reads).
-
Aligners are compared visually via ROC and PPV-vs-cutoff curves computed on the same read sets.↳ Could also: A paired comparison method such as McNemar's test or DeLong's test for correlated ROC curves could also be used, since each aligner is evaluated on the identical simulated reads. — Because the aligners are compared on matched (paired) data rather than independent samples, a paired statistical test could formally quantify whether observed differences in accuracy exceed what might arise by chance.
-
Mapping-quality calibration is summarized using Pearson's correlation coefficient between assigned and actual quality scores.↳ Could also: Alongside Pearson's r, a confidence interval for the correlation (e.g., via Fisher's z-transformation) or a Spearman rank correlation could also be reported. — A CI would indicate the precision of the correlation estimate, and a rank-based correlation would serve as a complementary check in case the assigned/actual quality relationship is monotonic but not strictly linear.
-
Benchmarking results are based on a single simulation run per condition (e.g., one set of 12 million simulated reads from chromosome 20).↳ Could also: Multiple independent simulation replicates, with metrics summarized as a mean and spread (e.g., mean ± SD) across runs, could also be used. — Repeating the simulation would let run-to-run variability be distinguished from genuine differences in aligner performance.
-
No correction for multiple comparisons is described, even though many aligner-pair and condition comparisons (read lengths, technologies, mapping-quality cutoffs) are presented.↳ Could also: If formal hypothesis tests were added to these comparisons, a family-wise error correction (e.g., Bonferroni) or false-discovery-rate control (e.g., Benjamini-Hochberg) could also be applied. — This would help control the overall false-positive rate across the many pairwise comparisons being made.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.