Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

OptiType: precision HLA typing from next-generation sequencing data.

Bioinformatics · 2014
L1 95/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
95/100
Reproducibility score
1.2 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 89% of all assessed papers rank 105 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced OptiType's own 16-sample colorectal-cancer (SRP010181) HLA class I typing benchmark end-to-end: downloaded all 64 SRA runs, built OptiType+RazerS3+GLPK from bioconda on «infra» scratch, ran OptiTypePipeline.py --rna on all 16 samples via SLURM, and compared results against the paper's own Supplementary Table S4 (predictions + PCR/Sanger ground truth). Mean two-digit accuracy reproduced exactly (0.98956 vs 0.98956 reported); mean four-digit accuracy reproduced within tolerance (0.97708 vs 0.96663 reported, reproduction slightly higher). 14/16 samples' per-sample predictions are byte-identical to the paper's own reported predictions; the other 2 differ by a single allele call each, one of which is actually more accurate against ground truth than the originally published result. No claim graded mismatch or error. Other OptiType paper benchmarks (1000G, HapMap, simulated reads) and the original Warren et al. ground-truth source (inaccessible, HTTP 403) were out of scope for this pass.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-08-02
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-08-03
no human curator yet
Last updated
2026-08-03

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can the four-digit HLA class I genotype be determined accurately and purely computationally from routine NGS data (RNA-Seq, exome, whole-genome) that were not specifically enriched for the HLA cluster? The authors hypothesize that the correct genotype is the allele combination explaining the largest number of mapped reads when all major and minor HLA-I loci are considered simultaneously, a problem formulated as an integer linear program.

Core claims
  • OptiType, an ILP-based HLA genotyping algorithm, produces accurate four-digit HLA-I predictions from NGS data not enriched for the HLA cluster. method
  • OptiType significantly outperforms previously published in silico HLA typing approaches, achieving an overall accuracy of 97%. finding
  • Considering all major (A, B, C) and minor (G, H, J) HLA-I loci simultaneously resolves ambiguous read alignments caused by inter-locus sequence homology, a likely cause of low accuracy in locus-independent methods. mechanism
  • Missing intronic sequence for partially sequenced HLA alleles can be imputed from the closest phylogenetic relative, because intronic variability in HLA follows highly systematic, lineage-reflecting mutations. mechanism
  • A homozygosity regularization term (weight beta) is required because the plain read-maximization objective favors heterozygous allele combinations owing to spurious hits such as sequencing errors. method
  • A comprehensive public benchmark dataset spanning RNA, exome and whole-genome sequencing data with PCR-verified HLA genotypes is provided. resource
  • A penalization term gamma prioritizes alleles with full sequence information over reconstructed alleles among equally good solutions. method
  • OptiType is applicable in a clinical setting, validated on in-house exome-sequenced acute lymphoblastic leukemia patient samples. finding
Experimental setups
Assay System Perturbation Readout Platform
RNA-Seq (paired-end, 2x100-102 bp) 16 colorectal cancer samples (SRP010181) none percentage of correctly predicted two-digit and four-digit HLA-I alleles vs PCR-verified genotypes Illumina HiSeq 2000
RNA-Seq (paired-end, 37 nt reads) 50 lymphoblastic cell line samples of CEU HapMap individuals (ERA002336) none percentage of correctly predicted HLA-I alleles Illumina Genome Analyzer II
Low-coverage whole-genome sequencing (paired-end, 2x100-102 bp) 20 HapMap Project samples (plus 12 HapMap WGS runs from the Major et al. benchmark) none percentage of correctly predicted HLA-I alleles Illumina HiSeq 2000
Exome sequencing 253 1000 Genomes Project runs (including the 161 fully typed by Major et al.) and 11 1000 Genomes samples from the ATHLATES benchmark none percentage of correctly predicted HLA-I alleles; also used for beta cross-validation Illumina HiSeq 2000 and Genome Analyzer II
Exome sequencing (paired-end, 76 bp) 10 in-house acute lymphoblastic leukemia (ALL) patients with experimentally determined HLA types none (clinical samples) concordance of predicted HLA type with experimentally determined type Illumina Genome Analyzer IIx; SureSelect Human All Exon V2 (Agilent) or SeqCap EZ Human Exome Library V2 enrichment
HLA-targeted vs standard exome capture sequencing (paired-end, 101 bp) two samples from a single patient HLA-region enrichment (custom SureSelect HLA kit) vs standard exome enrichment (SureSelectXT Human All Exon V5) effect of coverage depth on prediction performance Illumina HiSeq 2500
In silico read downsampling / coverage-depth simulation all 1000 Genomes Project exome sequencing benchmark samples randomized read subsets of decreasing size down to ~0.2x coverage of HLA-I loci prediction accuracy as a function of coverage depth
In silico leave-one-out intron reconstruction validation (sequence alignment / distance matrices) fully sequenced HLA-I alleles from IMGT/HLA Release 3.14.0 introns 1, 2 and 3 discarded and reconstructed from nearest neighbor using only exon 2 and 3 sequences sequence similarity and edit distance of reconstructed vs original intron sequences Clustal Omega 1.2.0
Key results
  • OptiType achieved an overall accuracy of 97% and significantly outperformed previously published in silico HLA typing approaches. 97% overall accuracy
  • Leave-one-out reconstructed intron sequences closely matched their original counterparts. 99.89% (+/- 0.43%) sequence similarity, ~1.2 average edit distance over three introns
  • Intron sequence similarity between alleles of the same locus was substantially lower than reconstruction accuracy, indicating reconstruction adds information beyond locus-level consensus. 97.36% (+/- 2.15%) similarity, 29 nt differences on average
  • Cross-validation of the homozygosity regularization weight identified an optimal value of beta = 0.009. beta = 0.009 (tested 0.000-0.050, step 0.001)
  • Partial HLA alleles had few unique nearest neighbors, yielding a manageable set of reconstructed reference sequences. 1.66 (+/- 1.04) nearest neighbors; 10 779 reconstructed sequences for 6489 partial alleles
  • Prior methods (Kim et al., Warren et al.) achieved only 85-90% correct four-digit HLA genotypes on RNA-Seq, with lower accuracy on short-read RNA-Seq and WGS data. 85-90%
  • Major et al. achieved 94% accuracy on exome samples, but could fully type only 161 of 217 samples considered. 94% accuracy; 161/217 samples fully typed
  • The majority of IMGT HLA sequences are incomplete, motivating phylogeny-based imputation of missing intronic/exonic regions. 94.6% of IMGT HLA sequences lack parts of exonic or intronic sequence
Key statistics
  • other 97% (OptiType overall HLA typing accuracy)
  • other 99.89% (+/- 0.43%) (sequence similarity of leave-one-out reconstructed introns to originals)
  • other 97.36% (+/- 2.15%) (intron sequence similarity between alleles of the same locus (29 nt differences on average))
  • mean 1.66 (+/- 1.04) (average number of unique-intron nearest neighbors per partial allele)
  • count 10 779 reconstructed sequences for 6489 partial alleles (reference library construction)
  • other beta = 0.009; gamma = 0.01 (regularization weights; beta from nested 5-fold cross-validation on 253 1000 Genomes runs)
  • other 94.6% (share of IMGT HLA sequences lacking parts of exonic or intronic sequence)
  • other 94 million reads per sample, 90x average whole-exome coverage (in-house ALL exome sequencing dataset)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper describes a computational method (OptiType) for HLA genotyping and reports performance primarily as percentage accuracy of correctly predicted alleles across multiple public and in-house sequencing datasets. A tunable regularization parameter (β) was selected via nested 5-fold cross-validation stratified by zygosity, and the quality of an intron-sequence-reconstruction procedure was assessed via leave-one-out validation, with results summarized using mean values and a '±' dispersion figure. No classical inferential hypothesis tests (e.g., t-test, ANOVA) or p-values are described in the provided text.

Replicationunclear Sample sizesample/run counts are stated per dataset (e.g., 16 RNA-Seq samples, 20 WGS samples, 50 lymphoblastic cell line samples, 12 HapMap WGS, 161 1000 Genomes exomes, 253 additional exome runs, 11 ATHLATES samples, 10 in-house ALL patients), but no formal power/sample-size calculation is described GroupsOptiType predictions vs. PCR-verified reference HLA genotypes, and vs. predictions from previously published in silico methods Pairingunclear Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnot stated
Statistical tests used
Test Applied to n Assumptions
Nested 5-fold cross-validation (stratified for heterozygous/homozygous cases) selection of the regularization parameter β in the ILP objective function 253 runs of the 1000 Genomes Project not stated
Leave-one-out validation assessing accuracy of reconstructed intron sequences (introns 1, 2, 3) against original sequences the set of fully sequenced HLA alleles (exact n not stated in this excerpt) not stated
Percentage accuracy (proportion of correctly predicted two-digit/four-digit HLA alleles) comparison of OptiType against HLAminer, seq2HLA, HLAforest, ATHLATES and Major et al. across RNA-Seq, exome and WGS benchmark datasets, and an in-house clinical dataset varies by dataset (e.g., 16, 20, 50, 12, 161, 253, 11, 10 samples/runs as described) na
Approaches that could also have been used
  • Accuracy is reported as a single point percentage for each dataset/method comparison.
    Could also: Reporting a binomial (e.g., Wilson score) or bootstrap confidence interval around each accuracy percentage — This would convey the precision of the estimate, which is particularly informative for datasets with smaller sample sizes (e.g., 10-12 samples), where a single percentage can be sensitive to a small number of typing outcomes.
  • Values such as sequence-reconstruction similarity and nearest-neighbor counts are reported with a '±' figure without specifying the dispersion measure.
    Could also: Explicitly labeling the statistic as SD, SEM, or a 95% CI — Naming the specific dispersion measure removes ambiguity about whether the spread reflects variability across alleles (SD) or the precision of an average (SEM/CI), which can otherwise be interpreted differently by readers.
  • The β regularization parameter was tuned using a single nested 5-fold cross-validation stratified by zygosity.
    Could also: Repeating the cross-validation procedure multiple times (repeated k-fold) or using bootstrap resampling — This would provide a distribution of accuracy estimates across repeats/folds, giving a sense of the variability in the selected parameter rather than a single-run estimate.
  • OptiType's per-sample accuracy is compared against other methods (HLAminer, seq2HLA, HLAforest, ATHLATES, Major et al.) as raw percentages on shared or overlapping samples.
    Could also: A paired comparison method for correct/incorrect binary outcomes on the same samples, such as McNemar's test — Since multiple methods are applied to the same or overlapping sample sets, a paired test could quantify whether the accuracy difference between methods exceeds what would be expected by chance, complementing the direct percentage comparison.
  • Reconstruction quality of intron sequences was validated using leave-one-out cross-validation.
    Could also: k-fold cross-validation (e.g., 10-fold) instead of leave-one-out — k-fold CV can reduce computational cost relative to leave-one-out while still yielding a robust, low-variance estimate of reconstruction accuracy, which can be useful as the reference allele set grows.
Software: RazerS3 (SeqAn C++ library) 3.1 · Clustal Omega 1.2.0

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

crc_mean_2digit_accuracy
Reported
0.98956
Reproduced
0.98956
exact
crc_mean_4digit_accuracy
Reported
0.96663
Reproduced
0.97708
within tolerance
crc_per_sample_predictions
Reported
see samples[] detail
Reproduced
see samples[] detail
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 95/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

Full, clean end-to-end reproduction. The exact 64 SRA runs of SRP010181 were identified from the paper's own Table S11, downloaded from ENA, and re-typed with a freshly built OptiType/RazerS3/GLPK stack; mean 2-digit accuracy reproduced exactly (0.98956 vs 0.98956) and mean 4-digit accuracy came out at 0.97708 vs the reported 0.96663. The single deviation source is one allele call in sample 66 (B44:02 here vs the published B44:27) plus a zygosity call in sample 20 — 14/16 samples are byte-identical to the published predictions. The difference lies on neither the authors' nor our methodological side but is a technically expected reference-version effect (OptiType bundles a fixed IMGT/HLA reference with no version flag), and notably it moves against the authors' interest: our run is more accurate against the PCR/Sanger ground truth than the published one, which is the strongest possible evidence against any fabrication concern.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.