RNA-Seq alignment to individualized genomes improves transcript abundance estimates in multiparent populations.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced the core computational claim of Walker et al. 2014 (PMID 25236449): aligning real CAST/EiJ founder-strain liver RNA-seq (SRR826332, 10.4M reads) to a Seqnature-built CAST-individualized transcriptome yields more aligned reads than aligning to the reference (C57BL6J) transcriptome, both at bowtie -v3 --best --strata (75.36% vs 77.73%, +2.37pp; paper reports ~2.1%) and at -v0 perfect-match (35.58% vs 47.64%, +12.06pp/+33.9% relative; paper reports ~23%) -- directionally and quantitatively consistent with the paper's reported improvement, graded 'partial' due to documented methodological substitutions (transcriptome- vs genome-level alignment; modern GRCm38/Ensembl102/bowtie1.3.1/mgp.v5-REL1505 vs the paper's original mm9/NCBIM37/bowtie0.12.8/REL-1211, since the original VCF release is no longer hosted). The Seqnature individualized-genome build itself required two genuine bug workarounds (VCF chromosome-order mismatch between current Sanger SNP/indel exports; 14/1.87M corrupted GTF coordinate lines from adjust_annotations.py) and surfaced one un-fixed, well-bounded tool limitation (the genome builder never emits 44 unplaced/alt scaffolds, causing a 0.07% transcript-count gap on scaffold-only genes). All 3 of GSE45684's processed expression matrices (DO-to-individualized, DO-to-NCBIM37, FounderStrain-to-strain-genomes) were retrieved and profiled; sample counts in all three match the paper's reported cohort sizes exactly (277 DO, 128 founder-strain); one file required a CGI-endpoint download workaround after HTTPS/FTP mirror failures specific to that file. Full technical detail in report_md. EXTENSION (2026-08-06): went beyond the alignment-rate claims above to test the paper's actual title-level claim -- transcript ABUNDANCE estimation, not just alignment rate -- by running RSEM gene-level quantification (rsem-calculate-expression) for SRR826332 against both the reference and CAST-individualized transcriptomes. Of 20,477 expressed genes, 4,572 (22.3%) differ by >10% in TPM between the two alignments, and 2,434/4,572 (53.2%) are higher under the CAST-individualized reference -- directionally consistent with the paper's real-data finding (714/12,248 genes differed >10%, 432/714 improved) though proportionally larger here, expectedly, since this uses a single near-homozygous founder-strain sample rather than the paper's heterozygous DO population and a simpler transcriptome-only RSEM pipeline rather than genome-level quantification. The paper's primary simulation-based validation claim (r=0.82 against 10M simulated-read ground truth) remains out of reach: the released Seqnature-master repo ships no simulation script. Graded 'partial' for the same reason as the alignment-rate claims: correct direction and mechanism, not exact magnitude agreement, due to disclosed and unavoidable sample-type/pipeline differences from the original 2014 study.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-08-07
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-07
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusDoes aligning RNA-seq reads to individualized diploid genomes/transcriptomes that incorporate known genetic variants — rather than to a single haploid reference genome — reduce read-misalignment bias and improve transcript abundance estimation in genetically diverse multiparent populations? The authors test whether this improves mapping accuracy, allele-specific expression estimation, and eQTL mapping.
- ★ Genetic variants distinguishing an individual genome from the reference cause read misalignment and biased transcript abundance estimates, and fine-tuning of alignment algorithms does not correct this problem. finding
- ★ Seqnature software constructs individualized diploid genomes, transcriptomes, and coordinate-adjusted gene annotations for inbred strains and multiparent populations (e.g., Diversity Outbred mice) from founder haplotypes and known SNPs/indels. resource
- ★ Alignment to individualized transcriptomes increases read mapping accuracy and improves estimation of transcript abundance in both simulated and real data sets. finding
- ★ The individualized diploid transcriptome enables direct estimation of allele-specific expression for genes with heterozygosity, including allocation of allelic multireads via an EM algorithm rather than discarding them. method
- ★ In eQTL mapping, individualized alignment corrects false-positive linkage signals and unmasks hidden (previously masked) associations, including extensive local genetic variation affecting gene expression. finding
- A complete analysis pipeline integrates Seqnature with existing tools — Bowtie for read alignment and RSEM for EM-based quantification — with other tools substitutable. method
- Reads are aligned at the isoform level but abundance is summarized at the gene level because isoform-proportion precision is low with current sequencing technologies. method
- The authors recommend individualized diploid genomes over reference-sequence alignment for all high-throughput sequencing applications in genetically diverse populations. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Bulk single-end mRNA-seq (100-bp reads) | Liver from male Diversity Outbred founder-strain mice (8 inbred strains), n = 128 | Diet: semipurified control diet (16.8% kcal fat, TD.08810) vs. high-fat/high-sucrose diet (44.6% kcal fat, 34% sucrose, TD.08811) | Isoform- and gene-level transcript abundance (read counts) | Illumina HiSeq 2000; Illumina TruSeq standard unstranded library prep; Trizol Plus RNA extraction kit; CASAVA v1.8.0 base calling; Agilent Bioanalyzer and Kapa Biosystems qPCR QC |
| Bulk single-end mRNA-seq (100-bp reads) | Liver from male and female Diversity Outbred (J:DO stock no. 009376) mice, n = 277, collected at 26 weeks | Diet: standard rodent chow (6% fat by weight, LabDiet 5K52) vs. high-fat/high-sucrose diet (TD.08811) | Gene-level and allele-level transcript abundance; 2–4 technical replicates per DO sample | Illumina HiSeq 2000; Illumina TruSeq unstranded protocol |
| High-density SNP genotyping array | Genomic DNA from each Diversity Outbred mouse | none | SNP probe intensities used to infer 36-state founder haplotypes and phased genotype transitions via hidden Markov model (DOQTL) | Mouse Universal Genotyping Array (MUGA, GeneSeek), 7664 SNPs |
| In silico RNA-seq read simulation | CAST/EiJ inbred genome and one reconstructed Diversity Outbred diploid individual | none (simulation) | Mapping accuracy and transcript abundance vs. known ground truth (realized simulated abundances); allelic abundances simulated independently for the DO sample | Flux Simulator v1.2; 10 and 30 million 100-bp single-end (CAST and DO) or paired-end (CAST only) reads; error model 0.028% average mutations/sequence, quality 34.96; paired-end fragment size 280 ± 50 |
| In silico RNA-seq read simulation for eQTL benchmarking | 277 simulated Diversity Outbred genomes | none (simulation) | Simulated gene-level counts for downstream eQTL mapping comparison | RSEM v1.2.8 rsem-simulate-reads; 30 million 100-bp single-end reads; θ0 = 0.018 |
| Read alignment to reference vs. individualized diploid transcriptomes | Simulated and real CAST, NCBIM37, and DO transcriptomes (Ensembl v67 annotation) | Reference NCBIM37 transcriptome alignment vs. individualized transcriptome alignment | Read mapping accuracy; isoform/gene/allele abundance estimates via EM allocation of multireads | Bowtie v0.12.8 (-v 3, best strata; -X 1000 and -y for paired-end); RSEM v1.2.1 |
| Expression QTL (eQTL) mapping | Liver RNA-seq from 277 male and female DO mice (high-fat or standard chow diet) | Alignment strategy: NCBIM37 reference transcriptome vs. individualized transcriptome | LOD scores / eQTL linkage signals from upper-quartile-normalized, normal-score-transformed gene counts; genes required to have nonzero counts in ≥85% (≥233) of samples | DOQTL R package; linear mixed model with sex, diet, sex-by-diet, and batch as additive covariates plus random polygenic term; QVALUE software |
- ▲ Alignment to individualized transcriptomes increases read mapping accuracy relative to reference-genome alignment in simulated and real data.
- ▲ Individualized alignment improves estimation of transcript abundance and enables direct estimation of allele-specific expression.
- – In eQTL mapping, individualized alignment corrects false-positive linkage signals and unmasks hidden associations, revealing extensive local genetic variation affecting gene expression.
- ▲ The seven nonreference DO founder strains differ substantially from the C57BL/6J (NCBIM37) reference, with variation especially high in the three wild-derived strains (CAST, PWK, WSB). CAST/EiJ: 17,673,726 genomic SNPs; PWK/PhJ: 17,202,436; vs. A/J: 4,198,324
- – Approximately half of all 100-bp reads from CAST are expected to contain at least one SNP or indel, based on transcriptome variant density. SNPs 1/217 bases; indels 1/1650 bases
- ▼ Haplotype transitions falling within genes are rare in DO genomes, supporting the validity of assigning whole genes to a single founder haplotype. 118 of 75,124 events = 0.16%
- – Comparable numbers of genes passed the expression filter under both alignment strategies, allowing a matched eQTL comparison. 17,125 (reference) vs. 16,985 (individualized); 16,924 common
- count 31,593,523 SNPs, 2,963,385 insertions, 3,213,340 deletions (genome); 746,993 SNPs, 56,354 insertions, 61,204 deletions (transcriptome) (Total variants segregating among the eight CC/DO founder strains vs. NCBIM37 (Table 1))
- other 1 SNP per 217 bases; 1 indel per 1650 bases (Variant density in the CAST transcriptome relative to the NCBIM37 reference)
- count 118 of 75,124 total recombination events (0.16%) (Intragenic haplotype transition events in a typical DO genome)
- count n = 277 (Diversity Outbred mice (male and female) profiled by liver RNA-seq and used for eQTL mapping)
- count n = 128 (Male DO founder-strain mice, eight biological replicates per strain:diet group)
- count 7664 SNPs (Markers on the MUGA array used for DO genotyping and haplotype reconstruction)
- count 17,125 / 16,985 / 16,924 genes (Genes with nonzero counts in ≥85% of DO samples: reference alignment / individualized alignment / common set)
- pvalue 1% false discovery rate; 100,000 permutations with extreme value distribution fit to maximum LOD scores (eQTL significance threshold and permutation procedure)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper describes a computational/genomics pipeline (Seqnature) for constructing individualized genomes for RNA-seq alignment in multiparent mouse populations, evaluated using simulated data, real Diversity Outbred (DO) mouse liver RNA-seq, and downstream expression QTL (eQTL) mapping. The primary inferential statistics reported are for the eQTL mapping step: a linear mixed model (implemented in DOQTL) with sex, diet, sex-by-diet, and batch as fixed covariates and a random polygenic term for genetic relatedness, with genome-wide significance assessed via permutation and multiple testing controlled via q-values. Read alignment and abundance estimation accuracy comparisons (reference vs. individualized genome) are described narratively in the excerpt provided but detailed statistical reporting for these comparisons is not visible in the supplied text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Linear mixed model (DOQTL) for expression QTL mapping, with random polygenic term for genetic relatedness | eQTL mapping of gene-level expression across 277 DO liver samples, comparing reference vs. individualized transcriptome alignment | 277 DO mice; 16,924 genes in the common analyzed set | not stated |
| Permutation test (100,000 permutations) with extreme value distribution fit to maximum LOD scores, for genome-wide significance thresholds | establishing per-scan significance thresholds in eQTL mapping | — | not stated |
-
Raw RNA-seq counts were normalized to the upper quartile and transformed to normal scores prior to mixed-model eQTL mapping.↳ Could also: A negative-binomial or Poisson generalized linear (mixed) model (e.g., as in DESeq2/edgeR-style frameworks) applied directly to counts — This would model the mean-variance relationship of RNA-seq count data directly, which can be an alternative to a rank-based normal-scores transformation followed by a Gaussian mixed model.
-
Genome-wide eQTL significance thresholds were derived from 100,000 permutations with an extreme value distribution fit to maximum LOD scores.↳ Could also: A Bonferroni correction based on the effective number of independent tests across the genome scan — This offers a simpler, more computationally direct way to control genome-wide type I error, though it can be more conservative than a permutation-calibrated threshold.
-
Multiple testing across genes in the eQTL scan was addressed by converting permutation-derived P-values to q-values with the bootstrap π0 estimation method.↳ Could also: The Benjamini-Hochberg step-up FDR procedure applied directly to nominal p-values — This is a widely used alternative for FDR control that does not require permutation-based p-value estimation, and can be simpler to implement when permutation is computationally costly.
-
The eQTL mixed model included sex, diet, sex-by-diet, and batch as fixed covariates with a single random polygenic term for genetic relatedness.↳ Could also: A Bayesian hierarchical model with shrinkage priors on covariate effects — This could also be used to borrow information across genes when estimating covariate and genetic effects, which may be useful given the large number of genes tested simultaneously.
-
Technical replicates (2-4 sequencing lanes per DO sample) appear to have been concatenated/pooled prior to quantification rather than modeled as a separate variance component.↳ Could also: A mixed-effects model explicitly including technical replicate (lane) as a random effect — This would allow technical and biological variance to be partitioned explicitly, which can be informative when assessing measurement reproducibility.
-
Filtering retained genes with nonzero counts in at least 85% of DO samples before eQTL analysis.↳ Could also: A minimum count-per-million (CPM) or mean-expression threshold combined with independent filtering approaches (e.g., as used in edgeR/DESeq2 workflows) — This is a standard alternative filtering strategy that accounts for both detection rate and expression magnitude when deciding which genes to include in downstream testing.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.