LaSSO, a strategy for genome-wide mapping of intronic lariats and branch points using RNA-seq.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓The central claim held under reproduction
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
LaSSO reproduced: GSE50246 primary headline numbers within-tolerance (all 4 metrics ~5-19% of paper), GSE48594 Awan reanalysis partial (improved recovery confirmed, exact counts differ), GSE30611 human validation not attempted (947GB DB infeasible, documented). Env deviation: Bowtie 1.2.3/R 4.2.3.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-31
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan intronic branch points be mapped precisely and genome-wide directly from RNA-seq data? The paper tests whether a data-driven algorithm (LaSSO), which builds a database of all theoretically possible lariat signatures and searches unmapped RNA-seq reads against it, can accurately pinpoint branch points and splicing events, using lariat-accumulating dbr1Δ fission yeast cells as a validation system.
- ★ LaSSO (Lariat Sequence Site Origin) identifies intronic lariat reads and pinpoints branch points genome-wide from RNA-seq data by considering every intronic base as a potential branch point and including all possible exon-skipping lariats. method
- ★ LaSSO accurately identifies canonical splicing events and also detects novel but rare exon-skipping events that may reflect aberrantly spliced transcripts. finding
- ★ Compromised intron turnover in dbr1Δ cells perturbs gene regulation at multiple levels, including splicing efficiency and protein translation. finding
- ★ Dbr1 function is critical for expression of mitochondrial genes and for processing of self-spliced group-II mitochondrial introns (cox1, cob1) in vivo. mechanism
- ★ LaSSO shows better sensitivity and accuracy than existing computational branch-point prediction algorithms and empirical branch-point determination methods (e.g., FELINES, 2D-Lariat-seq). finding
- ★ LaSSO works on human RNA-seq data acquired in the presence of debranching activity, identifying both canonical and exon-skipping branch points. finding
- ★ Adenine is the predominant primary branch-point base, and clustered neighboring branch points form the YURAY branch-site consensus motif; this 'fuzziness' most likely reflects reverse-transcriptase errors/base skipping across the 2'-5' bond. mechanism
- The method provides a high-resolution genome-wide map of branch-site sequences and intronic elements usable for validating introns in any eukaryote and integrable into RNA-seq pipelines. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RNA-seq (transcriptome sequencing) of proliferating cells, biological replicates | Schizosaccharomyces pombe (fission yeast), wild type and dbr1Δ | dbr1Δ deletion mutant (debranching enzyme knockout) vs wild type | Transcript, exon, and intron expression levels; differential expression; splicing efficiency from exon-intron and exon-exon junction reads | 49-bp reads, ~200-base library insert size; reads aligned with Bowtie allowing three mismatches |
| LaSSO computational lariat-database construction and alignment of unmapped RNA-seq reads | S. pombe genome/transcriptome (5361 introns) | none (in silico, applied to dbr1Δ and wild-type RNA-seq) | Diagnostic lariat reads, distinct lariats, branch-point base and position, branch-site consensus sequence | Bowtie alignment to lariat database allowing one mismatch; WebLogo for consensus plots |
| LaSSO applied to published data sets for benchmarking | S. pombe 2D-Lariat-seq data (Awan et al. 2013); human RNA-seq data set | none (human data acquired in presence of debranching activity) | Number of introns/lariats with detected branch points; canonical and exon-skipping branch points; comparison against FELINES predictions | — |
| Lariat-specific RT-PCR followed by Sanger sequencing (independent validation) | S. pombe dbr1Δ cells; 18 selected lariats of different lengths and read support | dbr1Δ deletion | Confirmation of lariat structure/branch point (17 of 18 confirmed) | Sanger sequencing |
| Lariat-specific RT-PCR of mitochondrial self-spliced introns | S. pombe dbr1Δ and wild type; group-II introns of cox1 and cob1 | dbr1Δ deletion | Abundance of lariat intermediates from self-spliced mitochondrial introns | — |
| Polysome profiling (translational profiling) | S. pombe dbr1Δ and wild type | dbr1Δ deletion | Polysome-to-monosome (P/M) ratio | — |
| Protein phosphorylation analysis of translation-initiation markers | S. pombe dbr1Δ and wild type | dbr1Δ deletion | Phosphorylation levels of S6 and eIF2α proteins | — |
| Gene Ontology enrichment / overlap statistical analysis of expression and splicing changes | S. pombe transcriptome data (5361 introns, 38 mitochondrial genes, 10 intron-embedded snoRNAs) | dbr1Δ vs wild type | Enriched GO terms; overlap significance between DESeq-increased introns and reduced-SE introns | DESeq (adjusted P < 0.05, |fold change| > 2); CMH test; hypergeometric, Wilcoxon rank sum, Fisher's exact tests |
- – 108,683 diagnostic lariat reads aligned to the S. pombe lariat database; ~99.8% originated from within 1655 single introns, defining 5060 distinct lariats 108,683 reads; 5060 lariats from 1655 introns
- – Restricting to adenine branch points yielded 2842 distinct lariats from 1584 introns; conservative filtering (≥3 lariat reads) gave 1236 nuclear introns plus one mitochondrial self-spliced intron (cob1) used for the branch-site consensus 93,845 adenine reads; 2842 lariats/1584 introns; 1236 introns at ≥3 reads
- ▲ Intronic expression was significantly higher in dbr1Δ than wild type, with ~25% of introns significantly increased and only 27 decreased; the increase was strongly biased toward long introns 1399/5361 introns increased; P < 2.2 × 10^-16
- ▼ Overall splicing efficiency was significantly lower in dbr1Δ than wild type, with 638 introns showing significant SE changes and no bias toward long introns; low-SE introns strongly overlapped DESeq-increased introns 638 introns (CMH test, Q < 0.05); P < 2.2 × 10^-16
- ▲ 50% of all mitochondrial genes were induced in dbr1Δ, and lariat intermediates of the self-spliced cox1/cob1 group-II introns increased, validated by lariat-specific RT-PCR 19/38 genes; P < 1.4 × 10^-15
- ▼ Most intron-embedded snoRNAs were among decreased transcripts, and decreased transcripts were enriched for ribosome-biogenesis GO categories 6/10 snoRNAs; P < 2.4 × 10^-7
- ▲ The polysome-to-monosome ratio was significantly higher in dbr1Δ than wild type, with no difference in S6 or eIF2α phosphorylation, suggesting translation is compromised at the elongation step P/M 1.83 (dbr1Δ) vs 1.14 (wild type)
- – 17 of 18 selected lariats were confirmed by lariat-specific RT-PCR followed by Sanger sequencing 17/18 (94%)
- count 347 transcripts significantly increased and 233 decreased in dbr1Δ; ~57% of increased transcripts were noncoding RNAs (Differential transcript expression, dbr1Δ vs wild type (DESeq, adjusted P < 0.05, |fold change| > 2))
- count 1399/5361 introns significantly increased; 27 decreased (Differential intronic expression in dbr1Δ (~25% of introns))
- pvalue P < 2.2 × 10^-16 (Wilcoxon rank sum test for higher intronic expression, its length bias, and lower splicing efficiency in dbr1Δ)
- pvalue P < 1.4 × 10^-15 (Hypergeometric test: 19/38 mitochondrial genes induced in dbr1Δ)
- pvalue P < 2.4 × 10^-7 (Hypergeometric test: 6/10 intron-embedded snoRNAs decreased in dbr1Δ)
- mean P/M ratio 1.83 (dbr1Δ) vs 1.14 (wild type) (Polysome-to-monosome ratio from translational profiling)
- count 108,683 diagnostic lariat reads; 93,845 adenine branch-point reads; 5060 lariats from 1655 introns; 2842 adenine-only lariats from 1584 introns; 1236 introns at ≥3 reads (LaSSO lariat detection in S. pombe RNA-seq)
- other 93% of S. pombe introns shorter than the ~200-base average insert size; ~36% smaller than the 49-bp read length; ~47% of S. pombe genes are spliced (Technical basis for preferential recovery of lariats from long introns)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper analyzes RNA-seq data from wild-type and dbr1Δ fission yeast biological replicates, using DESeq for differential transcript/intron expression (adjusted P<0.05, fold change>2), nonparametric rank-based tests (Wilcoxon rank sum) for comparing distributions of expression or splicing efficiency between conditions, a Cochran-Mantel-Haenszel (CMH) test for per-intron splicing-efficiency changes, hypergeometric tests for gene-set/overlap enrichment, and Fisher's exact test for a proportions comparison. Reproducibility between biological replicates was assessed with Pearson correlation coefficients. Results are reported mainly as P-value thresholds and fold-change/ratio values displayed in box plots, scatter plots, and MA plots, with select findings independently validated by RT-PCR and Sanger sequencing.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Wilcoxon rank sum test | intronic vs transcript expression comparison (Fig. 2A); intronic expression length-bias comparison between dbr1Δ and wild type (Fig. 2D); splicing efficiency (SE) comparison between dbr1Δ and wild type (Fig. 3C) | 5361 introns | not stated |
| DESeq (differential expression, adjusted P-value + fold-change cutoff) | differential transcript and intron expression between dbr1Δ and wild type (Fig. 2B) | not stated (whole transcriptome/intron set) | not stated |
| Hypergeometric test | enrichment of mitochondrial genes and snoRNAs among differentially expressed transcripts; overlap between introns with lower SE and introns with higher DESeq-based expression (Fig. 3F,G) | e.g., 19/38 mitochondrial genes; 6/10 intron-embedded snoRNAs | not stated |
| Cochran-Mantel-Haenszel (CMH) test | per-intron splicing efficiency changes between dbr1Δ and wild type (Fig. 3D) | 5361 introns; 638 introns called significant | not stated |
| Fisher's exact test | proportion of lariat reads relative to unmapped reads, dbr1Δ vs wild type (Fig. 5A) | counts of lariat reads vs total unmapped reads per sample | not stated |
| Pearson's correlation coefficient | reproducibility of intronic expression and splicing efficiency between biological replicates (Figs. 2C, 3A,B) | 5361 introns | not stated |
-
Several separate hypothesis tests (Wilcoxon rank sum, hypergeometric, Fisher's exact, CMH) are applied to different comparisons throughout the study, each evaluated largely on its own P-value threshold.↳ Could also: A joint multiple-testing correction (e.g., Benjamini-Hochberg FDR) applied across the full family of reported comparisons — Treating the whole set of tests performed in the paper as one family, rather than correcting within each analysis separately, would additionally control the overall false-discovery rate across the study.
-
Differential expression between dbr1Δ and wild type was called using DESeq with an adjusted P-value and a fixed fold-change cutoff (>2).↳ Could also: Shrinkage-based effect-size estimation (e.g., DESeq2/apeglm or edgeR with an effect-size-aware significance criterion, such as TREAT/glmTreat testing against a fold-change threshold) — Shrinking noisy fold-change estimates for low-count genes/introns before applying a fold-change cutoff can reduce the influence of large but imprecise fold changes, which is particularly relevant given the length-associated technical bias noted for intron read recovery.
-
Enrichment of specific GO categories (e.g., mitochondrial translation, ribosome biogenesis) among differentially expressed transcripts was assessed with a hypergeometric test on genes passing a significance cutoff.↳ Could also: Rank-based gene set enrichment analysis (e.g., GSEA) using the full ranked list of fold changes or test statistics — Using the continuous ranking of all transcripts rather than a hard significance cutoff can capture coordinated but individually sub-threshold shifts within a gene set.
-
Reproducibility between biological replicates was quantified using Pearson's correlation coefficient on intronic expression and splicing efficiency values.↳ Could also: Spearman's rank correlation or a concordance correlation coefficient (Lin's CCC) — Spearman's correlation is robust to non-normal or skewed count-based distributions (common in RNA-seq intensity data), and CCC additionally penalizes systematic deviation from the identity line, which the paper's replicate comparison is qualitatively assessing.
-
Splicing efficiency changes per intron were tested with a Cochran-Mantel-Haenszel test on the underlying exon-intron/exon-exon junction read counts.↳ Could also: Beta-binomial or logistic regression modeling of splicing efficiency as a proportion (e.g., as implemented in tools such as DEXSeq or MISO/rMATS-style frameworks) — A regression-based proportion model can incorporate overdispersion and covariates (e.g., intron length) directly, which is relevant given the length-related bias the authors separately investigate.
-
Distributional comparisons (e.g., intronic vs transcript expression, dbr1Δ vs wild-type SE) were summarized in box plots with associated Wilcoxon rank-sum P-values.↳ Could also: Reporting an accompanying effect-size measure (e.g., rank-biserial correlation or Hodges-Lehmann estimator of median difference) alongside the P-value — A nonparametric effect-size estimate would convey the magnitude of the shift between distributions in addition to its statistical significance, complementing the box-plot visualization.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Solid partial reproduction, deviations on our side. Two of the three datasets were downloaded and reprocessed end-to-end from public raw data: the primary GSE50246 headline numbers reproduce at 120,367 / 98,423 / 3,382 / 1,843 against the reported 108,683 / 93,845 / 2,842 / 1,584 (+5% to +19%, all same direction), and the Awan reanalysis confirms the improved-recovery claim multi-fold (103,907 vs 37,008 reads; 1,637 vs 813 introns) though the intron count overshoots the paper's 1,268 by 29%. The residual gap is best explained by the fact that the authors released only LaSSO.R (database construction) — the alignment, filtering and counting logic was reimplemented from Methods prose under Bowtie 1.2.3 instead of 0.12.7 — so this is our-method + partial-code-release, not evidence against the authors' numbers; the biological control (dbr1Δ carrying 119,552 of 120,367 reads vs 815 in WT) matches expectation exactly. The one genuine authors'-side gap is documentation: the human GSE30611 filtering that takes 51,469 candidates down to 586 reads is never specified, and that sub-result was transparently declared not-attempted (947 GB database) rather than fabricated. Overall yellow: reproducible in substance and conclusion, not 1:1 in counts, with the shortfall traceable to incomplete code/parameter release.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.