Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

LaSSO, a strategy for genome-wide mapping of intronic lariats and branch points using RNA-seq.

Genome Res · 2014
L1 68/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Same input data as the authors
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
68/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 32% of all assessed papers rank 765 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

LaSSO reproduced: GSE50246 primary headline numbers within-tolerance (all 4 metrics ~5-19% of paper), GSE48594 Awan reanalysis partial (improved recovery confirmed, exact counts differ), GSE30611 human validation not attempted (947GB DB infeasible, documented). Env deviation: Bowtie 1.2.3/R 4.2.3.

💻 Code ↗ 🗄 Data: GSE48594

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-31
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can intronic branch points be mapped precisely and genome-wide directly from RNA-seq data? The paper tests whether a data-driven algorithm (LaSSO), which builds a database of all theoretically possible lariat signatures and searches unmapped RNA-seq reads against it, can accurately pinpoint branch points and splicing events, using lariat-accumulating dbr1Δ fission yeast cells as a validation system.

Core claims
  • LaSSO (Lariat Sequence Site Origin) identifies intronic lariat reads and pinpoints branch points genome-wide from RNA-seq data by considering every intronic base as a potential branch point and including all possible exon-skipping lariats. method
  • LaSSO accurately identifies canonical splicing events and also detects novel but rare exon-skipping events that may reflect aberrantly spliced transcripts. finding
  • Compromised intron turnover in dbr1Δ cells perturbs gene regulation at multiple levels, including splicing efficiency and protein translation. finding
  • Dbr1 function is critical for expression of mitochondrial genes and for processing of self-spliced group-II mitochondrial introns (cox1, cob1) in vivo. mechanism
  • LaSSO shows better sensitivity and accuracy than existing computational branch-point prediction algorithms and empirical branch-point determination methods (e.g., FELINES, 2D-Lariat-seq). finding
  • LaSSO works on human RNA-seq data acquired in the presence of debranching activity, identifying both canonical and exon-skipping branch points. finding
  • Adenine is the predominant primary branch-point base, and clustered neighboring branch points form the YURAY branch-site consensus motif; this 'fuzziness' most likely reflects reverse-transcriptase errors/base skipping across the 2'-5' bond. mechanism
  • The method provides a high-resolution genome-wide map of branch-site sequences and intronic elements usable for validating introns in any eukaryote and integrable into RNA-seq pipelines. resource
Experimental setups
Assay System Perturbation Readout Platform
RNA-seq (transcriptome sequencing) of proliferating cells, biological replicates Schizosaccharomyces pombe (fission yeast), wild type and dbr1Δ dbr1Δ deletion mutant (debranching enzyme knockout) vs wild type Transcript, exon, and intron expression levels; differential expression; splicing efficiency from exon-intron and exon-exon junction reads 49-bp reads, ~200-base library insert size; reads aligned with Bowtie allowing three mismatches
LaSSO computational lariat-database construction and alignment of unmapped RNA-seq reads S. pombe genome/transcriptome (5361 introns) none (in silico, applied to dbr1Δ and wild-type RNA-seq) Diagnostic lariat reads, distinct lariats, branch-point base and position, branch-site consensus sequence Bowtie alignment to lariat database allowing one mismatch; WebLogo for consensus plots
LaSSO applied to published data sets for benchmarking S. pombe 2D-Lariat-seq data (Awan et al. 2013); human RNA-seq data set none (human data acquired in presence of debranching activity) Number of introns/lariats with detected branch points; canonical and exon-skipping branch points; comparison against FELINES predictions
Lariat-specific RT-PCR followed by Sanger sequencing (independent validation) S. pombe dbr1Δ cells; 18 selected lariats of different lengths and read support dbr1Δ deletion Confirmation of lariat structure/branch point (17 of 18 confirmed) Sanger sequencing
Lariat-specific RT-PCR of mitochondrial self-spliced introns S. pombe dbr1Δ and wild type; group-II introns of cox1 and cob1 dbr1Δ deletion Abundance of lariat intermediates from self-spliced mitochondrial introns
Polysome profiling (translational profiling) S. pombe dbr1Δ and wild type dbr1Δ deletion Polysome-to-monosome (P/M) ratio
Protein phosphorylation analysis of translation-initiation markers S. pombe dbr1Δ and wild type dbr1Δ deletion Phosphorylation levels of S6 and eIF2α proteins
Gene Ontology enrichment / overlap statistical analysis of expression and splicing changes S. pombe transcriptome data (5361 introns, 38 mitochondrial genes, 10 intron-embedded snoRNAs) dbr1Δ vs wild type Enriched GO terms; overlap significance between DESeq-increased introns and reduced-SE introns DESeq (adjusted P < 0.05, |fold change| > 2); CMH test; hypergeometric, Wilcoxon rank sum, Fisher's exact tests
Key results
  • 108,683 diagnostic lariat reads aligned to the S. pombe lariat database; ~99.8% originated from within 1655 single introns, defining 5060 distinct lariats 108,683 reads; 5060 lariats from 1655 introns
  • Restricting to adenine branch points yielded 2842 distinct lariats from 1584 introns; conservative filtering (≥3 lariat reads) gave 1236 nuclear introns plus one mitochondrial self-spliced intron (cob1) used for the branch-site consensus 93,845 adenine reads; 2842 lariats/1584 introns; 1236 introns at ≥3 reads
  • Intronic expression was significantly higher in dbr1Δ than wild type, with ~25% of introns significantly increased and only 27 decreased; the increase was strongly biased toward long introns 1399/5361 introns increased; P < 2.2 × 10^-16
  • Overall splicing efficiency was significantly lower in dbr1Δ than wild type, with 638 introns showing significant SE changes and no bias toward long introns; low-SE introns strongly overlapped DESeq-increased introns 638 introns (CMH test, Q < 0.05); P < 2.2 × 10^-16
  • 50% of all mitochondrial genes were induced in dbr1Δ, and lariat intermediates of the self-spliced cox1/cob1 group-II introns increased, validated by lariat-specific RT-PCR 19/38 genes; P < 1.4 × 10^-15
  • Most intron-embedded snoRNAs were among decreased transcripts, and decreased transcripts were enriched for ribosome-biogenesis GO categories 6/10 snoRNAs; P < 2.4 × 10^-7
  • The polysome-to-monosome ratio was significantly higher in dbr1Δ than wild type, with no difference in S6 or eIF2α phosphorylation, suggesting translation is compromised at the elongation step P/M 1.83 (dbr1Δ) vs 1.14 (wild type)
  • 17 of 18 selected lariats were confirmed by lariat-specific RT-PCR followed by Sanger sequencing 17/18 (94%)
Key statistics
  • count 347 transcripts significantly increased and 233 decreased in dbr1Δ; ~57% of increased transcripts were noncoding RNAs (Differential transcript expression, dbr1Δ vs wild type (DESeq, adjusted P < 0.05, |fold change| > 2))
  • count 1399/5361 introns significantly increased; 27 decreased (Differential intronic expression in dbr1Δ (~25% of introns))
  • pvalue P < 2.2 × 10^-16 (Wilcoxon rank sum test for higher intronic expression, its length bias, and lower splicing efficiency in dbr1Δ)
  • pvalue P < 1.4 × 10^-15 (Hypergeometric test: 19/38 mitochondrial genes induced in dbr1Δ)
  • pvalue P < 2.4 × 10^-7 (Hypergeometric test: 6/10 intron-embedded snoRNAs decreased in dbr1Δ)
  • mean P/M ratio 1.83 (dbr1Δ) vs 1.14 (wild type) (Polysome-to-monosome ratio from translational profiling)
  • count 108,683 diagnostic lariat reads; 93,845 adenine branch-point reads; 5060 lariats from 1655 introns; 2842 adenine-only lariats from 1584 introns; 1236 introns at ≥3 reads (LaSSO lariat detection in S. pombe RNA-seq)
  • other 93% of S. pombe introns shorter than the ~200-base average insert size; ~36% smaller than the 49-bp read length; ~47% of S. pombe genes are spliced (Technical basis for preferential recovery of lariats from long introns)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper analyzes RNA-seq data from wild-type and dbr1Δ fission yeast biological replicates, using DESeq for differential transcript/intron expression (adjusted P<0.05, fold change>2), nonparametric rank-based tests (Wilcoxon rank sum) for comparing distributions of expression or splicing efficiency between conditions, a Cochran-Mantel-Haenszel (CMH) test for per-intron splicing-efficiency changes, hypergeometric tests for gene-set/overlap enrichment, and Fisher's exact test for a proportions comparison. Reproducibility between biological replicates was assessed with Pearson correlation coefficients. Results are reported mainly as P-value thresholds and fold-change/ratio values displayed in box plots, scatter plots, and MA plots, with select findings independently validated by RT-PCR and Sanger sequencing.

Replicationbiological Sample sizeComparisons are based on the full sets of profiled transcripts/introns (e.g., 5361 introns); no explicit power calculation or sample-size justification is described beyond stating experiments were biologically repeated. Groupsdbr1Δ (debranching-enzyme deletion) vs wild-type fission yeast Pairingunclear Randomization/blindingnot stated DispersionIQR Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionDESeq adjusted P-values for differential expression; Q-value (FDR-type) cutoff for the CMH splicing-efficiency test
Statistical tests used
Test Applied to n Assumptions
Wilcoxon rank sum test intronic vs transcript expression comparison (Fig. 2A); intronic expression length-bias comparison between dbr1Δ and wild type (Fig. 2D); splicing efficiency (SE) comparison between dbr1Δ and wild type (Fig. 3C) 5361 introns not stated
DESeq (differential expression, adjusted P-value + fold-change cutoff) differential transcript and intron expression between dbr1Δ and wild type (Fig. 2B) not stated (whole transcriptome/intron set) not stated
Hypergeometric test enrichment of mitochondrial genes and snoRNAs among differentially expressed transcripts; overlap between introns with lower SE and introns with higher DESeq-based expression (Fig. 3F,G) e.g., 19/38 mitochondrial genes; 6/10 intron-embedded snoRNAs not stated
Cochran-Mantel-Haenszel (CMH) test per-intron splicing efficiency changes between dbr1Δ and wild type (Fig. 3D) 5361 introns; 638 introns called significant not stated
Fisher's exact test proportion of lariat reads relative to unmapped reads, dbr1Δ vs wild type (Fig. 5A) counts of lariat reads vs total unmapped reads per sample not stated
Pearson's correlation coefficient reproducibility of intronic expression and splicing efficiency between biological replicates (Figs. 2C, 3A,B) 5361 introns not stated
Approaches that could also have been used
  • Several separate hypothesis tests (Wilcoxon rank sum, hypergeometric, Fisher's exact, CMH) are applied to different comparisons throughout the study, each evaluated largely on its own P-value threshold.
    Could also: A joint multiple-testing correction (e.g., Benjamini-Hochberg FDR) applied across the full family of reported comparisons — Treating the whole set of tests performed in the paper as one family, rather than correcting within each analysis separately, would additionally control the overall false-discovery rate across the study.
  • Differential expression between dbr1Δ and wild type was called using DESeq with an adjusted P-value and a fixed fold-change cutoff (>2).
    Could also: Shrinkage-based effect-size estimation (e.g., DESeq2/apeglm or edgeR with an effect-size-aware significance criterion, such as TREAT/glmTreat testing against a fold-change threshold) — Shrinking noisy fold-change estimates for low-count genes/introns before applying a fold-change cutoff can reduce the influence of large but imprecise fold changes, which is particularly relevant given the length-associated technical bias noted for intron read recovery.
  • Enrichment of specific GO categories (e.g., mitochondrial translation, ribosome biogenesis) among differentially expressed transcripts was assessed with a hypergeometric test on genes passing a significance cutoff.
    Could also: Rank-based gene set enrichment analysis (e.g., GSEA) using the full ranked list of fold changes or test statistics — Using the continuous ranking of all transcripts rather than a hard significance cutoff can capture coordinated but individually sub-threshold shifts within a gene set.
  • Reproducibility between biological replicates was quantified using Pearson's correlation coefficient on intronic expression and splicing efficiency values.
    Could also: Spearman's rank correlation or a concordance correlation coefficient (Lin's CCC) — Spearman's correlation is robust to non-normal or skewed count-based distributions (common in RNA-seq intensity data), and CCC additionally penalizes systematic deviation from the identity line, which the paper's replicate comparison is qualitatively assessing.
  • Splicing efficiency changes per intron were tested with a Cochran-Mantel-Haenszel test on the underlying exon-intron/exon-exon junction read counts.
    Could also: Beta-binomial or logistic regression modeling of splicing efficiency as a proportion (e.g., as implemented in tools such as DEXSeq or MISO/rMATS-style frameworks) — A regression-based proportion model can incorporate overdispersion and covariates (e.g., intron length) directly, which is relevant given the length-related bias the authors separately investigate.
  • Distributional comparisons (e.g., intronic vs transcript expression, dbr1Δ vs wild-type SE) were summarized in box plots with associated Wilcoxon rank-sum P-values.
    Could also: Reporting an accompanying effect-size measure (e.g., rank-biserial correlation or Hodges-Lehmann estimator of median difference) alongside the P-value — A nonparametric effect-size estimate would convey the magnitude of the shift between distributions in addition to its statistical significance, complementing the box-plot visualization.
Software: DESeq · Bowtie · WebLogo

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

gse50246_primary_headline_lariat_numbers
Reported
diagnostic_lariat_reads_total=108683; adenine_specific_reads=93845; distinct_lariats=2842; introns_with_branchpoint=1584
Reproduced
diagnostic_lariat_reads_total=120367; adenine_specific_reads=98423; distinct_lariats=3382; introns_with_branchpoint=1843
within tolerance
gse48594_awan_reanalysis_recovery
Reported
paper_LaSSO_reanalysis_reads=97072; paper_LaSSO_reanalysis_introns=1268; original_Awan_2013_reads=37008; original_Awan_2013_introns=813
Reproduced
total_lariat_reads_all_branchpoints_short_plus_long=135522; adenine_specific_reads_short_plus_long=103907; introns_with_adenine_branchpoint=1637; introns_with_any_branchpoint=2064
partial
gse30611_human_crossspecies_validation
Reported
unmapped_reads_used_as_pipeline_input=592367167; diagnostic_lariat_reads_before_filtering=51469; adenine_based_lariat_reads_after_stringent_filtering=586; introns_identified=287; exon_skipping_events_identified=34; custom_human_lariat_database_size_reported=~947.26 GB, 4,206,823,861 lariat sequence entries
Reproduced
nicht durchgefuehrt — This is a secondary cross-species validation, not the paper's primary headline result (which is GSE50246, already reproduced above with a strong within-tol grade). The paper's own Methods describe a ~947 GB custom human lariat sequence database and alignment of 3.57 billion raw reads (46 ENA runs, 592M of them carried forward as unmapped input to the lariat search) across 16 human tissues -- roughly 3 orders of magnitude larger in storage and at least 10x larger in read volume than either S. pombe analysis performed in this room, which already used the full available job time/scratch budget fo
m.public.grade.not-attempted

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 68/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

Solid partial reproduction, deviations on our side. Two of the three datasets were downloaded and reprocessed end-to-end from public raw data: the primary GSE50246 headline numbers reproduce at 120,367 / 98,423 / 3,382 / 1,843 against the reported 108,683 / 93,845 / 2,842 / 1,584 (+5% to +19%, all same direction), and the Awan reanalysis confirms the improved-recovery claim multi-fold (103,907 vs 37,008 reads; 1,637 vs 813 introns) though the intron count overshoots the paper's 1,268 by 29%. The residual gap is best explained by the fact that the authors released only LaSSO.R (database construction) — the alignment, filtering and counting logic was reimplemented from Methods prose under Bowtie 1.2.3 instead of 0.12.7 — so this is our-method + partial-code-release, not evidence against the authors' numbers; the biological control (dbr1Δ carrying 119,552 of 120,367 reads vs 815 in WT) matches expectation exactly. The one genuine authors'-side gap is documentation: the human GSE30611 filtering that takes 51,469 candidates down to 586 reads is never specified, and that sub-result was transparently declared not-attempted (947 GB database) rather than fabricated. Overall yellow: reproducible in substance and conclusion, not 1:1 in counts, with the shortfall traceable to incomplete code/parameter release.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.