IsoSCM: improved and alternative 3' UTR annotation using multiple change-point inference.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
IsoSCM (a third-party but paper-authored GitHub tool) was built from source and run per its documented usage (assemble -s reverse_forward -coverage true) on the paper's own calibration dataset (mouse brain, SRR594393 from GSE41637/SRP016501), after TopHat2 alignment to Ensembl GRCm38.73 exactly as specified in Methods. The pipeline runs end-to-end and produces a full genome-scale 3' UTR assembly (brain.gtf etc). This is a 1:1 methodological reproduction of the core tool + its primary input, but the single headline number the paper reports for this run ('11,000 3' terminal exon models') could not be matched exactly: our best-effort interpretations give either 41,208 raw 3p_exon GTF entries or 17,044 distinct loci with >=1 3p_exon, both the right order of magnitude but off by 1.5x-3.7x, because the repo does not ship the script that produced the published count. The paper's headline comparison figures (IsoSCM vs Cufflinks vs Scripture counts; TPR curves; PAS/3'-seq enrichment) are NOT reproduced at all -- they require custom benchmark/matching code that is not part of the public repository, so attempting them would mean silently reimplementing undocumented methodology rather than genuinely reproducing it. Dataset profiling found the underlying data (GSE41637 mouse subset) fully accessible and of good quality (>90% mapping rate), with one honest discrepancy: the paper's methods text implies 27 mouse RNA-seq datasets were used (1 calibration + 26 additional) but only 26 are actually deposited (heart tissue has 2 of 3 expected replicates). NOT attempted in this pass: assembling the other 25 mouse tissue datasets, running Cufflinks/Scripture, and any of the simulation-based benchmarking -- all out of scope for a single-room pass given the missing comparison scripts, and explicitly flagged rather than fabricated.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-08-06
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-06
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe authors ask whether multiple change-point inference applied to RNA-seq read-coverage patterns can annotate 3' terminal exon boundaries more accurately than existing transcript assemblers, and in particular whether it can resolve tandem 3' UTR isoforms generated by alternative cleavage and polyadenylation (APA) that current tools cannot represent.
- ★ Existing ab initio assemblers (Cufflinks, Scripture) annotate at most one 3' boundary per terminal exon and therefore cannot assemble coexpressed tandem 3' UTR isoforms. finding
- ★ IsoSCM, a transcript assembly method incorporating Bayesian multiple change-point inference over RNA-seq coverage, annotates 3' termini with higher sensitivity and specificity than existing methods on simulated and genuine data sets. method
- ★ Transitions ('change points') in RNA-seq coverage depth mark terminal exon boundaries, and nested tandem isoforms produce a characteristic 'step-like' coverage pattern. mechanism
- ★ Constraining change points to monotonically decreasing coverage across sequential segments (fold-change threshold phi) suppresses spurious change points caused by local sequencing/mappability biases while retaining true polyadenylation sites. method
- ★ IsoSCM recovers known patterns of tissue-regulated APA from conventional RNA-seq data. finding
- ★ Local gaps in RNA-seq coverage arising from repetitive elements, sequence-specific and positional biases, and multimapping ambiguity fragment long 3' UTR assemblies produced by existing tools. finding
- IsoSCM is released as stand-alone open software and source code at https://github.com/shenkers/isoscm. resource
- Coverage is modeled as Negative Binomial NB(p,r) with an uninformative conjugate beta(1,1) prior, permitting analytical integration of the segment marginal likelihood, and the optimal segmentation is computed by dynamic programming in O(n^2) operations. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Computational transcript assembly / 3' UTR annotation (IsoSCM change-point segmentation) | Mapped RNA-seq reads from mouse and human genomes | none | Inferred change-point locations, terminal exon models, and splice graph of exon boundaries and spliced-read connections | IsoSCM (https://github.com/shenkers/isoscm); requires splice-aware alignment software upstream |
| Benchmarking on simulated RNA-seq data | Simulated coverage data sets (including toy simulated 3' UTR coverage with a local low-coverage region) | none | Sensitivity and specificity of 3' terminus annotation; constrained vs. unconstrained change-point configurations | — |
| Benchmarking on genuine RNA-seq data sets vs. Cufflinks and Scripture | Tissue RNA-seq (mouse/human genes including Kif3a, Hdlbp, Ict1) | none | Accuracy/completeness of assembled 3' terminal exon models compared with Ensembl 73 annotation | Cufflinks; Scripture; Ensembl 73 annotation |
| 3'-seq (3' end sequencing) used as orthogonal support for inferred polyadenylation sites | Ict1 locus (tandem alternative 3' ends) | none | Presence of 3' end reads at inferred tandem polyadenylation sites | — |
| Tissue-specific APA analysis from RNA-seq | Mouse and human tissues | none (tissue comparison) | Recovery of known tissue-regulated alternative polyadenylation / relative tandem 3' UTR isoform usage | — |
| Genome browser visualization of RNA-seq coverage with repeat annotation | Kif3a and Hdlbp loci | none | Read coverage depth across 3' UTRs, coincidence of coverage gaps with annotated repetitive elements | — |
- ▲ IsoSCM annotates 3' termini with higher sensitivity and specificity than existing methods on both simulated and genuine data sets.
- – For Hdlbp, which has tandem polyadenylation sites annotated in Ensembl 73, neither Cufflinks nor Scripture assembles the short isoform; both report only the long isoform.
- ▼ At Kif3a, a low-coverage region coinciding with an annotated repetitive element causes Cufflinks and Scripture to report fragmented 3' UTR assemblies.
- – On toy data, unconstrained change-point inference reports both a spurious change point at a local coverage dip and the true distal polyadenylation site, whereas the constrained formulation reports only the polyadenylation site.
- – Tandem alternative 3' ends inferred at Ict1 from RNA-seq coverage are both supported by 3'-seq data.
- – IsoSCM recovers known patterns of tissue-regulated APA.
- ▲ In prior related work, refinement of terminal exons extended thousands of 3' UTR models in the mouse and human genomes, with a substantial proportion of extensions showing tissue-specific expression. thousands of 3' UTR models
- – In the RGASP comparative assessment, transcript termini outputs from all tested algorithms were sufficiently inaccurate that a relaxed exon-correctness criterion evaluating only the 5' boundary of the 3' terminal exon was used.
- count 14 (Number of transcript assembly / exon identification algorithms compared in the RGASP assessment)
- fold_change twofold (Parameterization phi = 2 requires coverage of sequential segments to drop at least twofold at each change point (phi = 1 requires strictly decreasing coverage))
- other O(n^2) (Operations required to recursively compute the dynamic programming tables Q(t) and R(t) for n coverage observations)
- other NB(p,r) with prior beta(1,1) (Negative Binomial emission distribution with uninformative conjugate prior used to model per-position coverage)
- count thousands of repetitive elements genome wide (Repetitive elements harbored in UTRs, cited as a source of coverage gaps)
- other Ensembl 73 (Reference annotation containing short/long 3' UTR isoform transcript models for Hdlbp)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a computational-methods paper describing IsoSCM, an algorithm for annotating 3′ UTR boundaries and alternative polyadenylation isoforms from RNA-seq read coverage. Rather than classical inferential hypothesis testing, it uses a Bayesian multiple change-point framework (based on Fearnhead 2006) with a Negative Binomial likelihood for per-position read depth and a dynamic-programming solution for the maximum marginal-likelihood segmentation. Performance is evaluated by comparing IsoSCM's annotations to existing tools (Cufflinks, Scripture) on simulated and genuine RNA-seq data using sensitivity and specificity, along with illustrative genome-browser examples; the excerpted text does not present a quantitative results/statistics section with numeric outcomes.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Bayesian multiple change-point inference (marginal-likelihood maximization via dynamic programming) | identification of 3′ UTR boundary/polyadenylation change points from RNA-seq coverage | — | stated |
| sensitivity and specificity comparison against existing assembly tools (Cufflinks, Scripture) | evaluation of 3′ terminal exon/UTR annotation accuracy on simulated and genuine RNA-seq data sets | — | not stated |
-
Per-position read coverage was modeled with a Negative Binomial distribution within the Bayesian change-point framework.↳ Could also: count-based RNA-seq dispersion-shrinkage frameworks such as DESeq2 or edgeR's negative binomial GLMs, or a quasi-Poisson model — these are widely used alternatives for modeling overdispersed sequencing counts and could also provide a familiar statistical framework if downstream differential isoform-usage testing were desired.
-
Change-point locations were found by maximizing marginal likelihood via a custom dynamic-programming recursion (following Fearnhead 2006).↳ Could also: other established multiple change-point algorithms such as PELT, binary segmentation with a BIC-type penalty, or circular binary segmentation — these methods address the same class of multiple change-point problems and could also be applied to detect coverage transitions, offering different trade-offs in computational complexity versus exactness.
-
An uninformative conjugate Beta(1,1) prior was placed on the Negative Binomial's proportion parameter.↳ Could also: a weakly informative prior calibrated from genome-wide read-depth distributions — an empirically informed prior could also help stabilize change-point estimates in low-coverage regions while still letting the observed data dominate the posterior.
-
A monotonic fold-change constraint (ϕ) was used to distinguish genuine boundaries from local coverage fluctuations.↳ Could also: a hidden Markov model with state-dependent emission and transition probabilities — an HMM formulation could also encode structural constraints on the sequence of coverage states while probabilistically smoothing local noise, as an alternative way to separate true transitions from artifacts.
-
Method accuracy relative to Cufflinks and Scripture was summarized via sensitivity and specificity plus illustrative locus examples (e.g., Kif3a, Hdlbp).↳ Could also: reporting precision-recall curves, F1 score, or Matthews correlation coefficient alongside sensitivity/specificity — these metrics can also summarize classifier-style performance and are often informative when true boundary positions are rare relative to non-boundary positions (class imbalance).
-
Comparisons between IsoSCM and other assembly tools were presented qualitatively (genome-browser tracks, aggregate sensitivity/specificity) rather than with a formal paired statistical test across loci.↳ Could also: a paired statistical comparison such as McNemar's test or a paired Wilcoxon signed-rank test across evaluated loci — such tests could also quantify whether observed accuracy differences between tools are unlikely to arise by chance, complementing the descriptive comparison.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.