Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
EGASP: the human ENCODE Genome Annotation Assessment Project.
PMID 16925836 · PMC1810551 · Genome biology · 2006 · 8 claims · 6 setups
Best-performing computational gene prediction methods correctly predict at least one transcript for close to 70% of annotated genes in the ENCODE regions.
-
Has reproduction · 67
Generative and integrative modeling for transcriptomics with formalin fixed paraffin embedded material.
PMID 41029822 · PMC12486589 · Journal of translational medicine · 2025 · 8 claims · 5 setups
fRNA-seq transcript counts are best fit by the negative binomial distribution, with little evidence supporting zero-inflated extensions
-
Full-text index only
Exogean: a framework for annotating protein-coding genes in eukaryotic genomic DNA.
PMID 16925841 · PMC1810556 · Genome biology · 2006 · 8 claims · 5 setups
Exogean is a framework using directed acyclic coloured multigraphs (DACMs) to represent biological objects (mRNA, ESTs, protein alignments, exons) and iteratively combine them into complex protein-coding transcript models.
-
Full-text index only
ASPIC: a web resource for alternative splicing prediction and transcript isoforms characterization.
PMID 16845044 · PMC1538898 · Nucleic acids research · 2006 · 8 claims · 2 setups
The ASPIC algorithm, using an optimization procedure that minimizes splice site predictions and transcript isoforms from multiple EST-genome alignments, outperforms other similar AS-prediction tools in sensitivity and selectivity
-
Has reproduction · 68
LaSSO, a strategy for genome-wide mapping of intronic lariats and branch points using RNA-seq.
PMID 24709818 · PMC4079972 · Genome research · 2014 · 8 claims · 8 setups
LaSSO (Lariat Sequence Site Origin) identifies intronic lariat reads and pinpoints branch points genome-wide from RNA-seq data by considering every intronic base as a potential branch point and including all possible exon-skipping lariats.
-
Has reproduction · 50
Viewing RNA-seq data on the entire human genome.
PMID 28979763 · PMC5605993 · F1000Research · 2017 · 6 claims · 3 setups
RNA-Seq Viewer is a web application that visualizes genome-wide expression data from NCBI's SRA and GEO databases using an ideogram across the entire human genome.
-
Has reproduction · 78
annotate_my_genomes: an easy-to-use pipeline to improve genome annotation and uncover neglected genes by hybrid RNA sequencing.
PMID 36472574 · PMC9724561 · GigaScience · 2022 · 7 claims · 8 setups
annotate_my_genomes is an easy-to-use genome-guided pipeline that uses hybrid (PacBio+Illumina) assembled transcripts to distinguish coding genes from long non-coding RNAs and reconcile them with prior annotations.
-
Has reproduction · 71
Hyb: a bioinformatics pipeline for the analysis of CLASH (crosslinking, ligation and sequencing of hybrids) data.
PMID 24211736 · PMC3969109 · Methods (San Diego, Calif.) · 2014 · 8 claims · 6 setups
The 'hyb' pipeline detects, calls, folds and annotates chimeric reads from CLASH high-throughput sequencing data.
-
Full-text index only
Pairagon+N-SCAN_EST: a model-based gene annotation pipeline.
PMID 16925839 · PMC1810554 · Genome biology · 2006 · 7 claims · 5 setups
Pairagon+N-SCAN_EST, using only native alignments, was as accurate as ENSEMBL and ExoGean in the EGASP mRNA/EST evidence assessment
-
Full-text index only
GENCODE: producing a reference annotation for ENCODE.
PMID 16925838 · PMC1810553 · Genome biology · 2006 · 8 claims · 8 setups
GENCODE annotation combines initial manual annotation by HAVANA, experimental validation, and refinement based on results to identify protein-coding genes in ENCODE regions
-
Has reproduction · 79
Enriched domain detector: a program for detection of wide genomic enrichment domains robust against local variations.
PMID 24782521 · PMC4066758 · Nucleic acids research · 2014 · 8 claims · 5 setups
EDD is a new algorithm that detects broad (megabase-size) enrichment domains from ChIP-seq data of widely distributed chromatin proteins such as A- and B-type lamins.
-
Full-text index only
Functional annotation and identification of candidate disease genes by computational analysis of normal tissue gene expression data.
PMID 18560577 · PMC2409962 · PloS one · 2008 · 7 claims · 5 setups
Ranked Coexpression Groups (RCG) built from k=6 nearest coexpressed genes, combined with a majority-rule functional characterization, integrate multiple datasets/coexpression measures to generate high-confidence functional annotation predictions
-
Has reproduction · 68
Bayesian transcriptome assembly.
PMID 25367074 · PMC4397945 · Genome biology · 2014 · 8 claims · 8 setups
Bayesembler, a probabilistic transcriptome assembler built on a Bayesian model of the RNA sequencing process with Gibbs sampling over expressed candidates, abundances and read assignments, is introduced.