Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 62
Application of alternative de novo motif recognition models for analysis of structural heterogeneity of transcription factor binding sites: a case study of FOXA2 binding sites.
PMID 34547062 · PMC8408018 · Vavilovskii zhurnal genetiki i selektsii · 2021 · 6 claims · 7 setups
Combining four de novo models (PWM, diPWM, BaMM, InMoDe) significantly increases the fraction of recognized peaks versus PWM alone (by 26.3%).
-
Full-text index only
Fast and systematic genome-wide discovery of conserved regulatory elements using a non-alignment based approach.
PMID 15693947 · PMC551538 · Genome biology · 2005 · 7 claims · 8 setups
FastCompare, a non-alignment-based, linear-time algorithm, computes a genome-wide conservation score for all k-mers (7-9 nt) between two genomes to identify conserved regulatory elements
-
Full-text index only
The stem cell population of the human colon crypt: analysis via methylation patterns.
PMID 17335343 · PMC1808490 · PLoS computational biology · 2007 · 8 claims · 3 setups
A coalescent-based, full probabilistic model with MCMC Bayesian inference provides a more powerful alternative to prior forward-simulation approaches for analyzing methylation pattern data from crypts.
-
Full-text index only
Importance sampling for the infinite sites model.
PMID 18976228 · PMC2832804 · Statistical applications in genetics and molecular biology · 2008 · 7 claims · 2 setups
A new importance sampling proposal distribution for the ISM, derived from a new result on exact sampling from a single segregating site, generally shows greater efficiency than the GT and SD proposals.
-
Full-text index only
Aberrant 5' splice sites in human disease genes: mutation pattern, nucleotide structure and comparison of computational tools that predict their utilization.
PMID 17576681 · PMC1934990 · Nucleic acids research · 2007 · 8 claims · 4 setups
Cryptic 5'ss are best predicted by computational algorithms that accommodate nucleotide dependencies (e.g., Markov model, maximum entropy, maximum dependence decomposition) rather than by weight-matrix models
-
Full-text index only
Target SNP selection in complex disease association studies.
PMID 15248903 · PMC487897 · BMC bioinformatics · 2004 · 7 claims · 3 setups
A computational pipeline can retrieve gene sequence, collect SNP variation data, and annotate SNPs falling in functional motifs (promoter, exon-intron structure, AU-rich elements, TF binding sites, splice sites) with expression in target tissue
-
Full-text index only
Computational identification of transcriptional regulatory elements in DNA sequence.
PMID 16855295 · PMC1524905 · Nucleic acids research · 2006 · 8 claims · 3 setups
Weight matrix (PWM/PSSM) models of TF binding sites are grounded in biophysical theory of protein-DNA interactions, with position weights corresponding to log-odds contributions to binding free energy
-
Has reproduction · 79
Interpretable prediction models for widespread m6A RNA modification across cell lines and tissues.
PMID 37995291 · PMC10697738 · Bioinformatics (Oxford, England) · 2023 · 7 claims · 6 setups
CLSM6A, a CNN-based model set, predicts single-nucleotide-resolution m6A RNA modification sites across eight cell lines and three tissues in H. sapiens
-
Full-text index only
Genome-wide survey of allele-specific splicing in humans.
PMID 18518984 · PMC2427040 · BMC genomics · 2008 · 8 claims · 5 setups
A genome-wide computational scan identified 30,977 SNPs located within predicted splicing regulatory sequences (donor sites, acceptor sites, branch points, and ESEs)
-
Full-text index only
CTCF binding site classes exhibit distinct evolutionary, genomic, epigenomic and transcriptomic features.
PMID 19922652 · PMC3091324 · Genome biology · 2009 · 8 claims · 8 setups
CTCF binding sites can be classified into three occupancy-based classes (LowOc, MedOc, HighOc) based on similarity to the CTCF PWM motif
-
Full-text index only
SNP@Promoter: a database of human SNPs (single nucleotide polymorphisms) within the putative promoter regions.
PMID 18315851 · PMC2259403 · BMC bioinformatics · 2008 · 8 claims · 4 setups
SNP@Promoter is a database of human SNPs within putative promoter regions and predicted transcription factor binding sites
-
Has reproduction · 74
Transcriptome profiling of Giardia intestinalis using strand-specific RNA-seq.
PMID 23555231 · PMC3610916 · PLoS computational biology · 2013 · 8 claims · 8 setups
Most of the G. intestinalis genome is transcribed in in vitro-grown trophozoites, but at vastly different expression levels.
-
Full-text index only
miRGator: an integrated system for functional annotation of microRNAs.
PMID 17942429 · PMC2238850 · Nucleic acids research · 2008 · 8 claims · 8 setups
miRGator integrates target prediction, functional enrichment analysis (GO/pathway/disease), and expression data (miRNA/mRNA/protein) into one system for functional annotation of miRNAs
-
Full-text index only
In silico whole-genome screening for cancer-related single-nucleotide polymorphisms located in human mRNA untranslated regions.
PMID 17201911 · PMC1774567 · BMC genomics · 2007 · 8 claims · 5 setups
A computational EST-based pipeline can identify UTR-SNPs that are statistically over-represented in cancerous versus normal tissue libraries
-
Full-text index only
Decoding of superimposed traces produced by direct sequencing of heterozygous indels.
PMID 18654614 · PMC2429969 · PLoS computational biology · 2008 · 7 claims · 3 setups
A dynamic programming method (implemented as web app Indelligent) can decode superimposed allelic sequences from a single mixed trace, using only the observed string of ambiguous peak calls, without a reference sequence or reverse trace.
-
Full-text index only
TRED: a Transcriptional Regulatory Element Database and a platform for in silico gene regulation studies.
PMID 15608156 · PMC539958 · Nucleic acids research · 2005 · 8 claims · 5 setups
TRED is a database collecting both cis-regulatory elements (promoters) and trans-regulatory elements (transcription factor binding/regulation data) with linked access.
-
Full-text index only
EGASP: Introduction.
PMID 16925831 · PMC1810546 · Genome biology · 2006 · 8 claims · 5 setups
Computational gene finding methods, when compared to the GENCODE golden standard annotation, show that the human genome annotation is nearly complete in terms of novel protein-coding loci.
-
Full-text index only
Computational analysis of splicing errors and mutations in human transcripts.
PMID 18194514 · PMC2234086 · BMC genomics · 2008 · 8 claims · 4 setups
Retained introns are significantly shorter than constitutively spliced introns
-
Full-text index only
Finding signals that regulate alternative splicing in the post-genomic era.
PMID 12429065 · PMC244920 · Genome biology · 2002 · 8 claims · 8 setups
Alternative splicing generates protein and regulatory diversity from a limited number of genes and modulates isoform levels in a cell-context-specific manner
-
Full-text index only
Non-EST based prediction of exon skipping and intron retention events using Pfam information.
PMID 16204458 · PMC1243800 · Nucleic acids research · 2005 · 7 claims · 5 setups
A novel ab initio method predicts exon skipping and intron retention events using only Pfam domain annotation, via a Viterbi-like dynamic programming algorithm applied to the Pfam alignment.