Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
CorGen--measuring and generating long-range correlations for DNA sequence analysis.
PMID 16845099 · PMC1538783 · Nucleic acids research · 2006 · 8 claims · 3 setups
CorGen is a web server that measures long-range correlations in DNA sequences and generates random sequences with the same (or user-specified) correlation and composition parameters
-
Full-text index only
Skittle: a 2-dimensional genome visualization tool.
PMID 20042093 · PMC2817707 · BMC bioinformatics · 2009 · 7 claims · 6 setups
Skittle is a 2D genome visualization tool combining a color-coded Nucleotide Display, a Repeat Map, a Repeat Overview, and an Alignment Cylinder to reveal genomic patterns at multiple scales
-
Full-text index only
Inverse symmetry in complete genomes and whole-genome inverse duplication.
PMID 19898631 · PMC2771390 · PloS one · 2009 · 8 claims · 5 setups
Reverse and complement symmetries are essentially absent in genomic sequences at all scales.
-
Full-text index only
FeatureScan: revealing property-dependent similarity of nucleotide sequences.
PMID 16845077 · PMC1538849 · Nucleic acids research · 2006 · 6 claims · 5 setups
FeatureScan transforms nucleotide sequences into numerical signals of physico-chemical/conformational properties and compares them via a convolution/correlation (Fourier transform) method rather than comparing letters
-
Full-text index only
Effect of the assignment of ancestral CpG state on the estimation of nucleotide substitution rates in mammals.
PMID 18826599 · PMC2576242 · BMC evolutionary biology · 2008 · 7 claims · 4 setups
CpG/non-CpG assignment based on presence/absence of a CpG dinucleotide seriously biases substitution rate estimates, overestimating CpG changes and underestimating non-CpG changes.
-
Full-text index only
Identifying the important HIV-1 recombination breakpoints.
PMID 18787691 · PMC2522274 · PLoS computational biology · 2008 · 8 claims · 3 setups
Local sequence identity between co-packaged parental RNAs strongly influences the probability of strand-transfer/breakpoint location, with fewer breakpoints occurring near mismatches
-
Full-text index only
Grammar-based distance in progressive multiple sequence alignment.
PMID 18616828 · PMC2478692 · BMC bioinformatics · 2008 · 7 claims · 3 setups
A grammar-based (LZ complexity) distance metric can be used to determine the order in which sequences are progressively pairwise aligned
-
Has reproduction · 86
RNASEQR--a streamlined and accurate RNA-seq sequence analysis program.
PMID 22199257 · PMC3315322 · Nucleic acids research · 2012 · 8 claims · 7 setups
RNASEQR is a new RNA-seq mapper/aligner that combines a BWT-based (Bowtie) transcriptomic/genomic alignment with hash-based BLAT local alignment in three sequential steps: transcriptome mapping, novel exon detection, and anchor-and-align novel splice junction identification.
-
Full-text index only
Improving the specificity of exon prediction using comparative genomics.
PMID 18831778 · PMC2559877 · BMC genomics · 2008 · 8 claims · 6 setups
A log-odds ratio scoring method based on codon conservation across human-mouse/human-dog alignments and adjacent-codon dependency can classify putative exons as coding vs non-coding.
-
Full-text index only
Applications for protein sequence-function evolution data: mRNA/protein expression analysis and coding SNP scoring tools.
PMID 16912992 · PMC1538848 · Nucleic acids research · 2006 · 7 claims · 8 setups
PANTHER HMMs built from family/subfamily multiple sequence alignments can classify novel protein sequences into functional groups based on statistically significant HMM match scores
-
Full-text index only
The use of coded PCR primers enables high-throughput sequencing of multiple homolog amplification products by 454 parallel sequencing.
PMID 17299583 · PMC1797623 · PloS one · 2007 · 6 claims · 4 setups
5′-tagged PCR primers enable pooling of homologous PCR products from multiple sources into a single GS20 run with accurate post-hoc assignment of sequences to source
-
Full-text index only
An empirical study of choosing efficient discriminative seeds for oligonucleotide design.
PMID 19958494 · PMC2788383 · BMC genomics · 2009 · 8 claims · 3 setups
The spaced seed is the most efficient discriminative seed for oligonucleotide design among the five algorithms tested.
-
Has reproduction · 65
FusionQ: a novel approach for gene fusion detection and quantification from paired-end RNA-Seq.
PMID 23768108 · PMC3691734 · BMC bioinformatics · 2013 · 8 claims · 8 setups
FusionQ is a novel tool that detects gene fusions, constructs chimerical transcript structures, and estimates their abundances from paired-end RNA-Seq data.
-
Full-text index only
A computational screen for type I polyketide synthases in metagenomics shotgun data.
PMID 18953415 · PMC2568958 · PloS one · 2008 · 8 claims · 6 setups
Combining HMM domain searches with maximum-likelihood phylogenetic trees can discriminate true PKS I sequences from evolutionarily related but functionally different enzymes (e.g., FAS I) in metagenomic data.
-
Full-text index only
Genetic variation at hair length candidate genes in elephants and the extinct woolly mammoth.
PMID 19747392 · PMC2754481 · BMC evolutionary biology · 2009 · 8 claims · 5 setups
The coding sequence of FGF5 is not the critical determinant of hair length differences among elephantids, including the woolly mammoth.
-
Full-text index only
Sequence variation in G-protein-coupled receptors: analysis of single nucleotide polymorphisms.
PMID 15784611 · PMC1069129 · Nucleic acids research · 2005 · 7 claims · 8 setups
Position-specific phylogenetic features describing evolutionary conservation at a site (e.g. SIFT score, normalized site entropy, residue frequency change) are the best individual discriminators of disease-causing versus neutral GPCR mutations.
-
Has reproduction · 100
Integrative transcriptome sequencing identifies trans-splicing events with important roles in human embryonic stem cell pluripotency.
PMID 24131564 · PMC3875859 · Genome research · 2014 · 8 claims · 8 setups
TSscan, a computational pipeline integrating long- and short-read transcriptome sequencing from multiple hESC lines, can detect trans-splicing while minimizing false positives from experimental artifacts and genetic rearrangements.
-
Has reproduction · 67
A consensus approach to vertebrate de novo transcriptome assembly from RNA-seq data: assembly of the duck (Anas platyrhynchos) transcriptome.
PMID 25009556 · PMC4070175 · Frontiers in genetics · 2014 · 8 claims · 8 setups
Multiple k-mer (MK) assemblies are more complete than single k-mer (SK) assemblies, showing higher reads-mapped-back-to-transcripts (RMBT) and higher CEGMA complete-gene percentages for all three tools.
-
Has reproduction · 67
Evaluating native-like structures of RNA-protein complexes through the deep learning method.
PMID 36828844 · PMC9958188 · Nature communications · 2023 · 8 claims · 7 setups
DRPScore identifies native-like RNA-protein structures with higher success rates than ITScore-PR, DARS-RNP, and 3dRPC across bound and unbound testing sets.
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes