Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
EGenBio: a data management system for evolutionary genomics and biodiversity.
PMID 17118150 · PMC1683573 · BMC bioinformatics · 2006 · 7 claims · 7 setups
EGenBio is a web-based system for integrated management, filtering, curation, and visualization of large-scale genomic sequences, alignments, and phylogenetic trees for evolutionary genomics and biodiversity research.
-
Has reproduction · 73
Detecting aberrant DNA methylation in Illumina DNA methylation arrays: a toolbox and recommendations for its use.
PMID 37218167 · PMC10208159 · Epigenetics · 2023 · 8 claims · 7 setups
Probe-specific upper and lower thresholds for flagging aberrant DNA methylation can be derived from a reference database of >2,000 normal and tumour-adjacent normal samples spanning 25 tissue types.
-
Has reproduction · 71
RNAmountAlign: Efficient software for local, global, semiglobal pairwise and multiple RNA sequence/structure alignment.
PMID 31978147 · PMC6980424 · PloS one · 2020 · 8 claims · 6 setups
RNAmountAlign is the first RNA sequence/structure pairwise alignment algorithm based on incremental ensemble mountain distance, running in O(n^3) time and O(n^2) space for two sequences of length n.
-
Has reproduction · 81
Enabling Single-Cell Drug Response Annotations from Bulk RNA-Seq Using SCAD.
PMID 36762572 · PMC10104628 · Advanced science (Weinheim, Baden-Wurttemberg, Germany) · 2023 · 7 claims · 7 setups
SCAD, a transfer learning framework integrating adversarial discriminative domain adaptation (ADDA), can infer single-cell drug sensitivities by transferring knowledge from bulk RNA-seq pharmacogenomic data (GDSC) to scRNA-seq target domains
-
Full-text index only
ORFer--retrieval of protein sequences and open reading frames from GenBank and storage into relational databases or text files.
PMID 12493080 · PMC139979 · BMC bioinformatics · 2002 · 6 claims · 6 setups
ORFer retrieves protein and nucleic acid sequences and annotations from NCBI GenBank using the XML sequence format
-
Full-text index only
Haplotype analysis of common variants in the BRCA1 gene and risk of sporadic breast cancer.
PMID 15743496 · PMC1064127 · Breast cancer research : BCR · 2005 · 7 claims · 5 setups
A common BRCA1 haplotype (haplotype 2, C A G G) is associated with a modest increase in sporadic breast cancer risk
-
Full-text index only
GeneSeer: a sage for gene names and genomic resources.
PMID 16176584 · PMC1266031 · BMC genomics · 2005 · 7 claims · 4 setups
GeneSeer aggregates gene name synonyms from GenBank, FlyBase, ExPASy, HUGO, ENSEMBL, UCSC and Gene Ontology into a name-translation database that maps any familiar name to a reference (SOFAR) identifier.
-
Full-text index only
Ab initio identification of putative human transcription factor binding sites by comparative genomics.
PMID 15865625 · PMC1097714 · BMC bioinformatics · 2005 · 8 claims · 5 setups
An integrated algorithm combining human-mouse genomic comparison, motif overrepresentation, and coregulation filters (GO annotation and microarray coexpression) can identify candidate transcription factor binding sites genome-wide
-
Full-text index only
GeneTide--Terra Incognita Discovery Endeavor: a new transcriptome focused member of the GeneCards/GeneNote suite of databases.
PMID 15608261 · PMC540076 · Nucleic acids research · 2005 · 8 claims · 7 setups
GeneTide integrates UniGene, DoTS, AceView, BLAT/GeneLoc genomic alignment, and GeneAnnot probe-set data into a unified Consensus/Uniqueness/Score scheme to associate ESTs with GeneCards genes
-
Has reproduction · 76
Bayesian prediction of microbial oxygen requirement.
PMID 26913185 · PMC4743139 · F1000Research · 2013 · 7 claims · 8 setups
A naive Bayesian classifier based on presence/absence of class-associated Pfam-A domains can distinguish three oxygen requirement classes (aerobe, anaerobe, facultative anaerobe) from genome sequence, unlike prior studies that only made pairwise distinctions.
-
Full-text index only
In silico and in vivo splicing analysis of MLH1 and MSH2 missense mutations shows exon- and tissue-specific effects.
PMID 16995940 · PMC1590028 · BMC genomics · 2006 · 8 claims · 6 setups
In silico ESE-prediction algorithms (ESEfinder, RescueESE, PESX) do not reliably predict actual in vivo splicing behavior of missense mutations
-
Full-text index only
Computational disease gene identification: a concert of methods prioritizes type 2 diabetes and obesity candidate genes.
PMID 16757574 · PMC1475747 · Nucleic acids research · 2006 · 6 claims · 8 setups
Applying seven independent computational disease-gene prioritization methods in concert to 9556 positional candidate genes identifies a prioritized set of likely T2D and obesity candidate genes
-
Full-text index only
From genomics to chemical genomics: new developments in KEGG.
PMID 16381885 · PMC1347464 · Nucleic acids research · 2006 · 8 claims · 5 setups
KEGG BRITE has been formally added as a fourth main KEGG database to establish a logical foundation for functional interpretation and pathway reconstruction.
-
Full-text index only
X:Map: annotation and visualization of genome structure for Affymetrix exon array analysis.
PMID 17932061 · PMC2238884 · Nucleic acids research · 2008 · 7 claims · 4 setups
X:Map is a genome annotation database that maps every Affymetrix exon array probeset to Ensembl genome features (genes, ESTs, GenScan predictions) and supports both high-throughput and gene-centric analysis.
-
Full-text index only
piRNABank: a web resource on classified and clustered Piwi-interacting RNAs.
PMID 17881367 · PMC2238943 · Nucleic acids research · 2008 · 6 claims · 4 setups
piRNABank is a web-accessible database storing empirically known piRNA sequences and annotations for human, mouse and rat.
-
Full-text index only
Inference of transcriptional regulation using gene expression data from the bovine and human genomes.
PMID 17683551 · PMC1978505 · BMC genomics · 2007 · 7 claims · 8 setups
Using human reference promoter sequences is a useful approach for studying gene expression regulation in species with limited or non-existing genomic sequence, such as cattle.
-
Full-text index only
The 10 sea urchin receptor for egg jelly proteins (SpREJ) are members of the polycystic kidney disease-1 (PKD1) family.
PMID 17629917 · PMC1934368 · BMC genomics · 2007 · 8 claims · 5 setups
Sea urchins possess 10 SpREJ (PKD1 family) genes, compared to five in humans, all defined by possession of a ~600 residue REJ domain
-
Full-text index only
The Genographic Project public participation mitochondrial DNA database.
PMID 17604454 · PMC1904368 · PLoS genetics · 2007 · 7 claims · 4 setups
The Genographic Project created the largest standardized human mtDNA database to date, comprising 78,590 genotypes from the first 18 months of public participation.
-
Full-text index only
Adaptive discriminant function analysis and reranking of MS/MS database search results for improved peptide identification in shotgun proteomics.
PMID 18788775 · PMC3744223 · Journal of proteome research · 2008 · 7 claims · 4 setups
PeptideProphet's fixed LDA coefficients for combining search scores (Xcorr', ΔCn, SpRank) may not be optimal under all search/instrument conditions.
-
Full-text index only
The mammalian phenotype ontology: enabling robust annotation and comparative analysis.
PMID 20052305 · PMC2801442 · Wiley interdisciplinary reviews. Systems biology and medicine · 2009 · 8 claims · 6 setups
The Mammalian Phenotype (MP) Ontology enables classification and organization of phenotypic data for mouse and other mammalian species in a computationally useful, standardized manner.