Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Integration of text- and data-mining using ontologies successfully selects disease gene candidates.
PMID 15767279 · PMC1065256 · Nucleic acids research · 2005 · 7 claims · 6 setups
Integrating eVOC anatomical ontology-based text-mining of PubMed abstracts with data-mining of gene expression annotation successfully selects and prioritizes candidate disease genes
-
Full-text index only
Metagenomic analysis of human diarrhea: viral detection and discovery.
PMID 18398449 · PMC2290972 · PLoS pathogens · 2008 · 8 claims · 7 setups
Micro-mass sequencing (minimal stool input, minimal purification, ~384 reads/sample) can detect known enteric viruses in diarrhea specimens
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Full-text index only
Clustering of phosphorylation site recognition motifs can be exploited to predict the targets of cyclin-dependent kinase.
PMID 17316440 · PMC1852407 · Genome biology · 2007 · 8 claims · 6 setups
CDK consensus motifs are frequently clustered (closely spaced) in known CDK substrate proteins rather than uniformly distributed
-
Full-text index only
Conserved elements with potential to form polymorphic G-quadruplex structures in the first intron of human genes.
PMID 18187510 · PMC2275096 · Nucleic acids research · 2008 · 8 claims · 6 setups
G-richness downstream of the TSS is strand-biased, concentrated on the nontemplate strand, with a peak at +200 to +300 bp
-
Full-text index only
Genomic views of distant-acting enhancers.
PMID 19741700 · PMC2923221 · Nature · 2009 · 8 claims · 8 setups
Meta-analysis of ~1200 top GWAS SNPs found that in 40% of cases (472/1170) no known exons overlap the linked SNP or its haplotype block, implying noncoding variation causally contributes to many traits.
-
Has reproduction · 30
Minimal metabolic pathway structure is consistent with associated biomolecular interactions.
PMID 24987116 · PMC4299494 · Molecular systems biology · 2014 · 8 claims · 8 setups
MinSpan, a mixed-integer linear optimization algorithm, computes the shortest, linearly independent pathways (sparsest basis of the null space of the stoichiometric matrix S) for genome-scale metabolic networks, which convex approaches (extreme pathways, elementary flux modes) cannot do at genome scale.
-
Has reproduction · 85
Prediction of condition-specific regulatory genes using machine learning.
PMID 32329779 · PMC7293043 · Nucleic acids research · 2020 · 8 claims · 6 setups
ConSReg integrates expression, DAP-seq TF-DNA binding, and ATAC-seq open chromatin data into machine learning models to predict condition-specific regulatory genes
-
Has reproduction · 85
PowerBacGWAS: a computational pipeline to perform power calculations for bacterial genome-wide association studies.
PMID 35338232 · PMC8956664 · Communications biology · 2022 · 8 claims · 8 setups
Two computational approaches (sub-sampling and phenotype-simulation) can be implemented to perform power calculations for bacterial GWAS using existing genome collections, packaged as the PowerBacGWAS pipeline
-
Full-text index only
SNAP: predict effect of non-synonymous polymorphisms on function.
PMID 17526529 · PMC1920242 · Nucleic acids research · 2007 · 7 claims · 8 setups
SNAP, a neural network-based method using sequence-derived information, predicts whether a non-synonymous SNP is neutral or non-neutral for protein function
-
Has reproduction · 58
A comparative study of techniques for differential expression analysis on RNA-Seq data.
PMID 25119138 · PMC4132098 · PloS one · 2014 · 8 claims · 8 setups
edgeR performs slightly better than DESeq and Cuffdiff2 in terms of the ability to uncover true positives.