Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
G-quadruplexes in promoters throughout the human genome.
PMID 17169996 · PMC1802602 · Nucleic acids research · 2007 · 8 claims · 6 setups
Promoter regions (1 kb upstream of TSS) are significantly enriched in quadruplex motifs (PQS) relative to the rest of the genome
-
Full-text index only
Searching for SNPs with cloud computing.
PMID 19930550 · PMC3091327 · Genome biology · 2009 · 8 claims · 4 setups
Crossbow combines the Bowtie short-read aligner and SOAPsnp SNP caller into a seamless, automatic Hadoop/MapReduce pipeline for whole-genome resequencing analysis
-
Full-text index only
Ab initio identification of human microRNAs based on structure motifs.
PMID 18088431 · PMC2238772 · BMC bioinformatics · 2007 · 8 claims · 7 setups
MiRPred predicts miRNA precursors ab initio using only predicted secondary structure motifs, ignoring nucleotide sequence
-
Full-text index only
ProMiR II: a web server for the probabilistic prediction of clustered, nonclustered, conserved and nonconserved microRNAs.
PMID 16845048 · PMC1538778 · Nucleic acids research · 2006 · 6 claims · 4 setups
ProMiR II improves on the original ProMiR by integrating free energy, G/C ratio, conservation score and entropy for more controllable miRNA prediction
-
Has reproduction · 85
Digital sorting of complex tissues for cell type-specific gene expression profiles.
PMID 23497278 · PMC3626856 · BMC bioinformatics · 2013 · 8 claims · 8 setups
The Digital Sorting Algorithm (DSA) deconvolves mixed tissue expression into cell type-specific profiles using only marker genes, without requiring prior knowledge of cell type frequencies or in vitro pure-cell profiles.
-
Full-text index only
Computational disease gene identification: a concert of methods prioritizes type 2 diabetes and obesity candidate genes.
PMID 16757574 · PMC1475747 · Nucleic acids research · 2006 · 6 claims · 8 setups
Applying seven independent computational disease-gene prioritization methods in concert to 9556 positional candidate genes identifies a prioritized set of likely T2D and obesity candidate genes
-
Full-text index only
X:Map: annotation and visualization of genome structure for Affymetrix exon array analysis.
PMID 17932061 · PMC2238884 · Nucleic acids research · 2008 · 7 claims · 4 setups
X:Map is a genome annotation database that maps every Affymetrix exon array probeset to Ensembl genome features (genes, ESTs, GenScan predictions) and supports both high-throughput and gene-centric analysis.
-
Full-text index only
Genome-wide prioritization of disease genes and identification of disease-disease associations from an integrated human functional linkage network.
PMID 19728866 · PMC2768980 · Genome biology · 2009 · 6 claims · 6 setups
Integrating 16 genomic features (32 sub-features) via a naïve Bayes classifier produces a genome-scale FLN of 21,657 human genes and 22,388,609 weighted links that outperforms any individual data source for inferring functional linkages.
-
Has reproduction · 85
Prediction of condition-specific regulatory genes using machine learning.
PMID 32329779 · PMC7293043 · Nucleic acids research · 2020 · 8 claims · 6 setups
ConSReg integrates expression, DAP-seq TF-DNA binding, and ATAC-seq open chromatin data into machine learning models to predict condition-specific regulatory genes
-
Full-text index only
Random amino acid mutations and protein misfolding lead to Shannon limit in sequence-structure communication.
PMID 18769673 · PMC2518838 · PloS one · 2008 · 8 claims · 6 setups
The protein sequence-structure map behaves as a noisy digital communication channel whose capacity C exceeds the transmission rate R for native structures, satisfying Shannon's noisy channel theorem
-
Full-text index only
Cancer-specific high-throughput annotation of somatic mutations: computational prediction of driver missense mutations.
PMID 19654296 · PMC2763410 · Cancer research · 2009 · 7 claims · 7 setups
CHASM, a Random Forest-based computational method, was developed to identify and prioritize missense mutations likely to be functional drivers of tumor cell proliferation.
-
Has reproduction · 90
Prioritized mass spectrometry increases the depth, sensitivity and data completeness of single-cell proteomics.
PMID 37012480 · PMC10172113 · Nature methods · 2023 · 8 claims · 5 setups
pSCoPE (prioritized precursor selection via MaxQuant.Live) increases sensitivity, data completeness, and proteome coverage more than twofold over shotgun single-cell proteomics
-
Has reproduction
Using random walks to identify cancer-associated modules in expression data.
PMID 24128261 · PMC4015830 · BioData mining · 2013 · 8 claims · 8 setups
Walktrap-GM, a random-walk community detection algorithm adapted with stopping criteria (maximum modularity, maximum size, maximum module score), identifies modules significantly enriched with cancer genes in expression-weighted interaction networks.
-
Full-text index only
Leveraging two-way probe-level block design for identifying differential gene expression with high-density oligonucleotide arrays.
PMID 15099405 · PMC411067 · BMC bioinformatics · 2004 · 7 claims · 2 setups
Two-way ANOVA and Mack-Skillings tests on probe-level data with FDR control are substantially more powerful than t-test/Wilcoxon on probe-set level data for detecting differential expression
-
Has reproduction · 58
A comparative study of techniques for differential expression analysis on RNA-Seq data.
PMID 25119138 · PMC4132098 · PloS one · 2014 · 8 claims · 8 setups
edgeR performs slightly better than DESeq and Cuffdiff2 in terms of the ability to uncover true positives.