Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
'Genome design' model and multicellular complexity: golden middle.
PMID 17062620 · PMC1635334 · Nucleic acids research · 2006 · 8 claims · 8 setups
Intermediately expressed human genes are the longest genes genome-wide, in both coding and intronic sequence, longer than housekeeping or tissue-specific genes.
-
Has reproduction · 78
Detecting tipping points of complex diseases by network information entropy.
PMID 38960408 · PMC11221888 · Briefings in bioinformatics · 2024 · 8 claims · 4 setups
NIEE can detect critical states or tipping points in diverse data types, including bulk and single-sample expression data
-
Full-text index only
Computational analysis of the synergy among multiple interacting genes.
PMID 17299419 · PMC1828751 · Molecular systems biology · 2007 · 8 claims · 3 setups
Multivariate synergy of a set of factors with respect to a phenotype can be defined via the maximum-information partition, i.e., comparing the mutual information of the full set to the best achievable sum of mutual information over any partition into disjoint subsets.
-
Has reproduction · 59
Application of Machine Learning in Predicting Hepatic Metastasis or Primary Site in Gastroenteropancreatic Neuroendocrine Tumors.
PMID 37887568 · PMC10605255 · Current oncology (Toronto, Ont.) · 2023 · 8 claims · 7 setups
Multi-gene random forest models classify primary tumor vs. liver metastasis samples with 100% accuracy in training/test cohorts and >90% accuracy in an independent validation cohort
-
Full-text index only
Using several pair-wise informant sequences for de novo prediction of alternatively spliced transcripts.
PMID 16925842 · PMC1810557 · Genome biology · 2006 · 8 claims · 4 setups
MARS, an extension of the Twinscan algorithm, uses multiple pairwise informant genomes to predict human alternatively spliced transcripts de novo without expressed sequence information.
-
Has reproduction · 99
getSequenceInfo: a suite of tools allowing to get genome sequence information from public repositories.
PMID 35804320 · PMC9264741 · BMC bioinformatics · 2022 · 8 claims · 8 setups
getSequenceInfo (gSeqI) allows programmatic (CLI) or GUI-based retrieval of sequence data and metadata from GenBank, RefSeq, and ENA across Linux, MacOS, and Windows.
-
Full-text index only
CONTRAST: a discriminative, phylogeny-free approach to multiple informant de novo gene prediction.
PMID 18096039 · PMC2246271 · Genome biology · 2007 · 8 claims · 5 setups
CONTRAST predicts exact coding region structures for 65% more human genes than the previous state-of-the-art de novo predictor (N-SCAN)
-
Full-text index only
Improved mutation tagging with gene identifiers applied to membrane protein stability prediction.
PMID 19758467 · PMC2745585 · BMC bioinformatics · 2009 · 8 claims · 4 setups
MutationTagger achieves 87% F-measure for the mutation retrieval task on a benchmark dataset
-
Has reproduction · 75
FEM: mining biological meaning from cell level in single-cell RNA sequencing data.
PMID 34909283 · PMC8641482 · PeerJ · 2021 · 7 claims · 5 setups
The FEM algorithm converts each cell's gene expression matrix (GEM) into a functional expression matrix by applying Fisher's exact test enrichment per cell and per gene set, then encoding adjusted p-values as information content.
-
Has reproduction · 75
An informatics research platform to make public gene expression time-course datasets reusable for more scientific discoveries.
PMID 33247935 · PMC7698665 · Database : the journal of biological databases and curation · 2020 · 8 claims · 6 setups
GETc enables discovery and visualization of time-course gene expression data and analytical results from GEO
-
Full-text index only
Policy implications of genetic information on regulation under the Clean Air Act: the case of particulate matter and asthmatics.
PMID 16507451 · PMC1392222 · Environmental health perspectives · 2006 · 8 claims · 4 setups
The Clean Air Act mandates protection of sensitive subpopulations, including asthmatics, from air pollution health effects, creating an opening for genetic susceptibility data in regulation.
-
Full-text index only
Evolutionary origins of human apoptosis and genome-stability gene networks.
PMID 18832373 · PMC2577361 · Nucleic acids research · 2008 · 8 claims · 8 setups
The entanglement of DNA repair, chromosome stability and apoptosis gene networks appears with the caspase gene family and the antiapoptotic gene BCL2.
-
Has reproduction · 85
ScLRTC: imputation for single-cell RNA-seq data via low-rank tensor completion.
PMID 34844559 · PMC8628418 · BMC genomics · 2021 · 8 claims · 8 setups
scLRTC imputes dropout entries closest to the original expression values on simulated datasets, outperforming other state-of-the-art methods by SSE and PCC.
-
Full-text index only
Coverage and characteristics of the Affymetrix GeneChip Human Mapping 100K SNP set.
PMID 16680197 · PMC1456318 · PLoS genetics · 2006 · 7 claims · 7 setups
SNPs in the Affymetrix 100K set are undersampled from coding regions (both synonymous and nonsynonymous) and oversampled from regions outside genes, relative to HapMap SNPs
-
Full-text index only
Comprehensive splice-site analysis using comparative genomics.
PMID 16914448 · PMC1557818 · Nucleic acids research · 2006 · 8 claims · 6 setups
Over half a million splice sites were collected from five species (H. sapiens, M. musculus, D. melanogaster, C. elegans, A. thaliana) and classified into four main subtypes: U2-type GT-AG and GC-AG, and U12-type GT-AG and AT-AC.
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Full-text index only
A human genome-wide library of local phylogeny predictions for whole-genome inference problems.
PMID 18710563 · PMC2556685 · BMC genomics · 2008 · 7 claims · 5 setups
A genome-wide library of nearly 16 million local maximum parsimony phylogenies was constructed from HapMap CEU and YRI SNP data across all human autosomes
-
Full-text index only
Network inference and network response identification: moving genome-scale data to the next level of biological discovery.
PMID 20174676 · PMC3087299 · Molecular bioSystems · 2010 · 8 claims · 8 setups
Cellular response to a signal is assumed to involve only specific TRN modules (conditionally active subnetworks) rather than the entire network, providing quantitative tractability
-
Full-text index only
Inferring combinatorial regulation of transcription in silico.
PMID 15647509 · PMC546154 · Nucleic acids research · 2005 · 8 claims · 5 setups
Combining Cluster-Buster (TFBS cluster prediction) with GOSSIP (rigorous GO enrichment statistics with multiple-testing/FDR correction) predicts biological functions controlled by combinatorial transcription factor action, without prior knowledge of factor targets
-
Full-text index only
ARED Organism: expansion of ARED reveals AU-rich element cluster variations between human and mouse.
PMID 17984078 · PMC2238997 · Nucleic acids research · 2008 · 6 claims · 4 setups
ARED Organism and ARED-Integrated are new/updated public databases cataloguing ARE-containing mRNAs/genes in human, mouse and rat