Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
A statistical approach for array CGH data analysis.
PMID 15705208 · PMC549559 · BMC bioinformatics · 2005 · 8 claims · 4 setups
Existing model-selection criteria (AIC, BIC, and prior ad hoc penalties) are not well adapted to estimating the number of segments in array CGH data
-
Full-text index only
ProMiR II: a web server for the probabilistic prediction of clustered, nonclustered, conserved and nonconserved microRNAs.
PMID 16845048 · PMC1538778 · Nucleic acids research · 2006 · 6 claims · 4 setups
ProMiR II improves on the original ProMiR by integrating free energy, G/C ratio, conservation score and entropy for more controllable miRNA prediction
-
Has reproduction · 86
Molecular Classification Models for Triple Negative Breast Cancer Subtype Using Machine Learning.
PMID 34575658 · PMC8472680 · Journal of personalized medicine · 2021 · 6 claims · 4 setups
A training gene set of 719 unique upregulated DEGs (subtype-specific) can be used to build ML models that classify TNBC into BLIA, BLIS, MES, and LAR subtypes.
-
Full-text index only
A model-based approach to selection of tag SNPs.
PMID 16776821 · PMC1525207 · BMC bioinformatics · 2006 · 7 claims · 5 setups
The Li and Stephens hidden Markov model outperforms other tested models (simple Markov, two-state HMM, HMM-4D, greedy GR-1/GR-2) in description code-length, tag set information content, and prediction of tagged SNPs.
-
Full-text index only
QuantiSNP: an Objective Bayes Hidden-Markov Model to detect and accurately map copy number variation using SNP genotyping data.
PMID 17341461 · PMC1874617 · Nucleic acids research · 2007 · 8 claims · 7 setups
QuantiSNP (OB-HMM) provides probabilistic quantification of copy number states and significantly improves accuracy of segmental aneuploidy identification and breakpoint mapping relative to existing tools (BeadStudio/Illumina)
-
Has reproduction · 30
Minimal metabolic pathway structure is consistent with associated biomolecular interactions.
PMID 24987116 · PMC4299494 · Molecular systems biology · 2014 · 8 claims · 8 setups
MinSpan, a mixed-integer linear optimization algorithm, computes the shortest, linearly independent pathways (sparsest basis of the null space of the stoichiometric matrix S) for genome-scale metabolic networks, which convex approaches (extreme pathways, elementary flux modes) cannot do at genome scale.
-
Full-text index only
BRCA1 and BRCA2 mutation predictions using the BOADICEA and BRCAPRO models and penetrance estimation in high-risk French-Canadian families.
PMID 16417652 · PMC1413985 · Breast cancer research : BCR · 2006 · 8 claims · 7 setups
BOADICEA predicts accurately the number of BRCA1 and BRCA2 mutations across family groups and discriminates well between carriers and noncarriers
-
Full-text index only
Consolidating the set of known human protein-protein interactions in preparation for large-scale mapping of the human interactome.
PMID 15892868 · PMC1175952 · Genome biology · 2005 · 8 claims · 6 setups
Two quantitative benchmarks (functional-annotation-based and physical-interaction-based log likelihood ratio scores) can measure relative accuracy of human PPI datasets
-
Full-text index only
Clustering of phosphorylation site recognition motifs can be exploited to predict the targets of cyclin-dependent kinase.
PMID 17316440 · PMC1852407 · Genome biology · 2007 · 8 claims · 6 setups
CDK consensus motifs are frequently clustered (closely spaced) in known CDK substrate proteins rather than uniformly distributed
-
Full-text index only
The UCSC Genome Browser Database: update 2006.
PMID 16381938 · PMC1347506 · Nucleic acids research · 2006 · 8 claims · 8 setups
The UCSC Genome Browser Database (GBD) provides integrated sequence and annotation data, with web tools (Genome Browser, Table Browser, Proteome Browser, Gene Sorter, BLAT, In Silico PCR) for visualizing and querying genomes of about a dozen vertebrate species and several model organisms.
-
Has reproduction · 79
Interpretable prediction models for widespread m6A RNA modification across cell lines and tissues.
PMID 37995291 · PMC10697738 · Bioinformatics (Oxford, England) · 2023 · 7 claims · 6 setups
CLSM6A, a CNN-based model set, predicts single-nucleotide-resolution m6A RNA modification sites across eight cell lines and three tissues in H. sapiens
-
Has reproduction · 62
Application of alternative de novo motif recognition models for analysis of structural heterogeneity of transcription factor binding sites: a case study of FOXA2 binding sites.
PMID 34547062 · PMC8408018 · Vavilovskii zhurnal genetiki i selektsii · 2021 · 6 claims · 7 setups
Combining four de novo models (PWM, diPWM, BaMM, InMoDe) significantly increases the fraction of recognized peaks versus PWM alone (by 26.3%).
-
Has reproduction
Genome-wide associations of aortic distensibility suggest causality for aortic aneurysms and brain white matter hyperintensities.
PMID 35922433 · PMC9349177 · Nature communications · 2022 · 8 claims · 7 setups
Genome-wide association of six CMR-derived aortic traits in up to 32,590 UK Biobank participants identifies 102 loci (including 27 novel associations) for aortic distensibility and area.
-
Full-text index only
The MAPPER database: a multi-genome catalog of putative transcription factor binding sites.
PMID 15608292 · PMC540057 · Nucleic acids research · 2005 · 8 claims · 6 setups
Built a library of 1134 HMM models (359 matrix-derived, 718 factor-derived, 57 JASPAR-derived), corresponding to 863 distinct TF names, from TRANSFAC and JASPAR binding site data
-
Full-text index only
Evolutionary modeling of rate shifts reveals specificity determinants in HIV-1 subtypes.
PMID 18989394 · PMC2566816 · PLoS computational biology · 2008 · 7 claims · 4 setups
A novel Bayesian method, RASER, can detect site-specific evolutionary rate shifts and the lineages in which they occurred without pre-specifying candidate lineages.
-
Full-text index only
Predicting candidate genes for human deafness disorders: a bioinformatics approach.
PMID 16854223 · PMC1564145 · BMC genomics · 2006 · 8 claims · 4 setups
A bioinformatic approach combining expression databases and protein interaction data narrows ~2400 candidate genes across deafness loci to a manageable set of candidates.
-
Full-text index only
Large-scale analysis of Macaca fascicularis transcripts and inference of genetic divergence between M. fascicularis and M. mulatta.
PMID 18294402 · PMC2287170 · BMC genomics · 2008 · 8 claims · 6 setups
Constructed full-length-enriched cDNA libraries and determined 85,721 EST sequences and 9407 full-insert sequences from cynomolgus macaque brain (7 regions), testis, and liver
-
Full-text index only
The UCSC Proteome Browser.
PMID 15608236 · PMC540054 · Nucleic acids research · 2005 · 8 claims · 5 setups
The UCSC Proteome Browser is tightly integrated with the UCSC Genome Browser, giving users simultaneous access to genome and proteome data.
-
Full-text index only
Toxicogenomics research consortium sails into uncharted waters.
PMID 12460811 · PMC1241122 · Environmental health perspectives · 2002 · 8 claims · 8 setups
The NIEHS-funded $37 million Toxicogenomics Research Consortium (TRC) combines the NIEHS Microarray Center with five academic institutions (UNC, Duke, Fred Hutchinson/UW, MIT, OHSU) to define genetic variability, set gene expression standards, and study environmental stress responses.
-
Has reproduction · 44
Detecting DNA modifications from SMRT sequencing data by modeling sequence context dependence of polymerase kinetic.
PMID 23516341 · PMC3597545 · PLoS computational biology · 2013 · 8 claims · 7 setups
Local sequence context strongly determines position-specific polymerase kinetic rate: roughly 80% of IPD variation is explained by a 10 bp context (7 bases upstream, 2 bases downstream of the incorporation site), saturating at 7 bases upstream.