Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
A macaque's-eye view of human insertions and deletions: differences in mechanisms.
PMID 17941704 · PMC1976337 · PLoS computational biology · 2007 · 7 claims · 4 setups
Insertion and deletion rates are differentially associated with replication- versus recombination-related genomic features, indicating the two mutation types are driven in part by distinct mechanisms
-
Full-text index only
Comparative genomic analysis of Campylobacter jejuni associated with Guillain-Barré and Miller Fisher syndromes: neuropathogenic and enteritis-associated isolates can share high levels of genomic similarity.
PMID 17919333 · PMC2174954 · BMC genomics · 2007 · 8 claims · 4 setups
GBS/MFS strains are genomically heterogeneous, falling into about six major lineages rather than a single clonal group
-
Has reproduction · 89
Comparative genomics of dairy-associated Staphylococcus aureus from selected sub-Saharan African regions reveals milk as reservoir for human-and animal-derived strains and identifies a putative animal-related clade with presumptive novel siderophore.
PMID 36046020 · PMC9421002 · Frontiers in microbiology · 2022 · 7 claims · 8 setups
Milk serves as a reservoir for both human- and animal-derived S. aureus strains in sub-Saharan Africa
-
Full-text index only
POCUS: mining genomic sequence annotation to predict disease genes.
PMID 14611661 · PMC329128 · Genome biology · 2003 · 8 claims · 6 setups
Genes predisposing to the same disease tend to share functional annotation IDs (GO/InterPro) more than expected by chance
-
Full-text index only
AUGUSTUS at EGASP: using EST, protein and genomic alignments for improved gene prediction in the human genome.
PMID 16925833 · PMC1810548 · Genome biology · 2006 · 8 claims · 5 setups
AUGUSTUS predicted significantly more genes correctly than any other ab initio program in EGASP
-
Full-text index only
Computer-aided identification of polymorphism sets diagnostic for groups of bacterial and viral genetic variants.
PMID 17672919 · PMC1973086 · BMC bioinformatics · 2007 · 6 claims · 8 setups
The Not-N algorithm, incorporated into the Minimum SNPs program, identifies small marker sets diagnostic for user-defined subgroups of genetic variants with 0% false negatives
-
Full-text index only
Global variation in copy number in the human genome.
PMID 17122850 · PMC2669898 · Nature · 2006 · 8 claims · 6 setups
A first-generation CNV map of the human genome was constructed from 270 HapMap individuals across four populations, identifying 1,447 CNV regions covering ~360 Mb (12%) of the genome.
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Full-text index only
SNPdetector: a software tool for sensitive and accurate SNP detection.
PMID 16261194 · PMC1274293 · PLoS computational biology · 2005 · 7 claims · 7 setups
SNPdetector, which models human visual inspection of sequencing traces, achieves low false positive and false negative rates in automated SNP and mutation detection
-
Full-text index only
Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.
PMID 16790052 · PMC1534064 · BMC bioinformatics · 2006 · 8 claims · 7 setups
The Sequential Minimal Optimization (SMO) SVM algorithm was the best-performing classifier among 26 WEKA classifiers for predicting catalytic residues
-
Full-text index only
Environmental Burkholderia cepacia complex isolates in human infections.
PMID 17552100 · PMC2725883 · Emerging infectious diseases · 2007 · 6 claims · 4 setups
More than 20% of clinical Bcc isolates examined are indistinguishable by MLST from environmental isolates, linking the natural environment to emergence of clinical infections
-
Full-text index only
Eurasian and African mitochondrial DNA influences in the Saudi Arabian population.
PMID 17331239 · PMC1810519 · BMC evolutionary biology · 2007 · 8 claims · 4 setups
The majority (85%) of Saudi Arab mtDNA lineages have a western Asia (Eurasian) provenance
-
Full-text index only
Relative impact of nucleotide and copy number variation on gene expression phenotypes.
PMID 17289997 · PMC2665772 · Science (New York, N.Y.) · 2007 · 8 claims · 5 setups
SNPs and CNVs capture largely non-overlapping signals of genetic variation affecting gene expression
-
Has reproduction · 59
Refining breast cancer biomarker discovery and drug targeting through an advanced data-driven approach.
PMID 38253993 · PMC10810249 · BMC bioinformatics · 2024 · 8 claims · 8 setups
The BGWO_SA_Ens algorithm (hybrid BGWO + simulated annealing with an ensemble classifier objective function) selects predictive breast cancer biomarker genes with high classification performance
-
Has reproduction · 71
Parsimonious Gene Correlation Network Analysis (PGCNA): a tool to define modular gene co-expression for refined molecular stratification in cancer.
PMID 30993001 · PMC6459838 · NPJ systems biology and applications · 2019 · 8 claims · 7 setups
Retaining only the top ~3 most correlated edges per gene (EPG3) combined with FastUnfold clustering (termed PGCNA) produces gene co-expression modules with significantly better separation and enrichment of known biology than using all edges or other clustering methods.
-
Full-text index only
Oncogenic mutations in GNAQ occur early in uveal melanoma.
PMID 18719078 · PMC2634606 · Investigative ophthalmology & visual science · 2008 · 8 claims · 7 setups
Activating GNAQ mutations at codon 209 occur in 33/67 (49%) of primary uveal melanomas, making it the most common known oncogenic mutation in UM
-
Full-text index only
Genetic variation in an individual human exome.
PMID 18704161 · PMC2493042 · PLoS genetics · 2008 · 8 claims · 7 setups
The ~12,500 nonsilent coding variants in the HuRef exome can be reduced ~8-fold to a set of ~1,600 variants most likely to affect protein function.
-
Full-text index only
Extraction of human kinase mutations from literature, databases and genotyping studies.
PMID 19758464 · PMC2745582 · BMC bioinformatics · 2009 · 7 claims · 6 setups
A literature mining pipeline combining MutationFinder, false-positive filtering, and SVM-based classification can extract and disambiguate single-point mutation mentions from abstracts and full text
-
Full-text index only
The UCSC Genome Browser Database: 2008 update.
PMID 18086701 · PMC2238835 · Nucleic acids research · 2008 · 8 claims · 8 setups
The UCSC Genome Browser Database (GBD) provides integrated sequence and annotation data for a large collection of vertebrate and model organism genomes.
-
Has reproduction · 96
Scalable Prediction of Acute Myeloid Leukemia Using High-Dimensional Machine Learning and Blood Transcriptomics.
PMID 31918046 · PMC6992905 · iScience · 2020 · 8 claims · 8 setups
Data-driven, high-dimensional ML approaches that learn multivariate signatures directly from genome-wide transcriptomic data (no prior gene selection) yield accurate and robust AML classifiers.