Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Gene prediction in eukaryotes with a generalized hidden Markov model that uses hints from external sources.
PMID 16469098 · PMC1409804 · BMC bioinformatics · 2006 · 7 claims · 3 setups
AUGUSTUS+ extends the AUGUSTUS GHMM by combining intrinsic sequence information with extrinsic hints via an extended emission alphabet, so the GHMM jointly models the DNA sequence, gene structure, and hint collection.
-
Full-text index only
SNiPer: improved SNP genotype calling for Affymetrix 10K GeneChip microarray data.
PMID 16262895 · PMC1280925 · BMC genomics · 2005 · 8 claims · 5 setups
Poorly performing SNPs (NoCall rate ≥25%) fail primarily due to inadequate training/localization of the MPAM statistical model call zone, not detection filter failure
-
Has reproduction · 83
Accurate prediction of metagenome-assembled genome completeness by MAGISTA, a random forest model built on alignment-free intra-bin statistics.
PMID 35248155 · PMC8898458 · Environmental microbiome · 2022 · 7 claims · 7 setups
MAGISTA, a random forest model built on alignment-free intra-bin distance-distribution statistics, can estimate MAG completeness and purity without relying on reference marker genes.
-
Full-text index only
Genomics and public health: development of Web-based training tools for increasing genomic awareness.
PMID 15888236 · PMC1327719 · Preventing chronic disease · 2005 · 7 claims · 4 setups
Web-based training tools (Genomics for Public Health Practitioners and Six Weeks to Genomic Awareness) can increase genomic awareness among public health practitioners nationwide.
-
Full-text index only
Identification of disease causing loci using an array-based genotyping approach on pooled DNA.
PMID 16197552 · PMC1262713 · BMC genomics · 2005 · 8 claims · 5 setups
Pooling genomic DNA and genotyping on SNP microarrays accurately predicts allelic frequencies relative to individual genotyping
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.
-
Full-text index only
A genomic approach to improve prognosis and predict therapeutic response in chronic lymphocytic leukemia.
PMID 19861443 · PMC2783430 · Clinical cancer research : an official journal of the American Association for Cancer Research · 2009 · 8 claims · 8 setups
A genomic signature derived from CLL patient leukemic cells significantly differentiates stable from progressive disease
-
Has reproduction
Genomic prediction based on selective linkage disequilibrium pruning of low-coverage whole-genome sequence variants in a pure Duroc population.
PMID 37853325 · PMC10583454 · Genetics, selection, evolution : GSE · 2023 · 8 claims · 6 setups
Selective linkage disequilibrium pruning (SLDP) refines whole-genome SNP sets using GWAS prior information to improve genomic prediction accuracy.
-
Full-text index only
QuantiSNP: an Objective Bayes Hidden-Markov Model to detect and accurately map copy number variation using SNP genotyping data.
PMID 17341461 · PMC1874617 · Nucleic acids research · 2007 · 8 claims · 7 setups
QuantiSNP (OB-HMM) provides probabilistic quantification of copy number states and significantly improves accuracy of segmental aneuploidy identification and breakpoint mapping relative to existing tools (BeadStudio/Illumina)
-
Full-text index only
Predicting failure rate of PCR in large genomes.
PMID 18492719 · PMC2441781 · Nucleic acids research · 2008 · 7 claims · 8 setups
The number of predicted primer-binding sites in genomic DNA is the most important factor determining PCR failure.
-
Full-text index only
Review of state Comprehensive Cancer Control plans for genomics content.
PMID 15888219 · PMC1327702 · Preventing chronic disease · 2005 · 8 claims · 2 setups
18 of 30 state CCC plans analyzed contained genomics-related components, with wide variability in content
-
Full-text index only
Inference of transcriptional regulation using gene expression data from the bovine and human genomes.
PMID 17683551 · PMC1978505 · BMC genomics · 2007 · 7 claims · 8 setups
Using human reference promoter sequences is a useful approach for studying gene expression regulation in species with limited or non-existing genomic sequence, such as cattle.
-
Has reproduction · 79
Genome-wide prediction of DNase I hypersensitivity using gene expression.
PMID 29051481 · PMC5715040 · Nature communications · 2017 · 6 claims · 3 setups
Gene expression substantially predicts genome-wide DNase I hypersensitivity (DH), demonstrating transcriptome-based prediction as a feasible approach for regulome mapping
-
Full-text index only
AUGUSTUS at EGASP: using EST, protein and genomic alignments for improved gene prediction in the human genome.
PMID 16925833 · PMC1810548 · Genome biology · 2006 · 8 claims · 5 setups
AUGUSTUS predicted significantly more genes correctly than any other ab initio program in EGASP
-
Has reproduction · 63
Comparative Genomics of Borderline Oxacillin-Resistant Staphylococcus aureus Detected during a Pseudo-outbreak of Methicillin-Resistant S. aureus in a Neonatal Intensive Care Unit.
PMID 35038924 · PMC8764539 · mBio · 2022 · 7 claims · 8 setups
Of 42 isolates flagged as MRSA by screening agar, only 9 were PBP2a- and mecA-positive true MRSA, while the remaining 33 were mecA-negative and largely met criteria for BORSA
-
Full-text index only
Comparison of prognostic gene expression signatures for breast cancer.
PMID 18717985 · PMC2533026 · BMC genomics · 2008 · 8 claims · 5 setups
The three prognostic signatures (70-gene, 76-gene, GGI) show similar prognostic performance for predicting DMFS despite differing gene identity and development approach
-
Has reproduction · 83
Gene-expression patterns in peripheral blood classify familial breast cancer susceptibility.
PMID 26538066 · PMC4634735 · BMC medical genomics · 2015 · 8 claims · 5 setups
A multigene peripheral-blood gene-expression biomarker accurately classifies which women from high-risk families develop familial breast cancer.
-
Full-text index only
Using ESTs to improve the accuracy of de novo gene prediction.
PMID 16817966 · PMC1534067 · BMC bioinformatics · 2006 · 8 claims · 8 setups
TWINSCAN_EST combines EST alignments with TWINSCAN via a trainable 'ESTseq' representation and improves exact gene structure prediction accuracy on the whole C. elegans genome
-
Full-text index only
A re-annotation pipeline for Illumina BeadArrays: improving the interpretation of gene expression data.
PMID 19923232 · PMC2817484 · Nucleic acids research · 2010 · 8 claims · 7 setups
A Perl-based pipeline that BLASTs/BLATs Illumina probe sequences against genomes and transcript databases (RefSeq, UCSC Known Genes, UniGene/GenBank, Ensembl) can classify probes by quality grade (Perfect/Good/Bad/No match) and is applicable across 8 BeadArray platforms and other array types
-
Full-text index only
Classification of real and pseudo microRNA precursors using local structure-sequence features and support vector machine.
PMID 16381612 · PMC1360673 · BMC bioinformatics · 2005 · 7 claims · 7 setups
A 32-dimensional triplet structure-sequence feature vector combined with SVM (triplet-SVM) can distinguish real human pre-miRNAs from pseudo pre-miRNA hairpins with ~90% accuracy.