Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Genome annotation errors in pathway databases due to semantic ambiguity in partial EC numbers.
PMID 16034025 · PMC1179732 · Nucleic acids research · 2005 · 7 claims · 4 setups
Partial EC numbers are semantically ambiguous, and databases that assign a gene to all reactions sharing the same partial EC number make a faulty inference, causing systematic misannotation.
-
Full-text index only
EGASP: Introduction.
PMID 16925831 · PMC1810546 · Genome biology · 2006 · 8 claims · 5 setups
Computational gene finding methods, when compared to the GENCODE golden standard annotation, show that the human genome annotation is nearly complete in terms of novel protein-coding loci.
-
Has reproduction · 75
Roar: detecting alternative polyadenylation with standard mRNA sequencing libraries.
PMID 27756200 · PMC5069797 · BMC bioinformatics · 2016 · 8 claims · 5 setups
Roar, a method using PRE/POST read counts around annotated APA sites to compute an m/M ratio and a ratio-of-ratios (roar) statistic, detects differential 3'UTR shortening/lengthening from standard RNA-seq libraries.
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Full-text index only
PolySearch: a web-based text mining system for extracting relationships between human diseases, genes, mutations, drugs and metabolites.
PMID 18487273 · PMC2447794 · Nucleic acids research · 2008 · 8 claims · 7 setups
PolySearch supports more than 50 different classes of queries against nearly a dozen types of text, abstract, or bioinformatic databases
-
Full-text index only
Towards alignment independent quantitative assessment of homology detection.
PMID 17205117 · PMC1762415 · PloS one · 2006 · 8 claims · 6 setups
The Fhom Estimator uses the prevalence of a conserved protein feature (X) in two protein sets to estimate the fraction of true homologs among paired proteins, independent of alignment quality.
-
Full-text index only
Undergraduate research. Genomics Education Partnership.
PMID 18974335 · PMC2953277 · Science (New York, N.Y.) · 2008 · 6 claims · 4 setups
A course-embedded, multi-institution undergraduate research model (the Genomics Education Partnership) can deliver authentic research experiences during the academic year rather than only in summer programs.
-
Has reproduction · 50
MEDUSA: A Pipeline for Sensitive Taxonomic Classification and Flexible Functional Annotation of Metagenomic Shotgun Sequences.
PMID 35330728 · PMC8940201 · Frontiers in genetics · 2022 · 6 claims · 6 setups
MEDUSA is an automated, Conda-installable and Snakemake-managed pipeline performing preprocessing, assembly, alignment, taxonomic classification, and functional annotation on shotgun data.
-
Full-text index only
Biotin tagging coupled with amino acid-coded mass tagging for efficient and precise screening of interaction proteome in mammalian cells.
PMID 19834888 · PMC4302342 · Proteomics · 2009 · 7 claims · 7 setups
BioCAT (biotin tagging + AACT) enables highly sensitive and accurate single-step screening of mammalian protein-protein interactions without establishing a stable cell line
-
Has reproduction · 60
TRAPID 2.0: a web application for taxonomic and functional analysis of de novo transcriptomes.
PMID 34197621 · PMC8464036 · Nucleic acids research · 2021 · 8 claims · 8 setups
TRAPID 2.0 is a web application performing global characterization of de novo transcriptomes via structural, functional, and taxonomic annotation in an initial processing phase, followed by an exploratory phase of downstream analyses.
-
Full-text index only
Discovery of protein-protein interactions using a combination of linguistic, statistical and graphical information.
PMID 15941473 · PMC1164402 · BMC bioinformatics · 2005 · 8 claims · 5 setups
A combined linguistic+statistical+rule-based method achieves precision 0.61 and recall 0.97 (f=0.74) detecting yeast protein-protein interactions across 12,300 Medline abstracts.
-
Full-text index only
Filtering high-throughput protein-protein interaction data using a combination of genomic features.
PMID 15833142 · PMC1127019 · BMC bioinformatics · 2005 · 8 claims · 8 setups
A combination of three genomic features (interacting Pfam domains, GO annotations, sequence homology) using naive Bayesian networks predicts true protein-protein interactions with high sensitivity and good specificity.
-
Full-text index only
Ensembl 2007.
PMID 17148474 · PMC1761443 · Nucleic acids research · 2007 · 8 claims · 7 setups
Ensembl added 18 new chordate genomes this year, increasing total genomes available from 15 to 33, the largest yearly increase to date.
-
Has reproduction · 67
binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets.
PMID 36239393 · PMC9677464 · Briefings in bioinformatics · 2022 · 8 claims · 8 setups
binny outperforms or is highly competitive with commonly used and state-of-the-art binning methods (MetaBAT2, MaxBin2, CONCOCT, VAMB, SemiBin, MetaDecoder)
-
Full-text index only
Metagenomic analysis of respiratory tract DNA viral communities in cystic fibrosis and non-cystic fibrosis individuals.
PMID 19816605 · PMC2756586 · PloS one · 2009 · 8 claims · 8 setups
CF phage communities are highly similar to each other, whereas Non-CF individuals have more distinct, variable phage communities reflecting transient environmental sampling
-
Full-text index only
Satellog: a database for the identification and prioritization of satellite repeats in disease association studies.
PMID 15949044 · PMC1181805 · BMC bioinformatics · 2005 · 7 claims · 6 setups
Satellog is a database cataloging all pure 1-16 unit satellite repeats in the human genome with supplementary polymorphism, gene-location, and expression data for prioritizing repeats in disease-association studies.
-
Has reproduction · 58
Genomic Correlates of Virulence Attenuation in the Deadly Amphibian Chytrid Fungus, Batrachochytrium dendrobatidis.
PMID 26333840 · PMC4632049 · G3 (Bethesda, Md.) · 2015 · 8 claims · 8 setups
Virulence attenuation in the longer-passaged Bd isolate (JEL427-P39) is associated with loss of chromosome copy number relative to the shorter-passaged, more virulent isolate (JEL427-P9)
-
Has reproduction · 59
Nucleosome regulatory dynamics in response to TGFβ.
PMID 24771338 · PMC4066760 · Nucleic acids research · 2014 · 8 claims · 7 setups
SuMMIt, a Bayesian strand-based mixture model requiring support from both ends of sequenced fragments, enables precise nucleosome mid-position calling, fuzziness scoring and between-condition change detection.
-
Full-text index only
Brain-specific proteins decline in the cerebrospinal fluid of humans with Huntington disease.
PMID 18984577 · PMC2649809 · Molecular & cellular proteomics : MCP · 2009 · 8 claims · 6 setups
Brain-specific proteins are 1.8 times more likely to be observed in CSF than in plasma
-
Has reproduction · 66
RNAseq analysis of the parasitic nematode Strongyloides stercoralis reveals divergent regulation of canonical dauer pathways.
PMID 23145190 · PMC3493385 · PLoS neglected tropical diseases · 2012 · 8 claims · 8 setups
S. stercoralis possesses homologs of nearly all C. elegans dauer genes, but with significant differences in protein structure, developmental regulation, and gene family expansion.