Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
VIRGO: computational prediction of gene functions.
PMID 16845022 · PMC1538839 · Nucleic acids research · 2006 · 8 claims · 6 setups
VIRGO constructs a functional linkage network (FLN) from gene expression and molecular interaction data, labels genes with GO annotations, and propagates these labels to predict functions of unlabelled genes
-
Full-text index only
SNAP predicts effect of mutations on protein function.
PMID 18757876 · PMC2562009 · Bioinformatics (Oxford, England) · 2008 · 8 claims · 3 setups
SNAP is a publicly available web-server implementation predicting functional effects (neutral/non-neutral) of single amino acid substitutions.
-
Full-text index only
Towards alignment independent quantitative assessment of homology detection.
PMID 17205117 · PMC1762415 · PloS one · 2006 · 8 claims · 6 setups
The Fhom Estimator uses the prevalence of a conserved protein feature (X) in two protein sets to estimate the fraction of true homologs among paired proteins, independent of alignment quality.
-
Full-text index only
The Princeton Protein Orthology Database (P-POD): a comparative genomics analysis tool for biologists.
PMID 17712414 · PMC1942082 · PloS one · 2007 · 8 claims · 5 setups
P-POD is the first comparative genomics database to combine results from multiple computational ortholog/homolog prediction methods with manually curated literature-derived experimental evidence of functional conservation.
-
Full-text index only
Gene prediction in eukaryotes with a generalized hidden Markov model that uses hints from external sources.
PMID 16469098 · PMC1409804 · BMC bioinformatics · 2006 · 7 claims · 3 setups
AUGUSTUS+ extends the AUGUSTUS GHMM by combining intrinsic sequence information with extrinsic hints via an extended emission alphabet, so the GHMM jointly models the DNA sequence, gene structure, and hint collection.
-
Full-text index only
In silico analysis of missense substitutions using sequence-alignment based methods.
PMID 18951440 · PMC3431198 · Human mutation · 2008 · 8 claims · 7 setups
Carefully validated PMSA-based computational algorithms can achieve predictive values of ~75-95% for classifying missense substitutions as pathogenic or neutral.
-
Full-text index only
Prediction by graph theoretic measures of structural effects in proteins arising from non-synonymous single nucleotide polymorphisms.
PMID 18654622 · PMC2447880 · PLoS computational biology · 2008 · 8 claims · 5 setups
Bongo identifies mutations causing local and global structural effects with a remarkably low false positive rate
-
Full-text index only
Using ESTs to improve the accuracy of de novo gene prediction.
PMID 16817966 · PMC1534067 · BMC bioinformatics · 2006 · 8 claims · 8 setups
TWINSCAN_EST combines EST alignments with TWINSCAN via a trainable 'ESTseq' representation and improves exact gene structure prediction accuracy on the whole C. elegans genome
-
Has reproduction · 88
Comprehensive benchmarking of large language models for RNA secondary structure prediction.
PMID 40205851 · PMC11982019 · Briefings in bioinformatics · 2025 · 7 claims · 4 setups
Existing RNA-LLMs had not previously been evaluated for secondary structure prediction in a unified, fair experimental setup with the same datasets and prediction model.
-
Full-text index only
Flanking p10 contribution and sequence bias in matrix based epitope prediction: revisiting the assumption of independent binding pockets.
PMID 18925947 · PMC2600787 · BMC structural biology · 2008 · 8 claims · 3 setups
The extended matrix PP10 (built from a proline-containing peptide library) shows significant improvement in binding prediction over the original nine-residue matrix P9
-
Has reproduction · 50
Comparative analysis of circular RNAs between soybean cytoplasmic male-sterile line NJCMS1A and its maintainer NJCMS1B by high-throughput sequencing.
PMID 30208848 · PMC6134632 · BMC genomics · 2018 · 8 claims · 7 setups
2867 circRNAs were identified in soybean flower buds via high-throughput sequencing with RNase R enrichment, of which 1009 were differentially expressed between NJCMS1A and NJCMS1B
-
Full-text index only
Comparative metagenomics revealed commonly enriched gene sets in human gut microbiomes.
PMID 17916580 · PMC2533590 · DNA research : an international journal for rapid publication of reports on genes and genomes · 2007 · 7 claims · 7 setups
Adult and weaned-children gut microbiota show high functional (gene-content) uniformity despite taxonomic differences, while unweaned infant microbiota show high inter-individual variation in both taxonomic and gene composition.
-
Has reproduction · 68
Improved precision of epigenetic clock estimates across tissues and its implication for biological ageing.
PMID 31443728 · PMC6708158 · Genome medicine · 2019 · 8 claims · 6 setups
The proportion of variance in chronological age explained by all DNA methylation probes is close to 1, so a near-perfect age predictor is in principle achievable with sufficient training data.
-
Full-text index only
GeneTide--Terra Incognita Discovery Endeavor: a new transcriptome focused member of the GeneCards/GeneNote suite of databases.
PMID 15608261 · PMC540076 · Nucleic acids research · 2005 · 8 claims · 7 setups
GeneTide integrates UniGene, DoTS, AceView, BLAT/GeneLoc genomic alignment, and GeneAnnot probe-set data into a unified Consensus/Uniqueness/Score scheme to associate ESTs with GeneCards genes
-
Full-text index only
ABS: a database of Annotated regulatory Binding Sites from orthologous promoters.
PMID 16381947 · PMC1347478 · Nucleic acids research · 2006 · 7 claims · 6 setups
ABS is a public database of experimentally identified TF binding sites conserved in orthologous vertebrate gene promoters, manually curated from the literature.
-
Full-text index only
The other side of comparative genomics: genes with no orthologs between the cow and other mammalian species.
PMID 20003425 · PMC2808326 · BMC genomics · 2009 · 7 claims · 4 setups
3,801 bovine genes have no orthologs in human, mouse and dog, and 1,010 human genes have no orthologs in cow despite having orthologs in mouse and dog
-
Full-text index only
Genetic variation at hair length candidate genes in elephants and the extinct woolly mammoth.
PMID 19747392 · PMC2754481 · BMC evolutionary biology · 2009 · 8 claims · 5 setups
The coding sequence of FGF5 is not the critical determinant of hair length differences among elephantids, including the woolly mammoth.
-
Has reproduction · 88
Human methylome variation across Infinium 450K data on the Gene Expression Omnibus.
PMID 33937763 · PMC8061458 · NAR genomics and bioinformatics · 2021 · 8 claims · 6 setups
Approximately two-thirds of compiled HM450K samples are from blood, one-quarter from brain, and roughly one-third from cancer patients.
-
Has reproduction · 71
Microbial diversity of plant pathogens and insect endosymbionts in Reptalus artemisiae.
PMID 41826827 · PMC13202766 · BMC microbiology · 2026 · 8 claims · 8 setups
R. artemisiae harbors six prokaryotic taxa: two plant pathogens ('Ca. P. solani', 'Ca. A. phytopathogenicus') and four insect endosymbionts ('Ca. Vidania', 'Ca. Purcelliella', 'Ca. Karelsulcia', and Wolbachia).
-
Full-text index only
Cruciform extrusion propensity of human translocation-mediating palindromic AT-rich repeats.
PMID 17264116 · PMC1851657 · Nucleic acids research · 2007 · 8 claims · 4 setups
Cruciform extrusion propensity of PATRRs depends on both length and central symmetry of the repeat.