Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
nf-core/viralmetagenome: A novel pipeline for untargeted viral genome reconstruction.
PMID 42057295 · PMC13141149 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 5 setups
nf-core/viralmetagenome is a Nextflow pipeline that automates untargeted reconstruction and variant analysis of eukaryotic DNA and RNA viruses from short-read metagenomic or hybridisation-capture data.
-
Full-text index only
Complex genetic diseases: controversy over the Croesus code.
PMID 11532206 · PMC138948 · Genome biology · 2001 · 8 claims · 3 setups
The common disease/common variant hypothesis is predicted by population genetic theory (founder population dynamics, mutation-drift-selection balance) and supported by empirical examples such as APOE*E4.
-
Full-text index only
The application of genomic technology to combat parasitic disease.
PMID 16454893 · PMC1914215 · Parasitology · 2004 · 8 claims · 8 setups
Genomic resources from mammalian hosts can be exploited to understand parasitic processes from the host's standpoint, illustrated by host cell penetration by Trypanosoma cruzi.
-
Has reproduction · 88
Comprehensive benchmarking of large language models for RNA secondary structure prediction.
PMID 40205851 · PMC11982019 · Briefings in bioinformatics · 2025 · 7 claims · 4 setups
Existing RNA-LLMs had not previously been evaluated for secondary structure prediction in a unified, fair experimental setup with the same datasets and prediction model.
-
Full-text index only
Imputation-based analysis of association studies: candidate regions and quantitative traits.
PMID 17676998 · PMC1934390 · PLoS genetics · 2007 · 8 claims · 2 setups
Imputation-based Bayesian regression increases power to detect association compared with standard single-SNP tests, even when the causal variant is directly typed
-
Has reproduction · 39
Glacier shrinkage will accelerate downstream decomposition of organic matter and alters microbiome structure and function.
PMID 35320603 · PMC9323552 · Global change biology · 2022 · 8 claims · 7 setups
Glacier shrinkage (decreasing glacier influence) accelerates downstream organic matter decomposition rates in glacier-fed streams
-
Has reproduction · 85
Prediction of condition-specific regulatory genes using machine learning.
PMID 32329779 · PMC7293043 · Nucleic acids research · 2020 · 8 claims · 6 setups
ConSReg integrates expression, DAP-seq TF-DNA binding, and ATAC-seq open chromatin data into machine learning models to predict condition-specific regulatory genes
-
Full-text index only
Iterative class discovery and feature selection using Minimal Spanning Trees.
PMID 15355552 · PMC520744 · BMC bioinformatics · 2004 · 7 claims · 5 setups
Iterating between MST-based clustering and t-statistic feature selection removes noise genes step-wise while sharpening the sample clustering
-
Full-text index only
Classification of real and pseudo microRNA precursors using local structure-sequence features and support vector machine.
PMID 16381612 · PMC1360673 · BMC bioinformatics · 2005 · 7 claims · 7 setups
A 32-dimensional triplet structure-sequence feature vector combined with SVM (triplet-SVM) can distinguish real human pre-miRNAs from pseudo pre-miRNA hairpins with ~90% accuracy.
-
Full-text index only
Identifying synonymous regulatory elements in vertebrate genomes.
PMID 15980499 · PMC1160227 · Nucleic acids research · 2005 · 7 claims · 4 setups
SynoR is a tool that performs de novo genome-wide identification of synonymous regulatory elements (SREs) using evolutionarily conserved TFBS modules as seeds
-
Full-text index only
MILANO--custom annotation of microarray results using automatic literature searches.
PMID 15661078 · PMC547913 · BMC bioinformatics · 2005 · 7 claims · 4 setups
MILANO annotates microarray gene lists by counting literature co-occurrences of each gene with user-defined secondary terms
-
Full-text index only
Computing Ka and Ks with a consideration of unequal transitional substitutions.
PMID 16740169 · PMC1552089 · BMC evolutionary biology · 2006 · 7 claims · 7 setups
MYN, a modified version of the Yang-Nielsen (YN) algorithm based on the Tamura-Nei Model, allows unequal transitional substitution rates between purines (κR) and pyrimidines (κY) plus codon frequency bias
-
Full-text index only
Identification and evolutionary analysis of novel exons and alternative splicing events using cross-species EST-to-genome comparisons in human, mouse and rat.
PMID 16536879 · PMC1479377 · BMC bioinformatics · 2006 · 8 claims · 6 setups
ENACE, a cross-species EST-to-genome comparison algorithm, can identify novel cassette-on exons and retained introns for EST-scanty species and distinguish conserved vs lineage-specific exons
-
Full-text index only
Genome-wide identification of human functional DNA using a neutral indel model.
PMID 16410828 · PMC1326222 · PLoS computational biology · 2006 · 8 claims · 8 setups
A neutral indel model predicting a geometric distribution of intergap segment (IGS) lengths fits human-mouse ancestral repeat (AR) alignment data excellently
-
Full-text index only
In silico analysis of missense substitutions using sequence-alignment based methods.
PMID 18951440 · PMC3431198 · Human mutation · 2008 · 8 claims · 7 setups
Carefully validated PMSA-based computational algorithms can achieve predictive values of ~75-95% for classifying missense substitutions as pathogenic or neutral.
-
Full-text index only
Estimation of relevant variables on high-dimensional biological patterns using iterated weighted kernel functions.
PMID 18509521 · PMC2396875 · PloS one · 2008 · 7 claims · 6 setups
wKIERA combines a weighted-kernel discriminant (kernel perceptron) with an iterative stochastic probability estimation-of-distribution algorithm to estimate a relevance distribution over variables
-
Full-text index only
Identification of somatically acquired rearrangements in cancer using genome-wide massively parallel paired-end sequencing.
PMID 18438408 · PMC2705838 · Nature genetics · 2008 · 8 claims · 8 setups
Massively parallel paired-end sequencing can characterize somatic and germline structural rearrangements to base-pair resolution across a whole cancer genome
-
Full-text index only
Protein function assignment through mining cross-species protein-protein interactions.
PMID 18253506 · PMC2216687 · PloS one · 2008 · 8 claims · 6 setups
CSIDOP predicts protein molecular function with 95.42% accuracy using 2,972 GO functional categories in H. sapiens
-
Full-text index only
Microbial genomics: from sequence to function.
PMID 10998380 · PMC2627950 · Emerging infectious diseases · 2000 · 8 claims · 4 setups
Whole-genome shotgun sequencing (sequencing and assembly of random genome fragments), first demonstrated with Haemophilus influenzae in 1995, is now the method of choice for sequencing most genomes, including the human genome.
-
Has reproduction · 77
SurvConvMixer: robust and interpretable cancer survival prediction based on ConvMixer using pathway-level gene expression images.
PMID 38539106 · PMC10967213 · BMC bioinformatics · 2024 · 6 claims · 5 setups
SurvConvMixer, using pathway-level gene expression images and ConvMixer, achieves strong internal validation AUC for overall survival prediction, especially on larger datasets like LUAD