Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Inference of transcriptional regulation using gene expression data from the bovine and human genomes.
PMID 17683551 · PMC1978505 · BMC genomics · 2007 · 7 claims · 8 setups
Using human reference promoter sequences is a useful approach for studying gene expression regulation in species with limited or non-existing genomic sequence, such as cattle.
-
Full-text index only
Genomic rearrangements by LINE-1 insertion-mediated deletion in the human and chimpanzee lineages.
PMID 16034026 · PMC1179734 · Nucleic acids research · 2005 · 8 claims · 6 setups
L1 insertions are directly responsible for genomic deletions (L1IMDs) confirmed in both human and chimpanzee genomes
-
Full-text index only
Natural variation of HIV-1 group M integrase: implications for a new class of antiretroviral inhibitors.
PMID 18687142 · PMC2546438 · Retrovirology · 2008 · 7 claims · 6 setups
Integrase displays significantly less inter- and intra-subtype amino acid diversity and lower Shannon's entropy than protease or RT.
-
Full-text index only
Distinctive pattern of sequence polymorphism in the NS3 protein of hepatitis C virus type 1b reflects conflicting evolutionary pressures.
PMID 18632963 · PMC2577380 · The Journal of general virology · 2008 · 7 claims · 6 setups
NS3 shows less evidence of purifying selection acting on its CTL epitopes than the other 9 HCV proteins, while outside the CTL epitopes NS3 is more conserved than the other proteins.
-
Full-text index only
The most frequent short sequences in non-coding DNA.
PMID 19966278 · PMC2831315 · Nucleic acids research · 2010 · 8 claims · 2 setups
Short frequent sequences (9-14 bases) in non-coding DNA may play a role in maintaining chromosome structure and function
-
Full-text index only
The Functional RNA Database 3.0: databases to support mining and annotation of functional RNAs.
PMID 18948287 · PMC2686472 · Nucleic acids research · 2009 · 8 claims · 5 setups
fRNAdb 3.0 is a completely rebuilt sequence database hosting a much larger collection of known/predicted non-coding RNA sequences with improved search functionality
-
Full-text index only
GenBank.
PMID 18940867 · PMC2686462 · Nucleic acids research · 2009 · 8 claims · 4 setups
GenBank is a comprehensive public database of nucleotide sequences with bibliographic and biological annotation, growing exponentially with a current doubling time of ~30 months.
-
Full-text index only
Direct evidence of extensive diversity of HIV-1 in Kinshasa by 1960.
PMID 18833279 · PMC3682493 · Nature · 2008 · 7 claims · 8 setups
Recovered and characterized HIV-1 sequences (DRC60) from a 1960 Bouin's-fixed paraffin-embedded lymph node biopsy from Léopoldville, Belgian Congo
-
Full-text index only
Designating eukaryotic orthology via processed transcription units.
PMID 18445630 · PMC2425467 · Nucleic acids research · 2008 · 8 claims · 5 setups
Existing ortholog databases discard/ignore alternative splicing via all-against-all protein comparisons, causing ambiguous ortholog calls and misclassification of AS isoforms as in-paralogs
-
Has reproduction · 88
Comprehensive benchmarking of large language models for RNA secondary structure prediction.
PMID 40205851 · PMC11982019 · Briefings in bioinformatics · 2025 · 7 claims · 4 setups
Existing RNA-LLMs had not previously been evaluated for secondary structure prediction in a unified, fair experimental setup with the same datasets and prediction model.
-
Full-text index only
Genomewide pattern of synonymous nucleotide substitution in two complete genomes of Mycobacterium tuberculosis.
PMID 12453367 · PMC2738538 · Emerging infectious diseases · 2002 · 8 claims · 6 setups
Genomewide comparison of two complete M. tuberculosis genomes reveals substantially more nucleotide diversity than prior studies based on few loci suggested
-
Full-text index only
Ab initio identification of putative human transcription factor binding sites by comparative genomics.
PMID 15865625 · PMC1097714 · BMC bioinformatics · 2005 · 8 claims · 5 setups
An integrated algorithm combining human-mouse genomic comparison, motif overrepresentation, and coregulation filters (GO annotation and microarray coexpression) can identify candidate transcription factor binding sites genome-wide
-
Full-text index only
Phylogenetic reconstruction of ancestral character states for gene expression and mRNA splicing data.
PMID 15921519 · PMC1166541 · BMC bioinformatics · 2005 · 6 claims · 4 setups
A minimum evolution algorithm (implemented in software 'phyrex') can reconstruct ancestral states of continuous characters like gene expression or splicing levels along a phylogeny
-
Full-text index only
Functional nsSNPs from carcinogenesis-related genes expressed in breast tissue: potential breast cancer risk alleles and their distribution across human populations.
PMID 16595073 · PMC3500178 · Human genomics · 2006 · 7 claims · 5 setups
A bioinformatics strategy cross-referencing carcinogenesis-related gene lists with breast-tissue expression data can identify candidate breast cancer risk nsSNPs.
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Full-text index only
High-throughput sequencing provides insights into genome variation and evolution in Salmonella Typhi.
PMID 18660809 · PMC2652037 · Nature genetics · 2008 · 7 claims · 8 setups
Evolution in the Typhi population is characterized by ongoing loss of gene function (pseudogene accumulation) rather than gain of function or diversifying selection.
-
Has reproduction · 68
Rfam 15: RNA families database in 2025.
PMID 39526405 · PMC11701678 · Nucleic acids research · 2025 · 8 claims · 6 setups
Rfamseq was expanded to 26 106 genomes, a 76% increase, by incorporating the latest UniProt reference proteomes and additional viral genomes
-
Has reproduction · 68
Mod(mdg4) variants repress telomeric retrotransposon HeT-A by blocking subtelomeric enhancers.
PMID 36373634 · PMC9723646 · Nucleic acids research · 2022 · 8 claims · 8 setups
Specific splice variants of Mod(mdg4) repress HeT-A by blocking subtelomeric enhancers in ovarian somatic cells (OSCs)
-
Full-text index only
Reconstruction of human protein interolog network using evolutionary conserved network.
PMID 17493278 · PMC1885812 · BMC bioinformatics · 2007 · 8 claims · 7 setups
A relative conservation score derived from maximal quasi-cliques in protein interaction networks, combined with other interaction features, can score and rank predicted human interologs for confidence.
-
Full-text index only
A highly polymorphic insertion in the Y-chromosome amelogenin gene can be used for evolutionary biology, population genetics and sexing in Cetacea and Artiodactyla.
PMID 18925953 · PMC2580767 · BMC genetics · 2008 · 8 claims · 6 setups
A 460–465 bp insertion is present in intron 4 of the Amel-Y locus in most Cetartiodactyla lineages (cetaceans and ruminants) but absent in pig