Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Diversity of the parB and repA genes of the Burkholderia cepacia complex and their utility for rapid identification of Burkholderia cenocepacia.
PMID 18328098 · PMC2324101 · BMC microbiology · 2008 · 8 claims · 7 setups
repA sequences show distinct clustering of B. cenocepacia relative to other Bcc species, enabling design of a species-specific multiplex PCR
-
Has reproduction · 50
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization.
PMID 41266599 · PMC12635123 · Communications biology · 2025 · 8 claims · 8 setups
BiRNA-BERT uses adaptive dual-tokenization that dynamically selects nucleotide-level (NUC) or byte-pair encoding (BPE) tokens based on input sequence length
-
Full-text index only
ECgene: genome annotation for alternative splicing.
PMID 15608289 · PMC540072 · Nucleic acids research · 2005 · 8 claims · 5 setups
ECgene combines genome-based EST clustering with a graph-theoretic transcript assembly procedure to predict gene models including alternative splicing events.
-
Full-text index only
Polymorphix: a sequence polymorphism database.
PMID 15608242 · PMC540030 · Nucleic acids research · 2005 · 8 claims · 5 setups
Polymorphix is an ACNUC-structured database that organizes EMBL/GenBank sequences into within-species homologous sequence families using similarity and bibliographic criteria, with alignments, outgroups and phylogenetic trees provided.
-
Full-text index only
Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space.
PMID 16481312 · PMC1373602 · Nucleic acids research · 2006 · 8 claims · 7 setups
The number of protein families continues to expand steadily as more genomes are sequenced, showing no sign of saturation.
-
Full-text index only
Identification and characterization of insect-specific proteins by genome data analysis.
PMID 17407609 · PMC1852559 · BMC genomics · 2007 · 8 claims · 7 setups
Comparative genome analysis across five holometabolous insects and three non-insect eukaryotes (opisthokonts) identifies 154 insect-specific orthologous groups (refined to 51 proteins) and 466 eukaryote/opisthokont-core orthologous groups
-
Full-text index only
In silico segmentations of lentivirus envelope sequences.
PMID 17376229 · PMC1847453 · BMC bioinformatics · 2007 · 8 claims · 8 setups
C and V regions of lentivirus SU sequences have distinct statistical (oligonucleotide/amino-acid) compositions that HMMs can learn and use to delimit them.
-
Full-text index only
RAId_DbS: mass-spectrometry based peptide identification web server with knowledge integration.
PMID 18954448 · PMC2605478 · BMC genomics · 2008 · 7 claims · 4 setups
Constructed enhanced protein databases integrating annotated SAPs, PTMs, and disease associations for 17 organisms.
-
Full-text index only
The TIGR Gene Indices: clustering and assembling EST and known genes and integration with eukaryotic genomes.
PMID 15608288 · PMC540018 · Nucleic acids research · 2005 · 8 claims · 8 setups
The TIGR Gene Indices (TGI) are a collection of 77 species-specific databases that cluster and assemble EST and known gene sequences into tentative consensus (TC) sequences to identify and characterize expressed transcripts.
-
Full-text index only
The Princeton Protein Orthology Database (P-POD): a comparative genomics analysis tool for biologists.
PMID 17712414 · PMC1942082 · PloS one · 2007 · 8 claims · 5 setups
P-POD is the first comparative genomics database to combine results from multiple computational ortholog/homolog prediction methods with manually curated literature-derived experimental evidence of functional conservation.
-
Full-text index only
Coiled-coil protein composition of 22 proteomes--differences and common themes in subcellular infrastructure and traffic control.
PMID 16288662 · PMC1322226 · BMC evolutionary biology · 2005 · 7 claims · 5 setups
Proteins with extended coiled-coil domains (>250 amino acids) are largely absent from bacterial genomes but present in archaea and eukaryotes.
-
Full-text index only
EPGD: a comprehensive web resource for integrating and displaying eukaryotic paralog/paralogon information.
PMID 17984073 · PMC2238967 · Nucleic acids research · 2008 · 8 claims · 8 setups
EPGD is a gene-centered, internet-accessible database integrating paralog family and paralogon information for 26 eukaryotic genomes.
-
Full-text index only
Differentiation of core promoter architecture between plants and mammals revealed by LDSS analysis.
PMID 17855401 · PMC2094075 · Nucleic acids research · 2007 · 7 claims · 8 setups
LDSS analysis identifies octamer sequences with localized distribution profiles as promoter constituents, classifiable into groups (REG, TATA, Inr, Kozak, CpG, Y Patch)
-
Full-text index only
BIPASS: BioInformatics Pipeline Alternative Splicing Services.
PMID 17584795 · PMC1933140 · Nucleic acids research · 2007 · 8 claims · 4 setups
BIPASS offers two complementary services for alternative splicing (AS) research: BIPAS-SpliceDB, a queryable pre-computed AS data warehouse, and BIPAS-Align&Splice, an online pipeline for user-submitted sequences.
-
Full-text index only
MBGD update 2010: toward a comprehensive resource for exploring microbial genome diversity.
PMID 19906735 · PMC2808943 · Nucleic acids research · 2010 · 8 claims · 6 setups
MBGD allows users to create ortholog groups using a specified subgroup of organisms, distinguishing it from other comparative genomics resources
-
Full-text index only
Shotgun haplotyping: a novel method for surveying allelic sequence variation.
PMID 16221968 · PMC1253838 · Nucleic acids research · 2005 · 8 claims · 7 setups
A novel shotgun haplotyping method generates haplotypic sequences from long PCR products by shotgun sequencing both alleles concurrently and using read-pair information to separate alleles during assembly
-
Full-text index only
EPD in its twentieth year: towards complete promoter coverage of selected model organisms.
PMID 16381980 · PMC1347508 · Nucleic acids research · 2006 · 7 claims · 4 setups
EPD is an annotated, non-redundant collection of experimentally defined eukaryotic POL II promoters accessed via genome position pointers.
-
Full-text index only
The RCSB PDB information portal for structural genomics.
PMID 16381872 · PMC1347482 · Nucleic acids research · 2006 · 7 claims · 5 setups
The RCSB PDB Structural Genomics Information Portal integrates three resources: Structural Genomics Initiatives, Targets (TargetDB/PepcDB), and Structures (functional coverage analysis).
-
Full-text index only
piRNABank: a web resource on classified and clustered Piwi-interacting RNAs.
PMID 17881367 · PMC2238943 · Nucleic acids research · 2008 · 6 claims · 4 setups
piRNABank is a web-accessible database storing empirically known piRNA sequences and annotations for human, mouse and rat.
-
Full-text index only
The protein-phosphatome of the human malaria parasite Plasmodium falciparum.
PMID 18793411 · PMC2559854 · BMC genomics · 2008 · 8 claims · 8 setups
P. falciparum possesses 27 putative protein phosphatase sequences across the four major PP families (PPP, PPM, PTP, NIF), plus 7 additional sequences predicted to dephosphorylate non-protein substrates, totaling 34.