Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Novel gene and gene model detection using a whole genome open reading frame analysis in proteomics.
PMID 16646984 · PMC1557991 · Genome biology · 2006 · 8 claims · 4 setups
A six-frame genomic ORF translation used as an MS search database can detect novel peptides absent from standard protein databases, revealing incomplete genome annotation.
-
Full-text index only
High throughput sequencing and proteomics to identify immunogenic proteins of a new pathogen: the dirty genome approach.
PMID 20037647 · PMC2793016 · PloS one · 2009 · 7 claims · 7 setups
A dirty genome approach using unfinished, unclosed genome sequences combined with proteomics can rapidly identify immunogenic proteins useful for diagnostic tool development
-
Full-text index only
A genome-wide survey demonstrates widespread non-linear mRNA in expressed sequences from multiple species.
PMID 16237125 · PMC1258171 · Nucleic acids research · 2005 · 8 claims · 6 setups
A genome-wide computational survey identifies 245 genes in mammals (264 across six species) that produce RREO events in expressed sequences
-
Full-text index only
Gene losses during human origins.
PMID 16464126 · PMC1361800 · PLoS biology · 2006 · 7 claims · 7 setups
A comparative genomic screen identified 67 new human-specific nonprocessed pseudogenes, bringing the total (with 13 from prior literature) to 80 human-specific pseudogenes.
-
Full-text index only
Evolutionary genomics reveals lineage-specific gene loss and rapid evolution of a sperm-specific ion channel complex: CatSpers and CatSperbeta.
PMID 18974790 · PMC2572835 · PloS one · 2008 · 8 claims · 6 setups
The CatSper channel complex (four CatSpers plus CatSperβ) originated as early as primitive metazoans such as the Cnidarian Nematostella vectensis
-
Full-text index only
Database resources of the National Center for Biotechnology Information.
PMID 17170002 · PMC1781113 · Nucleic acids research · 2007 · 8 claims · 8 setups
NCBI maintains an integrated suite of database resources (Entrez, PubMed, RefSeq, dbSNP, BLAST, etc.) for molecular biology data retrieval and analysis
-
Full-text index only
ORFer--retrieval of protein sequences and open reading frames from GenBank and storage into relational databases or text files.
PMID 12493080 · PMC139979 · BMC bioinformatics · 2002 · 6 claims · 6 setups
ORFer retrieves protein and nucleic acid sequences and annotations from NCBI GenBank using the XML sequence format
-
Full-text index only
L1Base: from functional annotation to prediction of active LINE-1 elements.
PMID 15608246 · PMC539998 · Nucleic acids research · 2005 · 7 claims · 6 setups
L1Base is a database of putatively active LINE-1 insertions in human, mouse and rat genomes, containing FLI-L1s (intact in both ORFs), ORF2-L1s (intact ORF2, disrupted ORF1), and FLnI-L1s (full-length, >6000 bp, non-intact)
-
Full-text index only
Human genomic diversity, viral genomics and proteomics, as exemplified by human papillomaviruses and H5N1 influenza viruses.
PMID 19706363 · PMC3525194 · Human genomics · 2009 · 8 claims · 6 setups
A novel HPV type (HPV-85) was identified and phylogenetically characterized, showing closest relatedness to HPV-70/39/18/45/59 within the A7 genital HPV group
-
Full-text index only
Proteomics data repositories.
PMID 19795424 · PMC2908408 · Proteomics · 2009 · 5 claims · 5 setups
The YRC Public Data Repository (YRC PDR) provides a single unified interface disseminating multi-technology proteomics data (mass spectrometry, yeast two-hybrid, fluorescence microscopy, structure prediction) linked to protein annotations from many source databases.
-
Full-text index only
hORFeome v3.1: a resource of human open reading frames representing over 10,000 human genes.
PMID 17207965 · PMC4647941 · Genomics · 2007 · 8 claims · 7 setups
hORFeome v3.1 is a resource of 12,212 cloned human ORFs representing 10,214 genes, a 51% expansion over hORFeome v1.1
-
Has reproduction · 93
Characterization of protein isoform diversity in human umbilical vein endothelial cells via long-read proteogenomics.
PMID 36457147 · PMC9721438 · RNA biology · 2022 · 8 claims · 7 setups
Long-read RNA-seq detected 53,863 transcript isoforms from 10,426 genes in HUVECs, of which 22,195 were novel
-
Full-text index only
Characterization of 954 bovine full-CDS cDNA sequences.
PMID 16305752 · PMC1314900 · BMC genomics · 2005 · 7 claims · 8 setups
954 bovine full-length insert cDNA (bFLIC) clones representing 762 distinct loci were sequenced and characterized
-
Full-text index only
Distribution and effects of nonsense polymorphisms in human genes.
PMID 18852891 · PMC2561068 · PloS one · 2008 · 8 claims · 8 setups
Nonsense SNPs occur at a lower density than nonsynonymous SNPs, indicating stronger purifying selection against premature stop codons than amino acid changes.
-
Full-text index only
Genomic analysis of the chromosome 15q11-q13 Prader-Willi syndrome region and characterization of transcripts for GOLGA8E and WHCD1L1 from the proximal breakpoint region.
PMID 18226259 · PMC2268926 · BMC genomics · 2008 · 8 claims · 7 setups
GOLGA8E and WHDC1L1 are characterized for the first time as protein-coding transcripts from the PWS proximal breakpoint region.
-
Has reproduction · 24
MiGPC: a comprehensive catalog of enzybiotics from environmental metagenomes.
PMID 41888223 · PMC13172421 · Scientific reports · 2026 · 8 claims · 8 setups
MiGPC is the first genome-resolved metagenomic gene and protein catalog specifically targeted to enzybiotics
-
Full-text index only
A catalog of human cDNA expression clones and its application to structural genomics.
PMID 15345055 · PMC522878 · Genome biology · 2004 · 8 claims · 7 setups
A high-throughput screening approach can identify human cDNA clones from the hEx1 library that express soluble protein in E. coli
-
Full-text index only
Using ESTs to improve the accuracy of de novo gene prediction.
PMID 16817966 · PMC1534067 · BMC bioinformatics · 2006 · 8 claims · 8 setups
TWINSCAN_EST combines EST alignments with TWINSCAN via a trainable 'ESTseq' representation and improves exact gene structure prediction accuracy on the whole C. elegans genome
-
Full-text index only
Comparative analysis of cancer genes in the human and chimpanzee genomes.
PMID 16438707 · PMC1382208 · BMC genomics · 2006 · 7 claims · 6 setups
All 333 examined human cancer genes have intact, highly conserved orthologs in the chimpanzee genome (99.38% protein identity).
-
Full-text index only
Comparative genomics search for losses of long-established genes on the human lineage.
PMID 18085818 · PMC2134963 · PLoS computational biology · 2007 · 8 claims · 6 setups
A novel comparative genomics method (TransMap-based syntenic mapping of gene structures between human, mouse, and dog) can detect losses of well-established single-copy genes without relying on sequence homology to a parental gene, distinguishing them from typical duplication- or retrotransposition-derived pseudogenes.