Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
DDBJ in collaboration with mass-sequencing teams on annotation.
PMID 15608189 · PMC539974 · Nucleic acids research · 2005 · 7 claims · 5 setups
DDBJ collected and released 1,066,084 entries (718,072,425 bases) in the past year, including the complete chimpanzee chromosome 22 sequence and silkworm whole-genome shotgun data
-
Full-text index only
Integrative annotation of 21,037 human genes validated by full-length cDNA clones.
PMID 15103394 · PMC393292 · PLoS biology · 2004 · 8 claims · 5 setups
41,118 full-length human cDNAs from six high-throughput sequencing projects were exhaustively integratively characterized
-
Full-text index only
Large-scale identification and characterization of alternative splicing variants of human gene transcripts using 56,419 completely sequenced and manually annotated full-length cDNAs.
PMID 16914452 · PMC1557807 · Nucleic acids research · 2006 · 8 claims · 8 setups
Analysis of 56,419 full-length cDNAs identified 6877 alternative splicing genes encoding 18,297 alternative splicing variants made of 37,670 exons.
-
Full-text index only
Molecular archeology of L1 insertions in the human genome.
PMID 12372140 · PMC134481 · Genome biology · 2002 · 8 claims · 4 setups
TSDfinder, a new algorithm, refines RepeatMasker-identified L1 boundaries by locating poly(A) tails, TSDs, and inversion breakpoints
-
Full-text index only
L1Base: from functional annotation to prediction of active LINE-1 elements.
PMID 15608246 · PMC539998 · Nucleic acids research · 2005 · 7 claims · 6 setups
L1Base is a database of putatively active LINE-1 insertions in human, mouse and rat genomes, containing FLI-L1s (intact in both ORFs), ORF2-L1s (intact ORF2, disrupted ORF1), and FLnI-L1s (full-length, >6000 bp, non-intact)
-
Has reproduction · 68
Rfam 15: RNA families database in 2025.
PMID 39526405 · PMC11701678 · Nucleic acids research · 2025 · 8 claims · 6 setups
Rfamseq was expanded to 26 106 genomes, a 76% increase, by incorporating the latest UniProt reference proteomes and additional viral genomes
-
Has reproduction · 51
Polyploidy and the petal transcriptome of Gossypium.
PMID 24393201 · PMC3890615 · BMC plant biology · 2014 · 8 claims · 8 setups
Most homoeologous gene pairs in polyploid cotton petals are expressed at equal levels, indicating a surprising level of expression homeostasis; only ~20% of expressed genes show significant genome bias.
-
Has reproduction · 94
Systematic assessment of pathway databases, based on a diverse collection of user-submitted experiments.
PMID 36088548 · PMC9487593 · Briefings in bioinformatics · 2022 · 8 claims · 6 setups
Well-established, hierarchically organized pathway annotation systems (e.g. GO, Reactome, KEGG) yield the best overall enrichment performance despite covering much of the human genome only in general terms.
-
Full-text index only
Characterization of 954 bovine full-CDS cDNA sequences.
PMID 16305752 · PMC1314900 · BMC genomics · 2005 · 7 claims · 8 setups
954 bovine full-length insert cDNA (bFLIC) clones representing 762 distinct loci were sequenced and characterized
-
Full-text index only
AceView: a comprehensive cDNA-supported gene and transcripts annotation.
PMID 16925834 · PMC1810549 · Genome biology · 2006 · 8 claims · 4 setups
At the mRNA level, AceView transcripts are the closest match to Gencode transcripts among all evaluated methods, including alternative splice variants
-
Full-text index only
MACSIMS: multiple alignment of complete sequences information management system.
PMID 16792820 · PMC1539025 · BMC bioinformatics · 2006 · 8 claims · 5 setups
MACSIMS is a multiple alignment-based information management system combining knowledge-based database mining with ab initio sequence predictions
-
Full-text index only
Low conservation and species-specific evolution of alternative splicing in humans and mice: comparative genomics analysis using well-annotated full-length cDNAs.
PMID 18838389 · PMC2582632 · Nucleic acids research · 2008 · 7 claims · 8 setups
Although 86% of individual human exons are conserved in the mouse genome, only a small fraction (431/20392, ~2%) of human AS variants are perfectly conserved AS variants in mice.
-
Has reproduction · 52
epiGBS2: Improvements and evaluation of highly multiplexed, epiGBS-based reduced representation bisulfite sequencing.
PMID 35178872 · PMC9311447 · Molecular ecology resources · 2022 · 8 claims · 8 setups
epiGBS2 provides a laboratory protocol and revised bioinformatics pipeline for de novo cytosine methylation and SNP calling in species with or without a reference genome
-
Has reproduction · 85
Chromosome-level genome assembly of Lilford's wall lizard, Podarcis lilfordi (Günther, 1874) from the Balearic Islands (Spain).
PMID 37137526 · PMC10214862 · DNA research : an international journal for rapid publication of reports on genes and genomes · 2023 · 8 claims · 8 setups
First high-quality chromosome-level genome assembly and annotation of P. lilfordi, generated via a mixed sequencing strategy (10X linked reads, ONT long reads, Hi-C) plus RNAseq/Iso-Seq
-
Full-text index only
Columba: an integrated database of proteins, structures, and annotations.
PMID 15801979 · PMC1087474 · BMC bioinformatics · 2005 · 8 claims · 6 setups
COLUMBA physically integrates data from twelve protein structure-related databases (PDB, KEGG, Swiss-Prot, CATH, SCOP, Gene Ontology, ENZYME, etc.) into a single PostgreSQL data warehouse.
-
Has reproduction · 88
nf-core/isoseq: simple gene and isoform annotation with PacBio Iso-Seq long-read sequencing.
PMID 36961337 · PMC10199315 · Bioinformatics (Oxford, England) · 2023 · 7 claims · 4 setups
nf-core/isoseq is a new automated Nextflow-based pipeline that processes raw Iso-Seq subreads through to genome annotation (BED format) without requiring transcriptome assembly.
-
Full-text index only
Performance assessment of promoter predictions on ENCODE regions in the EGASP experiment.
PMID 16925837 · PMC1810552 · Genome biology · 2006 · 6 claims · 3 setups
Promoter predictors that combine promoter prediction with gene prediction (N-SCAN, Fprom) achieve better performance than pure ab initio promoter predictors, mainly by reducing the promoter search space and false positives
-
Full-text index only
Molecular cloning, genomic characterization and over-expression of a novel gene, XRRA1, identified from human colorectal cancer cell HCT116Clone2_XRR and macaque testis.
PMID 12908878 · PMC194569 · BMC genomics · 2003 · 8 claims · 7 setups
XRRA1 is a novel gene down-regulated ~2-fold in XR-resistant HCT116 Clone2_XRR relative to HCT116 Clone10, identified via cDNA microarray
-
Full-text index only
The 10 sea urchin receptor for egg jelly proteins (SpREJ) are members of the polycystic kidney disease-1 (PKD1) family.
PMID 17629917 · PMC1934368 · BMC genomics · 2007 · 8 claims · 5 setups
Sea urchins possess 10 SpREJ (PKD1 family) genes, compared to five in humans, all defined by possession of a ~600 residue REJ domain
-
Has reproduction · 86
RNASEQR--a streamlined and accurate RNA-seq sequence analysis program.
PMID 22199257 · PMC3315322 · Nucleic acids research · 2012 · 8 claims · 7 setups
RNASEQR is a new RNA-seq mapper/aligner that combines a BWT-based (Bowtie) transcriptomic/genomic alignment with hash-based BLAT local alignment in three sequential steps: transcriptome mapping, novel exon detection, and anchor-and-align novel splice junction identification.