Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Simple models of genomic variation in human SNP density.
PMID 17553150 · PMC1919371 · BMC genomics · 2007 · 6 claims · 4 setups
Hierarchical Poisson model B, which allows both the mutation-rate proxy (Beta-distributed Λ) and the ARG-size proxy (Gamma-distributed T) to vary, fits the observed SNP density distribution significantly better than models with only one or neither varying.
-
Has reproduction · 80
VGEA: an RNA viral assembly toolkit.
PMID 34567846 · PMC8428259 · PeerJ · 2021 · 8 claims · 5 setups
VGEA is a Snakemake workflow that chains existing tools (fastp, BWA, SAMtools, IVA, shiver, SeqKit, QUAST, MultiQC) into an all-in-one RNA viral genome assembly pipeline
-
Has reproduction · 100
poreCov-An Easy to Use, Fast, and Robust Workflow for SARS-CoV-2 Genome Reconstruction via Nanopore Sequencing.
PMID 34394197 · PMC8355734 · Frontiers in genetics · 2021 · 8 claims · 8 setups
poreCov is an easy-to-use, fast, and robust Nextflow-based workflow for reference-based SARS-CoV-2 genome reconstruction and lineage determination from nanopore sequencing data
-
Has reproduction · 49
EDGE COVID-19: a web platform to generate submission-ready genomes from SARS-CoV-2 sequencing efforts.
PMID 35561186 · PMC9113274 · Bioinformatics (Oxford, England) · 2022 · 7 claims · 5 setups
EDGE COVID-19 (EC-19) is a web-based platform that automates QC, reference-based variant/consensus calling, lineage determination, and submission of SARS-CoV-2 genomes and metadata to GenBank, GISAID and INSDC for both Illumina and ONT data.
-
Full-text index only
Clustering of phosphorylation site recognition motifs can be exploited to predict the targets of cyclin-dependent kinase.
PMID 17316440 · PMC1852407 · Genome biology · 2007 · 8 claims · 6 setups
CDK consensus motifs are frequently clustered (closely spaced) in known CDK substrate proteins rather than uniformly distributed
-
Full-text index only
TFBScluster web server for the identification of mammalian composite regulatory elements.
PMID 16845063 · PMC1538905 · Nucleic acids research · 2006 · 7 claims · 5 setups
TFBScluster is a web server that identifies genome-wide clusters of TFBSs conserved in multiple mammalian species using human or mouse as the reference genome.
-
Full-text index only
Ensembl's 10th year.
PMID 19906699 · PMC2808936 · Nucleic acids research · 2010 · 8 claims · 8 setups
Ensembl provides comprehensive gene annotation and integrated genomic resources (variation, regulation, comparative genomics) across a growing set of chordate genomes
-
Has reproduction · 93
Population genomics of the Wolbachia endosymbiont in Drosophila melanogaster.
PMID 23284297 · PMC3527207 · PLoS genetics · 2012 · 8 claims · 8 setups
Wolbachia infection status can be accurately predicted in silico from whole-genome shotgun sequence of individual host strains, showing 99% concordance with diagnostic PCR.
-
Full-text index only
The diploid genome sequence of an Asian individual.
PMID 18987735 · PMC2716080 · Nature · 2008 · 8 claims · 8 setups
First diploid genome sequence of an Asian (Han Chinese) individual generated using massively parallel Illumina sequencing
-
Full-text index only
Accurate whole human genome sequencing using reversible terminator chemistry.
PMID 18987734 · PMC2581791 · Nature · 2008 · 8 claims · 7 setups
A novel sequencing platform using fluorescent reversible terminator nucleotides on clonally amplified single-molecule DNA clusters generates several billion bases of accurate sequence per experiment at low cost.
-
Has reproduction · 80
TP53 engagement with the genome occurs in distinct local chromatin environments via pioneer factor activity.
PMID 25391375 · PMC4315292 · Genome research · 2015 · 8 claims · 8 setups
TP53 binding events fall into three distinct categories defined by the local chromatin environment: TSS (H3K4me3+), enhancer (H3K4me1+/H3K4me3-), and distal (H3K4me1-/H3K4me3-) peaks.
-
Full-text index only
Mitochondrial diversity within modern human populations.
PMID 17439969 · PMC1888801 · Nucleic acids research · 2007 · 8 claims · 5 setups
Modern humans show extremely low divergence from the mitochondrial consensus sequence, differing on average by only 21.6 nucleotide sites
-
Full-text index only
PigGIS: Pig Genomic Informatics System.
PMID 17090590 · PMC1669765 · Nucleic acids research · 2007 · 7 claims · 7 setups
PigGIS identified 15,700 pig consensus sequences covering 18.5 Mb of homologous human exons
-
Full-text index only
A third approach to gene prediction suggests thousands of additional human transcribed regions.
PMID 16543943 · PMC1391917 · PLoS computational biology · 2006 · 8 claims · 7 setups
A third basic concept for gene prediction exists, based on detecting strand-specific 'transcription footprints' (mutational and selectional biases) rather than gene structure or sequence similarity.
-
Has reproduction · 68
Rfam 15: RNA families database in 2025.
PMID 39526405 · PMC11701678 · Nucleic acids research · 2025 · 8 claims · 6 setups
Rfamseq was expanded to 26 106 genomes, a 76% increase, by incorporating the latest UniProt reference proteomes and additional viral genomes
-
Full-text index only
Molecular archeology of L1 insertions in the human genome.
PMID 12372140 · PMC134481 · Genome biology · 2002 · 8 claims · 4 setups
TSDfinder, a new algorithm, refines RepeatMasker-identified L1 boundaries by locating poly(A) tails, TSDs, and inversion breakpoints
-
Full-text index only
Analysis of the glutathione S-transferase (GST) gene family.
PMID 15607001 · PMC3500200 · Human genomics · 2004 · 8 claims · 4 setups
The complete human GST gene family comprises 16 genes in six subfamilies: alpha (GSTA), mu (GSTM), omega (GSTO), pi (GSTP), theta (GSTT) and zeta (GSTZ).
-
Full-text index only
SARS-CoV genome polymorphism: a bioinformatics study.
PMID 16144519 · PMC5172477 · Genomics, proteomics & bioinformatics · 2005 · 8 claims · 6 setups
SARS-CoV isolates can be classified into groups/subgroups based on the number and distribution of SNVs and INDELs relative to a 'profile' sequence, and this classification aligns with phylogenetic tree relationships and epidemiological spread.
-
Full-text index only
Using several pair-wise informant sequences for de novo prediction of alternatively spliced transcripts.
PMID 16925842 · PMC1810557 · Genome biology · 2006 · 8 claims · 4 setups
MARS, an extension of the Twinscan algorithm, uses multiple pairwise informant genomes to predict human alternatively spliced transcripts de novo without expressed sequence information.
-
Full-text index only
Identification of the REST regulon reveals extensive transposable element-mediated binding site duplication.
PMID 16899447 · PMC1557810 · Nucleic acids research · 2006 · 8 claims · 8 setups
The RE1 PSSM identifies functional RE1 binding sites with greater sensitivity and selectivity than the previously used RE1 consensus sequence