Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
CpG_MI: a novel approach for identifying functional CpG islands in mammalian genomes.
PMID 19854943 · PMC2800233 · Nucleic acids research · 2010 · 8 claims · 6 setups
Functional ('bona fide') CGIs show distinct average/cumulative mutual information (AMI/CMI) distributions of neighboring CpG distances compared to non-functional CGIs and random genome segments
-
Has reproduction · 74
Wide-Open: Accelerating public data release by automating detection of overdue datasets.
PMID 28594819 · PMC5464523 · PLoS biology · 2017 · 6 claims · 5 setups
A general text-mining + API-query approach (Wide-Open) can automatically identify datasets that are overdue for public release in a repository
-
Full-text index only
Global distribution of rubella virus genotypes.
PMID 14720390 · PMC3034328 · Emerging infectious diseases · 2003 · 8 claims · 6 setups
Phylogenetic analysis of 103 E1 gene sequences from 17 countries confirms at least two rubella virus genotypes, RGI and RGII
-
Full-text index only
A genome-wide survey demonstrates widespread non-linear mRNA in expressed sequences from multiple species.
PMID 16237125 · PMC1258171 · Nucleic acids research · 2005 · 8 claims · 6 setups
A genome-wide computational survey identifies 245 genes in mammals (264 across six species) that produce RREO events in expressed sequences
-
Full-text index only
Recovery of bisulfite-converted genomic sequences in the methylation-sensitive QPCR.
PMID 17439964 · PMC1888819 · Nucleic acids research · 2007 · 8 claims · 7 setups
Bisulfite treatment causes DNA strand breakage (via abasic site formation and beta-elimination) in addition to cytosine deamination.
-
Full-text index only
Potential biomarkers of human salivary function: a modified proteomic approach.
PMID 18804197 · PMC2633945 · Archives of oral biology · 2009 · 6 claims · 6 setups
Two SDS-PAGE bands, identified by MS-MS as statherin and a truncated (N-terminal 8-aa-missing) cystatin S, are the strongest and most consistent predictors of HAA/LAA group membership and clinical/microbiological outcomes
-
Full-text index only
Analysis of recent segmental duplications in the bovine genome.
PMID 19951423 · PMC2796684 · BMC genomics · 2009 · 8 claims · 6 setups
Recently duplicated sequence (≥1 kb, ≥90% identity) comprises 3.11% (94.4 Mb) of the bovine genome assembly (Btau_4.0)
-
Full-text index only
Identifying cis-regulatory sequences by word profile similarity.
PMID 19730735 · PMC2731932 · PloS one · 2009 · 8 claims · 8 setups
WPH-finder identifies putative co-regulated CRMs by scanning the genome for sequences with word profiles similar to a known CRM, without explicitly defining binding sites
-
Full-text index only
TEPEAK: A novel method for identifying and characterizing polymorphic transposable elements in non-model species populations.
PMID 41494038 · PMC12788660 · PLoS computational biology · 2026 · 8 claims · 6 setups
TEPEAK identifies and characterizes polymorphic TEs in populations without any prior TE sequence or loci information, using only a chromosome-level reference assembly.
-
Full-text index only
An online database for brain disease research.
PMID 16594998 · PMC1489945 · BMC genomics · 2006 · 7 claims · 5 setups
SMRIDB is a comprehensive web-based database integrating gene expression data and clinical metadata to aid understanding of the genetic effects of brain disease (bipolar disorder, schizophrenia, depression)
-
Has reproduction · 53
Estimates of recent and historical effective population size in turbot, seabream, seabass and carp selective breeding programmes.
PMID 34742227 · PMC8572424 · Genetics, selection, evolution : GSE · 2021 · 8 claims · 8 setups
Current (recent) effective population size is equal to or less than 50 fish in all analysed farmed populations
-
Full-text index only
POCUS: mining genomic sequence annotation to predict disease genes.
PMID 14611661 · PMC329128 · Genome biology · 2003 · 8 claims · 6 setups
Genes predisposing to the same disease tend to share functional annotation IDs (GO/InterPro) more than expected by chance
-
Full-text index only
A map of human protein interactions derived from co-expression of human mRNAs and their orthologs.
PMID 18414481 · PMC2387231 · Molecular systems biology · 2008 · 8 claims · 6 setups
Comparing human mRNA co-expression with co-expression of orthologous gene pairs in five other organisms identifies proteins that physically associate
-
Full-text index only
Discovering cancer genes by integrating network and functional properties.
PMID 19765316 · PMC2758898 · BMC medical genomics · 2009 · 8 claims · 6 setups
Cancer genes have distinct PPI network topology (higher connectivity, higher clustering coefficient, shorter path length to known cancer genes) compared to non-cancer genes
-
Full-text index only
Retentive Network promotes efficient RNA language modeling of long sequences.
PMID 41814064 · PMC13111708 · Communications biology · 2026 · 8 claims · 6 setups
RNAret, a RetNet-based RNA language model with O(n) complexity, achieves training parallelism and low computational overhead while processing long RNA sequences
-
Has reproduction · 44
Detecting DNA modifications from SMRT sequencing data by modeling sequence context dependence of polymerase kinetic.
PMID 23516341 · PMC3597545 · PLoS computational biology · 2013 · 8 claims · 7 setups
Local sequence context strongly determines position-specific polymerase kinetic rate: roughly 80% of IPD variation is explained by a 10 bp context (7 bases upstream, 2 bases downstream of the incorporation site), saturating at 7 bases upstream.
-
Has reproduction · 76
Bayesian prediction of microbial oxygen requirement.
PMID 26913185 · PMC4743139 · F1000Research · 2013 · 7 claims · 8 setups
A naive Bayesian classifier based on presence/absence of class-associated Pfam-A domains can distinguish three oxygen requirement classes (aerobe, anaerobe, facultative anaerobe) from genome sequence, unlike prior studies that only made pairwise distinctions.
-
Full-text index only
Using ESTs to improve the accuracy of de novo gene prediction.
PMID 16817966 · PMC1534067 · BMC bioinformatics · 2006 · 8 claims · 8 setups
TWINSCAN_EST combines EST alignments with TWINSCAN via a trainable 'ESTseq' representation and improves exact gene structure prediction accuracy on the whole C. elegans genome
-
Full-text index only
Does distance matter? Variations in alternative 3' splicing regulation.
PMID 17704130 · PMC2018619 · Nucleic acids research · 2007 · 8 claims · 7 setups
Alternative 3' splice sites can be distinguished from constitutive splice sites by a combination of sequence/conservation properties that vary depending on the distance between the splice sites.
-
Has reproduction · 81
Enabling Single-Cell Drug Response Annotations from Bulk RNA-Seq Using SCAD.
PMID 36762572 · PMC10104628 · Advanced science (Weinheim, Baden-Wurttemberg, Germany) · 2023 · 7 claims · 7 setups
SCAD, a transfer learning framework integrating adversarial discriminative domain adaptation (ADDA), can infer single-cell drug sensitivities by transferring knowledge from bulk RNA-seq pharmacogenomic data (GDSC) to scRNA-seq target domains