Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 63
RummaGEO: Automatic mining of human and mouse gene sets from GEO.
PMID 39569206 · PMC11573963 · Patterns (New York, N.Y.) · 2024 · 8 claims · 7 setups
RummaGEO is a gene expression signature search engine built from automatically mined human and mouse RNA-seq perturbation studies in GEO
-
Full-text index only
Using ESTs to improve the accuracy of de novo gene prediction.
PMID 16817966 · PMC1534067 · BMC bioinformatics · 2006 · 8 claims · 8 setups
TWINSCAN_EST combines EST alignments with TWINSCAN via a trainable 'ESTseq' representation and improves exact gene structure prediction accuracy on the whole C. elegans genome
-
Full-text index only
Exogean: a framework for annotating protein-coding genes in eukaryotic genomic DNA.
PMID 16925841 · PMC1810556 · Genome biology · 2006 · 8 claims · 5 setups
Exogean is a framework using directed acyclic coloured multigraphs (DACMs) to represent biological objects (mRNA, ESTs, protein alignments, exons) and iteratively combine them into complex protein-coding transcript models.
-
Full-text index only
FEDRANN: effective long-read overlap detection based on dimensionality reduction and approximate nearest neighbors.
PMID 42102720 · PMC13201080 · GigaScience · 2026 · 8 claims · 6 setups
A pipeline combining IDF transformation, sparse random projection (SRP), and NNDescent (the FEDRANN strategy) enables accurate overlap detection across diverse long-read datasets
-
Full-text index only
Ensembl 2006.
PMID 16381931 · PMC1347495 · Nucleic acids research · 2006 · 8 claims · 5 setups
Ensembl now provides annotation for 19 genomes, up from 4 the previous year, including new mammalian (Rhesus macaque, Opossum), chordate (Ciona intestinalis), and yeast genomes.
-
Has reproduction · 80
SLDMS: A Tool for Calculating the Overlapping Regions of Sequences.
PMID 35046988 · PMC8761809 · Frontiers in plant science · 2021 · 8 claims · 5 setups
SLDMS is a novel method for computing overlapping regions of sequencing reads using suffix array (SA), longest common prefix (LCP) array, document array (DA), and a monotonic stack.
-
Full-text index only
RaMBat: Accurate identification of medulloblastoma subtypes from diverse data sources with severe batch effects.
PMID 41571436 · PMC13060657 · Molecular oncology · 2026 · 7 claims · 5 setups
RaMBat achieves a median accuracy of 99% across 13 independent benchmark datasets, significantly outperforming state-of-the-art MB subtyping methods and conventional ML classifiers
-
Full-text index only
GoMiner: a resource for biological interpretation of genomic and proteomic data.
PMID 12702209 · PMC154579 · Genome biology · 2003 · 8 claims · 4 setups
GoMiner organizes 'interesting' gene lists (e.g., differentially expressed genes) into the Gene Ontology hierarchy for biological interpretation, displaying results as both a tree and a directed acyclic graph (DAG).
-
Full-text index only
Designating eukaryotic orthology via processed transcription units.
PMID 18445630 · PMC2425467 · Nucleic acids research · 2008 · 8 claims · 5 setups
Existing ortholog databases discard/ignore alternative splicing via all-against-all protein comparisons, causing ambiguous ortholog calls and misclassification of AS isoforms as in-paralogs