Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
EGASP: the human ENCODE Genome Annotation Assessment Project.
PMID 16925836 · PMC1810551 · Genome biology · 2006 · 8 claims · 6 setups
Best-performing computational gene prediction methods correctly predict at least one transcript for close to 70% of annotated genes in the ENCODE regions.
-
Full-text index only
AUGUSTUS at EGASP: using EST, protein and genomic alignments for improved gene prediction in the human genome.
PMID 16925833 · PMC1810548 · Genome biology · 2006 · 8 claims · 5 setups
AUGUSTUS predicted significantly more genes correctly than any other ab initio program in EGASP
-
Full-text index only
GeneAlign: a coding exon prediction tool based on phylogenetical comparisons.
PMID 16845010 · PMC1538901 · Nucleic acids research · 2006 · 8 claims · 5 setups
GeneAlign predicts coding exons by using signal detection (GeneSplicer/WMM) combined with CORAL, a heuristic linear-time alignment tool, to align candidate signal-flanked regions against annotated exons of a homologous organism's genes
-
Full-text index only
Performance assessment of promoter predictions on ENCODE regions in the EGASP experiment.
PMID 16925837 · PMC1810552 · Genome biology · 2006 · 6 claims · 3 setups
Promoter predictors that combine promoter prediction with gene prediction (N-SCAN, Fprom) achieve better performance than pure ab initio promoter predictors, mainly by reducing the promoter search space and false positives
-
Full-text index only
EpiToolKit--a web server for computational immunomics.
PMID 18440979 · PMC2447732 · Nucleic acids research · 2008 · 7 claims · 3 setups
EpiToolKit is a web server integrating five MHC class I and two MHC class II epitope prediction methods in a unified, user-friendly interface.
-
Full-text index only
Using ESTs to improve the accuracy of de novo gene prediction.
PMID 16817966 · PMC1534067 · BMC bioinformatics · 2006 · 8 claims · 8 setups
TWINSCAN_EST combines EST alignments with TWINSCAN via a trainable 'ESTseq' representation and improves exact gene structure prediction accuracy on the whole C. elegans genome
-
Has reproduction · 75
ResnetAge: A Resnet-Based DNA Methylation Age Prediction Method.
PMID 38247911 · PMC10813502 · Bioengineering (Basel, Switzerland) · 2023 · 8 claims · 4 setups
ResnetAge, a ResNet-based neural network using 22,278 shared Illumina 27K/450K CpG sites, predicts DNA methylation age from beta values.
-
Has reproduction · 84
Foster thy young: enhanced prediction of orphan genes in assembled genomes.
PMID 34928390 · PMC9023268 · Nucleic acids research · 2022 · 7 claims · 8 setups
Each of the five gene-prediction pipelines under-predicts orphan genes, with detection as low as 11% under one prediction scenario.
-
Full-text index only
GENCODE: producing a reference annotation for ENCODE.
PMID 16925838 · PMC1810553 · Genome biology · 2006 · 8 claims · 8 setups
GENCODE annotation combines initial manual annotation by HAVANA, experimental validation, and refinement based on results to identify protein-coding genes in ENCODE regions
-
Full-text index only
Automatic annotation of eukaryotic genes, pseudogenes and promoters.
PMID 16925832 · PMC1810547 · Genome biology · 2006 · 8 claims · 6 setups
Fgenesh++ gene prediction pipeline identifies 91% of coding nucleotides with 90% specificity
-
Full-text index only
Evolutionary trace annotation of protein function in the structural proteome.
PMID 20036248 · PMC2831211 · Journal of molecular biology · 2010 · 8 claims · 7 setups
ET-ranked residue clusters can be used to build 3D templates that predict GO function in enzymes and non-enzymes alike, without prior knowledge of functional mechanism.
-
Has reproduction · 81
SEMdag: Fast learning of Directed Acyclic Graphs via node or layer ordering.
PMID 39775401 · PMC11709272 · PloS one · 2025 · 8 claims · 5 setups
SEMdag() is a two-step order-based algorithm for fast learning of high-dimensional linear SEMs, using knowledge-based (KB) or data-driven bottom-up (BU) node/layer ordering followed by penalized (L1) DAG estimation
-
Has reproduction · 84
Pharokka: a fast scalable bacteriophage annotation tool.
PMID 36453861 · PMC9805569 · Bioinformatics (Oxford, England) · 2023 · 8 claims · 5 setups
Pharokka is a one-line, fast, scalable bacteriophage annotation tool producing standards-compliant outputs, installable via a two-line bioconda command
-
Has reproduction · 96
Scalable Prediction of Acute Myeloid Leukemia Using High-Dimensional Machine Learning and Blood Transcriptomics.
PMID 31918046 · PMC6992905 · iScience · 2020 · 8 claims · 8 setups
Data-driven, high-dimensional ML approaches that learn multivariate signatures directly from genome-wide transcriptomic data (no prior gene selection) yield accurate and robust AML classifiers.
-
Has reproduction · 76
Bayesian prediction of microbial oxygen requirement.
PMID 26913185 · PMC4743139 · F1000Research · 2013 · 7 claims · 8 setups
A naive Bayesian classifier based on presence/absence of class-associated Pfam-A domains can distinguish three oxygen requirement classes (aerobe, anaerobe, facultative anaerobe) from genome sequence, unlike prior studies that only made pairwise distinctions.
-
Full-text index only
Exogean: a framework for annotating protein-coding genes in eukaryotic genomic DNA.
PMID 16925841 · PMC1810556 · Genome biology · 2006 · 8 claims · 5 setups
Exogean is a framework using directed acyclic coloured multigraphs (DACMs) to represent biological objects (mRNA, ESTs, protein alignments, exons) and iteratively combine them into complex protein-coding transcript models.
-
Full-text index only
MACSIMS: multiple alignment of complete sequences information management system.
PMID 16792820 · PMC1539025 · BMC bioinformatics · 2006 · 8 claims · 5 setups
MACSIMS is a multiple alignment-based information management system combining knowledge-based database mining with ab initio sequence predictions
-
Has reproduction · 95
Mouse-Geneformer: A deep learning model for mouse single-cell transcriptome and its cross-species utility.
PMID 40106407 · PMC11964219 · PLoS genetics · 2025 · 7 claims · 6 setups
Mouse-Geneformer, a Transformer Encoder model pre-trained via masked-token self-supervised learning on mouse-Genecorpus-20M, was successfully constructed following the original human Geneformer architecture.
-
Full-text index only
ASPIC: a web resource for alternative splicing prediction and transcript isoforms characterization.
PMID 16845044 · PMC1538898 · Nucleic acids research · 2006 · 8 claims · 2 setups
The ASPIC algorithm, using an optimization procedure that minimizes splice site predictions and transcript isoforms from multiple EST-genome alignments, outperforms other similar AS-prediction tools in sensitivity and selectivity
-
Full-text index only
Analysis of protein sequence and interaction data for candidate disease gene prediction.
PMID 17020920 · PMC1636487 · Nucleic acids research · 2006 · 8 claims · 7 setups
Combining CPS and CMP using known disease genes as input achieves sensitivity 0.52 and specificity 0.97, reducing candidate lists 13-fold