Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
The European Bioinformatics Institute's data resources: towards systems biology.
PMID 15608238 · PMC539980 · Nucleic acids research · 2005 · 8 claims · 5 setups
Since 2003 the EBI has launched new databases covering protein-protein interactions (IntAct), pathways (Reactome) and small molecules (ChEBI)
-
Has reproduction · 83
Macrel: antimicrobial peptide screening in genomes and metagenomes.
PMID 33384902 · PMC7751412 · PeerJ · 2020 · 8 claims · 8 setups
Macrel is an end-to-end pipeline that predicts high-quality AMP candidates from peptides, contigs, or reads of (meta)genomes
-
Full-text index only
MODBASE, a database of annotated comparative protein structure models and associated resources.
PMID 18948282 · PMC2686492 · Nucleic acids research · 2009 · 8 claims · 8 setups
MODBASE contains 5,152,695 reliable comparative protein structure models for 1,593,209 unique protein sequences.
-
Full-text index only
Wiggle-predicting functionally flexible regions from primary sequence.
PMID 16839194 · PMC1500818 · PLoS computational biology · 2006 · 7 claims · 6 setups
A GNM-derived, correlation-weighted 'FF score' can objectively define functionally flexible regions (FFRs) that match experimentally confirmed flexible/functional regions (hinges, recognition loops, catalytic loops).
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Has reproduction · 50
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization.
PMID 41266599 · PMC12635123 · Communications biology · 2025 · 8 claims · 8 setups
BiRNA-BERT uses adaptive dual-tokenization that dynamically selects nucleotide-level (NUC) or byte-pair encoding (BPE) tokens based on input sequence length
-
Has reproduction · 78
GenTB: A user-friendly genome-based predictor for tuberculosis resistance powered by machine learning.
PMID 34461978 · PMC8407037 · Genome medicine · 2021 · 8 claims · 6 setups
GenTB is a free, open, web-based application offering two ML predictors (Random Forest and WDNN) that predict resistance to 13 and 10 anti-TB drugs, respectively.
-
Full-text index only
Ab initio identification of human microRNAs based on structure motifs.
PMID 18088431 · PMC2238772 · BMC bioinformatics · 2007 · 8 claims · 7 setups
MiRPred predicts miRNA precursors ab initio using only predicted secondary structure motifs, ignoring nucleotide sequence
-
Full-text index only
An SVM-based system for predicting protein subnuclear localizations.
PMID 16336650 · PMC1325059 · BMC bioinformatics · 2005 · 7 claims · 3 setups
New kernels defined on k-peptide vectors mapped by BLOSUM62-based high-scored pair matrices (D1, D2, D3) improve SVM discrimination of protein subnuclear localization compared to conventional k-peptide encodings.
-
Full-text index only
MODBASE: a database of annotated comparative protein structure models and associated resources.
PMID 16381869 · PMC1347422 · Nucleic acids research · 2006 · 8 claims · 7 setups
MODBASE is a database of automatically calculated comparative protein structure models covering all UniProt sequences matchable to a known structure
-
Has reproduction · 82
Reusable building blocks in biological systems.
PMID 30958230 · PMC6303794 · Journal of the Royal Society, Interface · 2018 · 8 claims · 5 setups
Biological systems can be decomposed into phenotypic building blocks (PBBs) whose reusability ranges from single-use (condition-specific) to constitutive
-
Has reproduction · 67
Adaptive learning embedding features to improve the predictive performance of SARS-CoV-2 phosphorylation sites.
PMID 37847658 · PMC10628388 · Bioinformatics (Oxford, England) · 2023 · 8 claims · 6 setups
PSPred-ALE outperforms state-of-the-art SARS-CoV-2 phosphorylation site predictors (e.g. DeepIPs) and handcrafted feature-based methods in benchmarking comparisons
-
Full-text index only
Filtering high-throughput protein-protein interaction data using a combination of genomic features.
PMID 15833142 · PMC1127019 · BMC bioinformatics · 2005 · 8 claims · 8 setups
A combination of three genomic features (interacting Pfam domains, GO annotations, sequence homology) using naive Bayesian networks predicts true protein-protein interactions with high sensitivity and good specificity.
-
Full-text index only
Evolutionary trace annotation of protein function in the structural proteome.
PMID 20036248 · PMC2831211 · Journal of molecular biology · 2010 · 8 claims · 7 setups
ET-ranked residue clusters can be used to build 3D templates that predict GO function in enzymes and non-enzymes alike, without prior knowledge of functional mechanism.
-
Has reproduction · 89
MirDIP 5.2: tissue context annotation and novel microRNA curation.
PMID 36453996 · PMC9825511 · Nucleic acids research · 2023 · 7 claims · 6 setups
mirDIP 5.2 removed eight outdated resources, added miRNATIP, and ran five prediction algorithms against miRBase and mirGeneDB miRNAs to expand and improve interaction coverage
-
Has reproduction · 79
Interpretable prediction models for widespread m6A RNA modification across cell lines and tissues.
PMID 37995291 · PMC10697738 · Bioinformatics (Oxford, England) · 2023 · 8 claims · 8 setups
CLSM6A is a set of CNN-based deep learning models that predict single-nucleotide-resolution m6A RNA modification sites across eight cell lines and three tissues in H. sapiens
-
Has reproduction · 85
PowerBacGWAS: a computational pipeline to perform power calculations for bacterial genome-wide association studies.
PMID 35338232 · PMC8956664 · Communications biology · 2022 · 8 claims · 8 setups
Two computational approaches (sub-sampling and phenotype-simulation) can be implemented to perform power calculations for bacterial GWAS using existing genome collections, packaged as the PowerBacGWAS pipeline