Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Prodepth: predict residue depth by support vector regression approach from protein sequences only.
PMID 19759917 · PMC2742725 · PloS one · 2009 · 8 claims · 8 setups
Residue depth can be reliably predicted solely from protein primary sequence using support vector regression on sequence-derived features.
-
Has reproduction · 87
ddRAD-seq reveals the genetic structure and detects signals of selection in Italian brown trout.
PMID 35100964 · PMC8805291 · Genetics, selection, evolution : GSE · 2022 · 7 claims · 8 setups
Italian brown trout populations are genetically differentiated but show strong admixture introduced by stocking, especially with the Atlantic lineage.
-
Full-text index only
MODBASE, a database of annotated comparative protein structure models and associated resources.
PMID 18948282 · PMC2686492 · Nucleic acids research · 2009 · 8 claims · 8 setups
MODBASE contains 5,152,695 reliable comparative protein structure models for 1,593,209 unique protein sequences.
-
Full-text index only
Columba: an integrated database of proteins, structures, and annotations.
PMID 15801979 · PMC1087474 · BMC bioinformatics · 2005 · 8 claims · 6 setups
COLUMBA physically integrates data from twelve protein structure-related databases (PDB, KEGG, Swiss-Prot, CATH, SCOP, Gene Ontology, ENZYME, etc.) into a single PostgreSQL data warehouse.
-
Full-text index only
nsSNPAnalyzer: identifying disease-associated nonsynonymous single nucleotide polymorphisms.
PMID 15980516 · PMC1160133 · Nucleic acids research · 2005 · 6 claims · 4 setups
nsSNPAnalyzer is a web server that predicts whether a query nsSNP is disease-associated or functionally neutral using a Random Forest classifier combining structural and evolutionary information
-
Full-text index only
Information extraction from full text scientific articles: where are the keywords?
PMID 12775220 · PMC166134 · BMC bioinformatics · 2003 · 8 claims · 5 setups
The keyword content of the five article sections (A, I, M, R, D) is heterogeneous, i.e., different sections carry different kinds of information.
-
Full-text index only
Functional coverage of the human genome by existing structures, structural genomics targets, and homology models.
PMID 16118666 · PMC1188274 · PLoS computational biology · 2005 · 8 claims · 5 setups
Existing PDB structures provide single-domain coverage for 37% of functional classes in the human genome and complete (whole-protein) structure coverage for 25%.
-
Has reproduction · 86
Prediction, syntax and semantic grounding in the brain and large language models.
PMID 41807493 · PMC12979642 · Scientific reports · 2026 · 8 claims · 6 setups
Nouns show significant pre-onset neural activity, suggesting enhanced anticipatory processing of this word class.
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes
-
Has reproduction · 85
Predicting the pathogenicity of missense variants using features derived from AlphaFold2.
PMID 37084271 · PMC10203375 · Bioinformatics (Oxford, England) · 2023 · 6 claims · 8 setups
AlphaFold2-derived structural features (solvent accessibility, amino acid network features, physicochemical environment, pLDDT) can be used to train a random forest classifier (AlphScore) that distinguishes proxy-benign from proxy-pathogenic missense variants.
-
Full-text index only
Non-linear mapping for exploratory data analysis in functional genomics.
PMID 15661072 · PMC548129 · BMC bioinformatics · 2005 · 8 claims · 8 setups
A relaxation method for non-linear mapping adapts one pair of points per step rather than all points at once, and was originally shown by Chang and Lee to outperform Sammon's mapping in cluster detection effectiveness and computational efficiency.
-
Full-text index only
GeneMark: web software for gene finding in prokaryotes, eukaryotes and viruses.
PMID 15980510 · PMC1160247 · Nucleic acids research · 2005 · 8 claims · 2 setups
The GeneMark website provides web interfaces to the GeneMark family of ab initio gene-finding programs for prokaryotic, eukaryotic and viral genomic sequences
-
Has reproduction · 49
Insights into the differentiation and adaptation within Circaeasteraceae from Circaeaster agrestis genome sequencing and resequencing.
PMID 36895650 · PMC9988679 · iScience · 2023 · 8 claims · 8 setups
C. agrestis and K. uniflora are sister species with contrasting reproductive modes, providing a natural system to test effects of sexual vs asexual reproduction on genome evolution
-
Full-text index only
A procedure for the detection of linkage with high density SNP arrays in a large pedigree with colorectal cancer.
PMID 17222328 · PMC1784097 · BMC cancer · 2007 · 7 claims · 8 setups
A workflow combining Alohomora, Mega2, MENDEL, SNPLINK and SimWalk2 enables linkage analysis with high-density SNP arrays in large pedigrees (>35-40 bits) that exceed the capacity of single existing programs
-
Full-text index only
The genome of Brugia malayi - all worms are not created equal.
PMID 18952001 · PMC2668601 · Parasitology international · 2009 · 8 claims · 8 setups
Comparative genome analysis shows conserved long-range synteny but divergent local gene order between B. malayi and C. elegans, reflecting distinct evolutionary trajectories of parasitic and free-living lineages.
-
Full-text index only
Evolutionary history of the UCP gene family: gene duplication and selection.
PMID 18980678 · PMC2584656 · BMC evolutionary biology · 2008 · 8 claims · 8 setups
The UCP gene family arose through two ancestral gene duplications early in vertebrate evolution, producing the UCP1, UCP2 and UCP3 lineages.