Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Including microbiome information in a multi-trait genomic evaluation: a case study on longitudinal growth performance in beef cattle.
PMID 38491422 · PMC10943865 · Genetics, selection, evolution : GSE · 2024 · 8 claims · 5 setups
The host genome's influence on the functional rumen microbiome contributes to temporal variation in average daily gain (ADG1-ADG4) across finishing months in beef cattle.
-
Has reproduction · 76
Bayesian prediction of microbial oxygen requirement.
PMID 26913185 · PMC4743139 · F1000Research · 2013 · 7 claims · 8 setups
A naive Bayesian classifier based on presence/absence of class-associated Pfam-A domains can distinguish three oxygen requirement classes (aerobe, anaerobe, facultative anaerobe) from genome sequence, unlike prior studies that only made pairwise distinctions.
-
Full-text index only
Evidence for positive selection in putative virulence factors within the Paracoccidioides brasiliensis species complex.
PMID 18820744 · PMC2553485 · PLoS neglected tropical diseases · 2008 · 8 claims · 8 setups
Positive selection has played an important role in the molecular evolution of putative virulence factors of P. brasiliensis
-
Has reproduction · 86
Molecular Classification Models for Triple Negative Breast Cancer Subtype Using Machine Learning.
PMID 34575658 · PMC8472680 · Journal of personalized medicine · 2021 · 6 claims · 4 setups
A training gene set of 719 unique upregulated DEGs (subtype-specific) can be used to build ML models that classify TNBC into BLIA, BLIS, MES, and LAR subtypes.
-
Full-text index only
A survey of integral alpha-helical membrane proteins.
PMID 19760129 · PMC2780624 · Journal of structural and functional genomics · 2009 · 8 claims · 8 setups
An automated annotation pipeline defines the integral membrane genome and family associations for 21,379 proteins from 34 genomes, most belonging to 598 Pfam-derived membrane protein families.
-
Full-text index only
The origins of lactase persistence in Europe.
PMID 19714206 · PMC2722739 · PLoS computational biology · 2009 · 8 claims · 5 setups
The −13,910*T allele first underwent selection among dairying farmers around 7,500 years ago in a region between the central Balkans and central Europe, possibly linked to the Linearbandkeramik culture.
-
Has reproduction · 87
R2DT is a framework for predicting and visualising RNA secondary structure using templates.
PMID 34108470 · PMC8190129 · Nature communications · 2021 · 8 claims · 6 setups
R2DT is a template-based computational framework/pipeline that predicts and visualises RNA 2D structure in standardised, community-accepted layouts
-
Full-text index only
A model-based approach to selection of tag SNPs.
PMID 16776821 · PMC1525207 · BMC bioinformatics · 2006 · 7 claims · 5 setups
The Li and Stephens hidden Markov model outperforms other tested models (simple Markov, two-state HMM, HMM-4D, greedy GR-1/GR-2) in description code-length, tag set information content, and prediction of tagged SNPs.
-
Full-text index only
Genome-wide identification of human functional DNA using a neutral indel model.
PMID 16410828 · PMC1326222 · PLoS computational biology · 2006 · 8 claims · 8 setups
A neutral indel model predicting a geometric distribution of intergap segment (IGS) lengths fits human-mouse ancestral repeat (AR) alignment data excellently
-
Has reproduction · 78
Machine learning and free energy clustering reveal PAH protein binding linked to AD risk.
PMID 41953002 · PMC13053772 · iScience · 2026 · 7 claims · 8 setups
An integrated framework of bioinformatics, machine learning, and ΔG clustering can prioritize PAHs for AD-associated neurotoxicity.
-
Full-text index only
Predicting positive p53 cancer rescue regions using Most Informative Positive (MIP) active learning.
PMID 19756158 · PMC2742196 · PLoS computational biology · 2009 · 8 claims · 4 setups
MIP active learning is a novel active learning method that preferentially seeks informative Positive (functionally active) examples rather than only maximizing classifier accuracy.
-
Full-text index only
The evolution of human influenza A viruses from 1999 to 2006: a complete genome study.
PMID 18325125 · PMC2311284 · Virology journal · 2008 · 8 claims · 6 setups
H3N2 was the prevalent influenza A strain in Denmark from 1999 to 2006, except the 2000–2001 season when H1N1 dominated
-
Full-text index only
Identifying the important HIV-1 recombination breakpoints.
PMID 18787691 · PMC2522274 · PLoS computational biology · 2008 · 8 claims · 3 setups
Local sequence identity between co-packaged parental RNAs strongly influences the probability of strand-transfer/breakpoint location, with fewer breakpoints occurring near mismatches
-
Has reproduction · 68
Machine learning algorithm predicts fibrosis-related blood diagnosis markers of intervertebral disc degeneration.
PMID 37915003 · PMC10619283 · BMC medical genomics · 2023 · 7 claims · 7 setups
CEP120 and SPDL1 are fibrosis-related diagnostic genes for IDD, identified via a random forest model from 29 differentially expressed fibrosis-related genes
-
Has reproduction · 90
Inferring a spatial code of cell-cell interactions across a whole animal body.
PMID 36395331 · PMC9714814 · PLoS computational biology · 2022 · 8 claims · 6 setups
cell2cell computes cell-cell interaction (CCI) potential using a novel modified Bray-Curtis score based on complementary coexpression of ligand-receptor pairs between cells
-
Has reproduction · 81
Enabling Single-Cell Drug Response Annotations from Bulk RNA-Seq Using SCAD.
PMID 36762572 · PMC10104628 · Advanced science (Weinheim, Baden-Wurttemberg, Germany) · 2023 · 7 claims · 7 setups
SCAD, a transfer learning framework integrating adversarial discriminative domain adaptation (ADDA), can infer single-cell drug sensitivities by transferring knowledge from bulk RNA-seq pharmacogenomic data (GDSC) to scRNA-seq target domains
-
Has reproduction · 90
A Decentralized Kidney Transplant Biopsy Classifier for Transplant Rejection Developed Using Genes of the Banff-Human Organ Transplant Panel.
PMID 35619722 · PMC9128066 · Frontiers in immunology · 2022 · 6 claims · 6 setups
A random forest model trained solely on B-HOT panel genes (B-HOT Model) accurately classifies kidney transplant biopsies as NR, ABMR, or TCMR.
-
Full-text index only
Comparative genomics.
PMID 14624258 · PMC261895 · PLoS biology · 2003 · 8 claims · 7 setups
Conserved DNA between species tends to encode shared functional features, while divergent DNA underlies species differences
-
Has reproduction
Genome-wide signatures of convergent evolution in echolocating mammals.
PMID 24005325 · PMC3836225 · Nature · 2013 · 8 claims · 8 setups
Genome-wide convergent sequence evolution between echolocating lineages is not rare but widespread and continuously distributed, with signatures consistent with convergence in nearly 200 loci out of 2,326 examined.
-
Full-text index only
SePaCS--a web-based application for classification of seroreactivity profiles.
PMID 17478503 · PMC1933220 · Nucleic acids research · 2007 · 8 claims · 4 setups
SePaCS is a freely available web-based tool that trains and applies multiple classification methods (4 Naive Bayes variants, SVM with RBF kernel, LDA, DLDA) to seroreactivity profiles and outputs results as a summary table plus a detailed PDF report