Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
The whole alignment and nothing but the alignment: the problem of spurious alignment flanks.
PMID 18796526 · PMC2566872 · Nucleic acids research · 2008 · 8 claims · 4 setups
Some common scoring schemes tend to overextend alignments, generating spurious alignment flanks up to hundreds of bp/amino acids in length
-
Full-text index only
A comparison of random sequence reads versus 16S rDNA sequences for estimating the biodiversity of a metagenomic library.
PMID 18682527 · PMC2532719 · Nucleic acids research · 2008 · 8 claims · 7 setups
Biodiversity observed by RSR analysis is consistent with that obtained by 16S rDNA analysis
-
Has reproduction · 60
Deconvolution of the hematopoietic stem cell microenvironment reveals a high degree of specialization and conservation.
PMID 35494238 · PMC9046238 · iScience · 2022 · 8 claims · 6 setups
A customized bootstrapping/random-forest divide-and-conquer clustering pipeline integrating three scRNA-seq datasets robustly resolves cell states despite high cell-to-cell similarity within compartments
-
Has reproduction · 59
Application of Machine Learning in Predicting Hepatic Metastasis or Primary Site in Gastroenteropancreatic Neuroendocrine Tumors.
PMID 37887568 · PMC10605255 · Current oncology (Toronto, Ont.) · 2023 · 8 claims · 7 setups
Multi-gene random forest models classify primary tumor vs. liver metastasis samples with 100% accuracy in training/test cohorts and >90% accuracy in an independent validation cohort
-
Full-text index only
A surrogate-based approach for post-genomic partner identification.
PMID 11602024 · PMC57814 · BMC biotechnology · 2001 · 8 claims · 5 setups
Peptide surrogates derived from random phage display libraries contain amino acid sequence information that identifies the natural biological partner of the panned target via database searching.
-
Full-text index only
Protein interaction networks by proteome peptide scanning.
PMID 14737190 · PMC314469 · PLoS biology · 2004 · 8 claims · 7 setups
WISE (combining phage display-derived relaxed consensus patterns with SPOT peptide synthesis arrays) can identify proteome-wide binding partners of a peptide-recognition domain
-
Full-text index only
Evolutionary distance estimation and fidelity of pair wise sequence alignment.
PMID 15840174 · PMC1087827 · BMC bioinformatics · 2005 · 8 claims · 8 setups
Evolutionary distance estimation is relatively unaffected by alignment error as long as 50% or more of homologous sites remain identical between sequences
-
Full-text index only
The fragile breakage versus random breakage models of chromosome evolution.
PMID 16501665 · PMC1378107 · PLoS computational biology · 2006 · 8 claims · 6 setups
Sankoff and Trinh's synteny block identification algorithm (ST-Synteny) is flawed, producing erroneous block identifications even in small toy examples.
-
Full-text index only
Integrated analysis of genetic and proteomic data identifies biomarkers associated with adverse events following smallpox vaccination.
PMID 18923431 · PMC2692715 · Genes and immunity · 2009 · 7 claims · 6 setups
A two-stage strategy (Random Forest filtering followed by decision tree modeling) can integrate categorical genetic and continuous proteomic data to identify biomarkers of AE risk
-
Full-text index only
Metagenomic analysis of human diarrhea: viral detection and discovery.
PMID 18398449 · PMC2290972 · PLoS pathogens · 2008 · 8 claims · 7 setups
Micro-mass sequencing (minimal stool input, minimal purification, ~384 reads/sample) can detect known enteric viruses in diarrhea specimens
-
Full-text index only
A statistical change point model approach for the detection of DNA copy number variations in array CGH data.
PMID 19875853 · PMC4154476 · IEEE/ACM transactions on computational biology and bioinformatics · 2009 · 7 claims · 4 setups
A novel mean and variance change point model (MVCM) is proposed to detect CNVs/breakpoints in aCGH data.
-
Full-text index only
Application of two machine learning algorithms to genetic association studies in the presence of covariates.
PMID 19014573 · PMC2620353 · BMC genetics · 2008 · 8 claims · 3 setups
The relative performance of RF and MARS for detecting genotype-trait associations depends on both the strategy used to handle covariates and the true underlying model of association (e.g., confounding vs. mediation vs. interaction).
-
Has reproduction · 65
High-throughput sequencing SELEX for the determination of DNA-binding protein specificities in vitro.
PMID 35776646 · PMC9243297 · STAR protocols · 2022 · 8 claims · 8 setups
HT-SELEX enables unbiased, in vitro determination of preferred DNA target motifs for DNA-binding proteins by iterative selection and PCR amplification of bound oligonucleotides
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Full-text index only
The association of Alu repeats with the generation of potential AU-rich elements (ARE) at 3' untranslated regions.
PMID 15610565 · PMC544599 · BMC genomics · 2004 · 6 claims · 4 setups
Alu repeats are a source of AREs at 3' UTRs of human mRNA, via poly-A regions of Alu generating complementary poly-T/poly-U regions that acquire regular adenine insertions to form ARE motifs.
-
Full-text index only
Deducing topology of protein-protein interaction networks from experimentally measured sub-networks.
PMID 18598366 · PMC2474618 · BMC bioinformatics · 2008 · 7 claims · 6 setups
Experimentally measured protein-protein interaction sub-networks are not random samples of their parent networks.
-
Full-text index only
Disease-aging network reveals significant roles of aging genes in connecting genetic diseases.
PMID 19779549 · PMC2739292 · PLoS computational biology · 2009 · 8 claims · 8 setups
Human disease genes are much closer to aging genes in the PPI network than expected by chance
-
Has reproduction · 83
Accurate prediction of metagenome-assembled genome completeness by MAGISTA, a random forest model built on alignment-free intra-bin statistics.
PMID 35248155 · PMC8898458 · Environmental microbiome · 2022 · 7 claims · 7 setups
MAGISTA, a random forest model built on alignment-free intra-bin distance-distribution statistics, can estimate MAG completeness and purity without relying on reference marker genes.
-
Has reproduction · 68
Molecular subtype of recurrent implantation failure reveals distinct endometrial etiology of female infertility.
PMID 40660214 · PMC12257665 · Journal of translational medicine · 2025 · 8 claims · 8 setups
RIF endometrial samples segregate into two reproducible molecular subtypes: an immune-driven subtype (RIF-I) and a metabolic-driven subtype (RIF-M)
-
Has reproduction · 48
Improved epigenetic age prediction models by combining sex chromosome and autosomal markers.
PMID 40665390 · PMC12261677 · Epigenetics & chromatin · 2025 · 7 claims · 5 setups
Combining sex chromosomal DNAm markers with autosomal age-informative markers can produce a high-accuracy age prediction model competitive with autosomal-only models