Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Towards the identification of essential genes using targeted genome sequencing and comparative analysis.
PMID 17052348 · PMC1624830 · BMC genomics · 2006 · 8 claims · 8 setups
Phyletic retention (ortholog presence across organisms) is the single most predictive feature of gene essentiality in both E. coli and S. cerevisiae.
-
Full-text index only
A comparison of classification methods for predicting Chronic Fatigue Syndrome based on genetic data.
PMID 19772600 · PMC2765429 · Journal of translational medicine · 2009 · 7 claims · 3 setups
The naive Bayes model with the wrapper-based feature selection approach performed best among all predictive models tested for distinguishing CFS from controls.
-
Full-text index only
A scale space approach for unsupervised feature selection in mass spectra classification for ovarian cancer detection.
PMID 19828085 · PMC2762074 · BMC bioinformatics · 2009 · 7 claims · 1 setups
A scale-space based unsupervised feature extraction method combined with SVM classification achieves high accuracy in ovarian cancer detection from serum mass spectra.
-
Has reproduction · 72
Prediction of prognostic signatures in triple-negative breast cancer based on the differential expression analysis via NanoString nCounter immune panel.
PMID 33138797 · PMC7607642 · BMC cancer · 2020 · 8 claims · 8 setups
edgeR-based DEG selection is more appropriate for feature selection than Elastic Net when sample sizes are small.
-
Full-text index only
Towards alignment independent quantitative assessment of homology detection.
PMID 17205117 · PMC1762415 · PloS one · 2006 · 8 claims · 6 setups
The Fhom Estimator uses the prevalence of a conserved protein feature (X) in two protein sets to estimate the fraction of true homologs among paired proteins, independent of alignment quality.
-
Has reproduction · 83
Gene-expression patterns in peripheral blood classify familial breast cancer susceptibility.
PMID 26538066 · PMC4634735 · BMC medical genomics · 2015 · 8 claims · 7 setups
A multigene expression biomarker from PBMCs accurately classifies familial breast cancer (FBC) status
-
Has reproduction · 75
stDyer-image improves clustering analysis of spatially resolved transcriptomics and proteomics with morphological images.
PMID 41692960 · PMC12960910 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 5 setups
stDyer-image directly associates the image modality with predicted cluster labels rather than using images to enhance/impute gene expression data
-
Full-text index only
MACSIMS: multiple alignment of complete sequences information management system.
PMID 16792820 · PMC1539025 · BMC bioinformatics · 2006 · 8 claims · 5 setups
MACSIMS is a multiple alignment-based information management system combining knowledge-based database mining with ab initio sequence predictions
-
Has reproduction · 100
A workflow reproducibility scale for automatic validation of biological interpretation results.
PMID 37150537 · PMC10164546 · GigaScience · 2022 · 8 claims · 4 setups
Comparing output files by checksum alone is insufficient to verify reproducibility, since checksums can differ even when the underlying biological interpretation is unchanged
-
Full-text index only
A comprehensive literature review of haplotyping software and methods for use with unrelated individuals.
PMID 15814067 · PMC3525117 · Human genomics · 2005 · 7 claims · 2 setups
Forty-six haplotyping programs were identified and reviewed, split into 43 designed for individual genotype data and three designed for pooled DNA samples.
-
Full-text index only
SNP-RFLPing: restriction enzyme mining for SNPs in genomes.
PMID 16503968 · PMC1386656 · BMC genomics · 2006 · 8 claims · 2 setups
SNP-RFLPing accepts three flexible input types (dbSNP rs#/ss# IDs, HUGO gene name/Entrez gene ID, or free-form SNP-in-sequence including IUPAC or [dNTP1/dNTP2] formats) for human, rat, and mouse genomes
-
Has reproduction · 59
Refining breast cancer biomarker discovery and drug targeting through an advanced data-driven approach.
PMID 38253993 · PMC10810249 · BMC bioinformatics · 2024 · 8 claims · 8 setups
The BGWO_SA_Ens algorithm (hybrid BGWO + simulated annealing with an ensemble classifier objective function) selects predictive breast cancer biomarker genes with high classification performance
-
Has reproduction · 100
Smart spatial omics (S2-omics) optimizes region of interest selection to capture molecular heterogeneity in diverse tissues.
PMID 41298871 · PMC12662399 · Nature cell biology · 2025 · 7 claims · 6 setups
S2-omics is an end-to-end workflow that automatically selects ROIs from H&E histology images to maximize molecular information content for spatial omics profiling.
-
Has reproduction · 44
An OMICs-based meta-analysis to support infection state stratification.
PMID 33560295 · PMC8388022 · Bioinformatics (Oxford, England) · 2021 · 7 claims · 6 setups
Multi-class Random Forest models built from meta-analyzed blood gene expression data can predict infection state (bacterial/viral/none) with high accuracy, correctly classifying 93% of bacterial and 89% of viral samples in the best model.
-
Full-text index only
Constructing support vector machine ensembles for cancer classification based on proteomic profiling.
PMID 16689692 · PMC5173238 · Genomics, proteomics & bioinformatics · 2005 · 7 claims · 4 setups
CSVME, built by selecting a subset of base SVMs via SVM-RFE ranking and fusing them with a trained upper-layer SVM, achieves better classification performance than an ensemble of all base SVMs.
-
Full-text index only
miRBase: tools for microRNA genomics.
PMID 17991681 · PMC2238936 · Nucleic acids research · 2008 · 8 claims · 6 setups
miRBase release 10.0 contains 5071 miRNA hairpin loci from 58 species, expressing 5922 distinct mature miRNA sequences, a growth of over 2000 sequences in 2 years
-
Full-text index only
Fast-evolving noncoding sequences in the human genome.
PMID 17578567 · PMC2394770 · Genome biology · 2007 · 8 claims · 6 setups
1,356 conserved noncoding sequences show human-specific accelerated substitution rates (ANC sequences) relative to chimpanzee
-
Full-text index only
Next-generation high-density self-assembling functional protein arrays.
PMID 18469824 · PMC3070491 · Nature methods · 2008 · 8 claims · 7 setups
A next-generation NAPPA method produces high-density protein microarrays displaying over 1500 unique proteins with >90% expression success
-
Full-text index only
VarDetect: a nucleotide sequence variation exploratory tool.
PMID 19091032 · PMC2638149 · BMC bioinformatics · 2008 · 8 claims · 2 setups
VarDetect is a stand-alone software tool that automatically detects nucleotide variation (SNPs) from fluorescence-based chromatogram traces using pre-calculated peak content ratios and artifact-handling rules.
-
Has reproduction · 55
Identification of TYR, TYRP1, DCT and LARP7 as related biomarkers and immune infiltration characteristics of vitiligo via comprehensive strategies.
PMID 34107850 · PMC8806433 · Bioengineered · 2021 · 7 claims · 8 setups
131 robust DEGs (89 upregulated, 42 downregulated) were identified by RRA integration of three vitiligo datasets and are closely associated with melanogenesis and vitiligo development