Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
SNPdetector: a software tool for sensitive and accurate SNP detection.
PMID 16261194 · PMC1274293 · PLoS computational biology · 2005 · 7 claims · 7 setups
SNPdetector, which models human visual inspection of sequencing traces, achieves low false positive and false negative rates in automated SNP and mutation detection
-
Full-text index only
High-throughput chromatin information enables accurate tissue-specific prediction of transcription factor binding sites.
PMID 18988630 · PMC2662491 · Nucleic acids research · 2009 · 8 claims · 8 setups
Incorporating H3K4me3 chromatin modification estimates greatly improves the accuracy of in silico prediction of in vivo TF binding for a wide range of TFs in human and mouse
-
Has reproduction · 80
Differential analysis of RNA structure probing experiments at nucleotide resolution: uncovering regulatory functions of RNA structure.
PMID 35869080 · PMC9307511 · Nature communications · 2022 · 7 claims · 4 setups
DiffScan is a computational framework combining a Normalization module and a Scan module to identify SVRs at nucleotide resolution from SP data.
-
Full-text index only
HCK and ABAA: A Newly Designed Pipeline to Improve Fungi Metabarcoding Analysis.
PMID 34025601 · PMC8134036 · Frontiers in microbiology · 2021 · 8 claims · 8 setups
ABAA reduces the number of false-positives with all metabarcoding methods tested
-
Has reproduction · 62
Application of alternative de novo motif recognition models for analysis of structural heterogeneity of transcription factor binding sites: a case study of FOXA2 binding sites.
PMID 34547062 · PMC8408018 · Vavilovskii zhurnal genetiki i selektsii · 2021 · 8 claims · 4 setups
MultiDeNA pipeline combines PWM, diPWM, BaMM and InMoDe models to train, evaluate, threshold, and classify ChIP-seq peaks for TFBS structural heterogeneity
-
Full-text index only
The human urinary proteome contains more than 1500 proteins, including a large proportion of membrane proteins.
PMID 16948836 · PMC1794545 · Genome biology · 2006 · 8 claims · 6 setups
Identified 1543 proteins in urine from ten healthy donors while essentially eliminating false-positive identifications
-
Has reproduction · 83
ConNIS and labeling instability: New statistical methods for improving the detection of essential genes in TraDIS libraries.
PMID 41790830 · PMC12991369 · PLoS computational biology · 2026 · 7 claims · 4 setups
ConNIS provides an analytic probability distribution for the length of the longest insertion-free sequence within a gene, given gene length and expected insertion count.
-
Full-text index only
Functional annotation and identification of candidate disease genes by computational analysis of normal tissue gene expression data.
PMID 18560577 · PMC2409962 · PloS one · 2008 · 7 claims · 5 setups
Ranked Coexpression Groups (RCG) built from k=6 nearest coexpressed genes, combined with a majority-rule functional characterization, integrate multiple datasets/coexpression measures to generate high-confidence functional annotation predictions
-
Full-text index only
A machine learning approach uncovers principles and determinants of eukaryotic ribosome pausing.
PMID 39423268 · PMC11488575 · Science advances · 2024 · 8 claims · 5 setups
An unsupervised ML pipeline using the extended isolation forest (EIF) algorithm can reliably detect ribosome pausing sites from noisy, coverage-biased RiboSeq data across expression levels
-
Has reproduction · 87
CoINcIDE: A framework for discovery of patient subtypes across multiple datasets.
PMID 26961683 · PMC4784276 · Genome medicine · 2016 · 8 claims · 4 setups
CoINcIDE is a novel framework for discovering patient subtypes across multiple datasets that requires no between-dataset transformations (e.g., batch correction)
-
Full-text index only
Bayesian survival analysis in genetic association studies.
PMID 18617538 · PMC2530885 · Bioinformatics (Oxford, England) · 2008 · 7 claims · 5 setups
A novel Bayesian method (BETA-Surv) extends prior case-control haplotype-clustering work to censored survival outcomes by clustering haplotypes via gene tree/perfect phylogeny topology and relative mutation age.
-
Full-text index only
Benchmarking of methods to analyse data derived from GBS-MeDIP.
PMID 41555215 · PMC12829230 · BMC bioinformatics · 2026 · 7 claims · 4 setups
featureCounts is the most reliable tool for count matrix generation from GBS-MeDIP data, outperforming MEDIPS
-
Full-text index only
The use of edge-betweenness clustering to investigate biological function in protein interaction networks.
PMID 15740614 · PMC555937 · BMC bioinformatics · 2005 · 8 claims · 7 setups
Edge-Betweenness clustering separates protein interaction graphs into subgraphs whose GO term distributions show significant correlations, revealing biologically meaningful functional modules.
-
Full-text index only
Broad network-based predictability of Saccharomyces cerevisiae gene loss-of-function phenotypes.
PMID 18053250 · PMC2246260 · Genome biology · 2007 · 8 claims · 4 setups
Loss-of-function phenotypes in yeast are predictable from a gene's connections in a functional gene network via guilt-by-association.
-
Full-text index only
QuantiSNP: an Objective Bayes Hidden-Markov Model to detect and accurately map copy number variation using SNP genotyping data.
PMID 17341461 · PMC1874617 · Nucleic acids research · 2007 · 8 claims · 7 setups
QuantiSNP (OB-HMM) provides probabilistic quantification of copy number states and significantly improves accuracy of segmental aneuploidy identification and breakpoint mapping relative to existing tools (BeadStudio/Illumina)
-
Full-text index only
Exploiting noise in array CGH data to improve detection of DNA copy number change.
PMID 17272296 · PMC1994778 · Nucleic acids research · 2007 · 7 claims · 4 setups
When aberrations are present, noise in BAC, 19k oligo, and 385k oligo array-CGH data is highly non-Gaussian and shows long-range spatial correlations.
-
Has reproduction · 50
MEDUSA: A Pipeline for Sensitive Taxonomic Classification and Flexible Functional Annotation of Metagenomic Shotgun Sequences.
PMID 35330728 · PMC8940201 · Frontiers in genetics · 2022 · 7 claims · 6 setups
MEDUSA correctly identifies more species than MEGAN 6 CE, especially less abundant species.
-
Full-text index only
Effects of DNA mass on multiple displacement whole genome amplification and genotyping performance.
PMID 16168060 · PMC1249558 · BMC biotechnology · 2005 · 8 claims · 6 setups
Increased gDNA input into the MDA WGA reaction increases the proportion of double-stranded and human-specific PCR-amplifiable wgaDNA and improves genotyping performance.
-
Full-text index only
Quadratic regression analysis for gene discovery and pattern recognition for non-cyclic short time-course microarray experiments.
PMID 15850479 · PMC1127068 · BMC bioinformatics · 2005 · 8 claims · 8 setups
A step-down quadratic regression method (fitting quadratic, then linear, then null models per gene) identifies differentially expressed genes and classifies them into 9 temporal expression patterns using continuous time information.
-
Full-text index only
F-SNP: computationally predicted functional SNPs for disease association studies.
PMID 17986460 · PMC2238878 · Nucleic acids research · 2008 · 6 claims · 8 setups
F-SNP is a database integrating functional effect predictions for SNPs from 16 bioinformatics tools/databases across four categories: splicing, transcription, translation, and post-translation