Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 50
SMAC, a computational system to link literature, biomedical and expression data.
PMID 31324861 · PMC6642118 · Scientific reports · 2019 · 8 claims · 8 setups
SMAC is a tool that extracts, prioritises, integrates and analyses biomedical and molecular data according to user-defined terms
-
Has reproduction · 71
RNAmountAlign: Efficient software for local, global, semiglobal pairwise and multiple RNA sequence/structure alignment.
PMID 31978147 · PMC6980424 · PloS one · 2020 · 8 claims · 6 setups
RNAmountAlign is the first RNA sequence/structure pairwise alignment algorithm based on incremental ensemble mountain distance, running in O(n^3) time and O(n^2) space for two sequences of length n.
-
Full-text index only
Automatic annotation of eukaryotic genes, pseudogenes and promoters.
PMID 16925832 · PMC1810547 · Genome biology · 2006 · 8 claims · 6 setups
Fgenesh++ gene prediction pipeline identifies 91% of coding nucleotides with 90% specificity
-
Full-text index only
LIMPIC: a computational method for the separation of protein MALDI-TOF-MS signals from noise.
PMID 17386085 · PMC1847688 · BMC bioinformatics · 2007 · 7 claims · 4 setups
LIMPIC is a computational method for detecting protein peaks from linear-mode MALDI-TOF-MS data using background noise reduction and baseline removal followed by non-uniform threshold peak detection and multi-spectra detection-rate classification.
-
Has reproduction · 63
RummaGEO: Automatic mining of human and mouse gene sets from GEO.
PMID 39569206 · PMC11573963 · Patterns (New York, N.Y.) · 2024 · 8 claims · 7 setups
RummaGEO is a gene expression signature search engine built from automatically mined human and mouse RNA-seq perturbation studies in GEO
-
Full-text index only
Analysis of protein sequence and interaction data for candidate disease gene prediction.
PMID 17020920 · PMC1636487 · Nucleic acids research · 2006 · 8 claims · 7 setups
Combining CPS and CMP using known disease genes as input achieves sensitivity 0.52 and specificity 0.97, reducing candidate lists 13-fold
-
Has reproduction · 45
Identifying and classifying trait linked polymorphisms in non-reference species by walking coloured de bruijn graphs.
PMID 23536903 · PMC3607606 · PloS one · 2013 · 8 claims · 9 setups
Bubbleparse detects sequence variants directly from NGS reads without a reference genome, using the coloured de Bruijn graph implementation of Cortex plus a new depth-first bubble-finding module.
-
Full-text index only
Designating eukaryotic orthology via processed transcription units.
PMID 18445630 · PMC2425467 · Nucleic acids research · 2008 · 8 claims · 5 setups
Existing ortholog databases discard/ignore alternative splicing via all-against-all protein comparisons, causing ambiguous ortholog calls and misclassification of AS isoforms as in-paralogs
-
Has reproduction · 97
CellFishing.jl: an ultrafast and scalable cell search method for single-cell RNA sequencing.
PMID 30744683 · PMC6371477 · Genome biology · 2019 · 8 claims · 8 setups
CellFishing.jl searches prebuilt databases for cells with similar expression patterns with high accuracy and throughput using locality-sensitive hashing.
-
Full-text index only
Functional annotation and identification of candidate disease genes by computational analysis of normal tissue gene expression data.
PMID 18560577 · PMC2409962 · PloS one · 2008 · 7 claims · 5 setups
Ranked Coexpression Groups (RCG) built from k=6 nearest coexpressed genes, combined with a majority-rule functional characterization, integrate multiple datasets/coexpression measures to generate high-confidence functional annotation predictions
-
Full-text index only
Evolutionary trace annotation of protein function in the structural proteome.
PMID 20036248 · PMC2831211 · Journal of molecular biology · 2010 · 8 claims · 7 setups
ET-ranked residue clusters can be used to build 3D templates that predict GO function in enzymes and non-enzymes alike, without prior knowledge of functional mechanism.
-
Full-text index only
Decoding of superimposed traces produced by direct sequencing of heterozygous indels.
PMID 18654614 · PMC2429969 · PLoS computational biology · 2008 · 7 claims · 3 setups
A dynamic programming method (implemented as web app Indelligent) can decode superimposed allelic sequences from a single mixed trace, using only the observed string of ambiguous peak calls, without a reference sequence or reverse trace.
-
Has reproduction · 88
Wochenende - modular and flexible alignment-based shotgun metagenome analysis.
PMID 36368923 · PMC9650795 · BMC genomics · 2022 · 8 claims · 6 setups
Wochenende is a modular, transparent alignment-based pipeline for shotgun metagenome analysis supporting short and long reads across all kingdoms of life
-
Has reproduction · 78
A case study for large-scale human microbiome analysis using JCVI's metagenomics reports (METAREP).
PMID 22719821 · PMC3374610 · PloS one · 2012 · 8 claims · 7 setups
METAREP version 1.3.1 is an open-source, scalable tool for querying, browsing and comparing extremely large volumes of metagenomic annotations, with an extended data model, dynamic weighting, distributed searches and advanced clustering.
-
Full-text index only
GENCODE: producing a reference annotation for ENCODE.
PMID 16925838 · PMC1810553 · Genome biology · 2006 · 8 claims · 8 setups
GENCODE annotation combines initial manual annotation by HAVANA, experimental validation, and refinement based on results to identify protein-coding genes in ENCODE regions
-
Has reproduction · 30
Minimal metabolic pathway structure is consistent with associated biomolecular interactions.
PMID 24987116 · PMC4299494 · Molecular systems biology · 2014 · 8 claims · 8 setups
MinSpan, a mixed-integer linear optimization algorithm, computes the shortest, linearly independent pathways (sparsest basis of the null space of the stoichiometric matrix S) for genome-scale metabolic networks, which convex approaches (extreme pathways, elementary flux modes) cannot do at genome scale.