Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 68
Cell-type annotation with accurate unseen cell-type identification using multiple references.
PMID 37379341 · PMC10335708 · PLoS computational biology · 2023 · 8 claims · 4 setups
mtANN integrates multiple reference datasets and eight gene selection methods via ensemble learning (multiple deep classification models + majority voting) to improve cell-type annotation accuracy
-
Has reproduction · 45
Identifying and classifying trait linked polymorphisms in non-reference species by walking coloured de bruijn graphs.
PMID 23536903 · PMC3607606 · PloS one · 2013 · 8 claims · 9 setups
Bubbleparse detects sequence variants directly from NGS reads without a reference genome, using the coloured de Bruijn graph implementation of Cortex plus a new depth-first bubble-finding module.
-
Full-text index only
Reference based annotation with GeneMapper.
PMID 16600017 · PMC1557983 · Genome biology · 2006 · 7 claims · 6 setups
GeneMapper transfers reference gene annotations to target genomes with higher accuracy than GeneWise and Projector
-
Full-text index only
Pathway projector: web-based zoomable pathway browser using KEGG atlas and Google Maps API.
PMID 19907644 · PMC2770834 · PloS one · 2009 · 8 claims · 6 setups
Existing pathway databases and tools do not satisfy all requirements for a generic, comprehensive pathway browser (integrated maps, data access, mapping/editing, export, installation-free availability).
-
Has reproduction · 91
A reference profile-free deconvolution method to infer cancer cell-intrinsic subtypes and tumor-type-specific stromal profiles.
PMID 32111252 · PMC7049190 · Genome medicine · 2020 · 8 claims · 8 setups
DeClust is a reference profile-free deconvolution method that simultaneously deconvolves bulk tumor expression into cancer, immune, and stromal compartments and clusters samples into cancer cell-intrinsic molecular subtypes, outputting subtype-specific reference profiles for the cohort rather than for individuals.
-
Has reproduction · 87
Ultra-deep sequencing data from a liquid biopsy proficiency study demonstrating analytic validity.
PMID 35418127 · PMC9008010 · Scientific data · 2022 · 6 claims · 5 setups
This dataset is the most comprehensive public-facing dataset of ultra-deep ctDNA sequencing data generated to date
-
Has reproduction · 82
Ultra-deep multi-oncopanel sequencing of benchmarking samples with a wide range of variant allele frequencies.
PMID 35680918 · PMC9184574 · Scientific data · 2022 · 8 claims · 8 setups
Four reference samples (Sample A, Sample B, Sample C, Sample Spike-in/AC5) were developed with large numbers of high-confidence positive and negative small variant positions to serve as known content for oncopanel performance assessment.
-
Has reproduction · 92
Missense variants in human forkhead transcription factors reveal determinants of forkhead DNA bispecificity.
PMID 41124077 · PMC12795473 · Cell reports · 2025 · 6 claims · 5 setups
Non-DNA-contacting residues, especially in the loop between helices 2 and 3 and in wing 2, control mono- vs. bispecificity of FH domains for the FKH and FHL motifs
-
Has reproduction · 51
Evaluation of the Available Variant Calling Tools for Oxford Nanopore Sequencing in Breast Cancer.
PMID 36140751 · PMC9498802 · Genes · 2022 · 7 claims · 6 setups
Clair3 and Human-SNP-wf (which incorporates Clair3) achieved the highest performance among the six variant callers tested.
-
Full-text index only
Allele quantification using molecular inversion probes (MIP).
PMID 16314297 · PMC1301601 · Nucleic acids research · 2005 · 8 claims · 5 setups
MIP technology at high multiplex (>20,000 SNPs) can provide copy number measurements while simultaneously obtaining allele information
-
Full-text index only
Satellog: a database for the identification and prioritization of satellite repeats in disease association studies.
PMID 15949044 · PMC1181805 · BMC bioinformatics · 2005 · 7 claims · 6 setups
Satellog is a database cataloging all pure 1-16 unit satellite repeats in the human genome with supplementary polymorphism, gene-location, and expression data for prioritizing repeats in disease-association studies.
-
Full-text index only
Computational verification of protein-protein interactions by orthologous co-expression.
PMID 15740634 · PMC555590 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Co-expression of orthologous protein pairs across multiple species can verify/predict S. cerevisiae PPIs with better performance than S. cerevisiae co-expression alone.
-
Full-text index only
MitoP2: the mitochondrial proteome database--now including mouse data.
PMID 16381964 · PMC1347489 · Nucleic acids research · 2006 · 8 claims · 8 setups
MitoP2 is a database integrating manually annotated mitochondrial reference proteins, functions, and disease associations for yeast, human, and mouse, with cross-species orthologue mapping
-
Full-text index only
Automated recognition of retroviral sequences in genomic data--RetroTector.
PMID 17636050 · PMC1976444 · Nucleic acids research · 2007 · 8 claims · 8 setups
RetroTector uses 'fragment threading' (detection of chains of conserved retroviral motifs satisfying distance constraints) combined with LTR detection and protein reconstruction to identify ERVs in genomic sequences
-
Full-text index only
Short tandem repeats in human exons: a target for disease mutations.
PMID 18789129 · PMC2543027 · BMC genomics · 2008 · 8 claims · 6 setups
STRs are present in exons of 92% of known human genes, unlike longer tandem repeats which are rare in exons
-
Full-text index only
A single-step sequencing method for the identification of Mycobacterium tuberculosis complex species.
PMID 18618024 · PMC2453075 · PLoS neglected tropical diseases · 2008 · 7 claims · 8 setups
ETR-D sequencing allows accurate, single-step identification of MTC species, circumventing the expensive polyphasic approach.
-
Full-text index only
BFAST: an alignment tool for large scale genome resequencing.
PMID 19907642 · PMC2770639 · PloS one · 2009 · 7 claims · 4 setups
BFAST is a new algorithm and freely available software tool for aligning large-scale short-read sequencing data to large reference genomes with user-customizable speed and accuracy
-
Full-text index only
BioAfrica's HIV-1 proteomics resource: combining protein data with bioinformatics tools.
PMID 15757512 · PMC555852 · Retrovirology · 2005 · 8 claims · 3 setups
BioAfrica's HIV-1 Proteomics Resource integrates protein structure, gene expression, post-translational modification, functional activity and protein-macromolecule interaction data with bioinformatics tools in a single website.
-
Full-text index only
NCBI Reference Sequences: current status, policy and new initiatives.
PMID 18927115 · PMC2686572 · Nucleic acids research · 2009 · 7 claims · 5 setups
RefSeq is a curated, non-redundant, explicitly linked database of nucleotide and protein sequences spanning genomes, transcripts and proteins across prokaryotes, eukaryotes and viruses
-
Has reproduction · 90
CONSULT: accurate contamination removal using locality-sensitive hashing.
PMID 34377979 · PMC8340999 · NAR genomics and bioinformatics · 2021 · 8 claims · 6 setups
CONSULT uses locality-sensitive hashing to test whether query k-mers fall within a user-defined Hamming distance of a reference k-mer database, allowing inexact matching against tens of thousands of microbial species.