Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
GENCODE: producing a reference annotation for ENCODE.
PMID 16925838 · PMC1810553 · Genome biology · 2006 · 8 claims · 8 setups
GENCODE annotation combines initial manual annotation by HAVANA, experimental validation, and refinement based on results to identify protein-coding genes in ENCODE regions
-
Full-text index only
Comparative Toxicogenomics Database: a knowledgebase and discovery tool for chemical-gene-disease networks.
PMID 18782832 · PMC2686584 · Nucleic acids research · 2009 · 8 claims · 5 setups
CTD is a manually curated knowledgebase that integrates chemical-gene interactions, chemical-disease relationships, and gene-disease relationships into a chemical-gene-disease triad
-
Full-text index only
MatchMiner: a tool for batch navigation among gene and gene product identifiers.
PMID 12702208 · PMC154578 · Genome biology · 2003 · 8 claims · 3 setups
MatchMiner's LookUp function automates batch translation of an input list of gene identifiers into a matching list of a different identifier type.
-
Full-text index only
InSite: a computational method for identifying protein-protein interaction binding sites on a proteome-wide scale.
PMID 17868464 · PMC2375030 · Genome biology · 2007 · 8 claims · 8 setups
InSite predicts protein-pair-specific binding motifs ('Motif M on protein A binds to protein B') by integrating heterogeneous PPI and motif-motif interaction evidence within a Bayesian network trained by EM
-
Full-text index only
An analysis of human microRNA and disease associations.
PMID 18923704 · PMC2559869 · PloS one · 2008 · 8 claims · 8 setups
MicroRNAs tend to show similar dysfunctional evidence (both up- or both down-regulated) for diseases within the same disease cluster, and different dysfunctional evidence between different disease clusters.
-
Full-text index only
Exogean: a framework for annotating protein-coding genes in eukaryotic genomic DNA.
PMID 16925841 · PMC1810556 · Genome biology · 2006 · 8 claims · 5 setups
Exogean is a framework using directed acyclic coloured multigraphs (DACMs) to represent biological objects (mRNA, ESTs, protein alignments, exons) and iteratively combine them into complex protein-coding transcript models.
-
Full-text index only
Discovery of protein-protein interactions using a combination of linguistic, statistical and graphical information.
PMID 15941473 · PMC1164402 · BMC bioinformatics · 2005 · 8 claims · 5 setups
A combined linguistic+statistical+rule-based method achieves precision 0.61 and recall 0.97 (f=0.74) detecting yeast protein-protein interactions across 12,300 Medline abstracts.
-
Full-text index only
SNPdetector: a software tool for sensitive and accurate SNP detection.
PMID 16261194 · PMC1274293 · PLoS computational biology · 2005 · 7 claims · 7 setups
SNPdetector, which models human visual inspection of sequencing traces, achieves low false positive and false negative rates in automated SNP and mutation detection
-
Full-text index only
Large-scale identification and characterization of alternative splicing variants of human gene transcripts using 56,419 completely sequenced and manually annotated full-length cDNAs.
PMID 16914452 · PMC1557807 · Nucleic acids research · 2006 · 8 claims · 8 setups
Analysis of 56,419 full-length cDNAs identified 6877 alternative splicing genes encoding 18,297 alternative splicing variants made of 37,670 exons.
-
Full-text index only
CGMIM: automated text-mining of Online Mendelian Inheritance in Man (OMIM) to identify genetically-associated cancers and candidate genes.
PMID 15796777 · PMC1274267 · BMC bioinformatics · 2005 · 8 claims · 2 setups
CGMIM is a Perl program that text-mines OMIM entries to identify cancer-gene associations and genetically-related cancer type pairs.
-
Has reproduction · 100
Smart spatial omics (S2-omics) optimizes region of interest selection to capture molecular heterogeneity in diverse tissues.
PMID 41298871 · PMC12662399 · Nature cell biology · 2025 · 7 claims · 6 setups
S2-omics is an end-to-end workflow that automatically selects ROIs from H&E histology images to maximize molecular information content for spatial omics profiling.
-
Full-text index only
Information extraction from full text scientific articles: where are the keywords?
PMID 12775220 · PMC166134 · BMC bioinformatics · 2003 · 8 claims · 5 setups
The keyword content of the five article sections (A, I, M, R, D) is heterogeneous, i.e., different sections carry different kinds of information.
-
Full-text index only
Extraction of human kinase mutations from literature, databases and genotyping studies.
PMID 19758464 · PMC2745582 · BMC bioinformatics · 2009 · 7 claims · 6 setups
A literature mining pipeline combining MutationFinder, false-positive filtering, and SVM-based classification can extract and disambiguate single-point mutation mentions from abstracts and full text
-
Full-text index only
IDEAL-Q, an automated tool for label-free quantitation analysis using an efficient peptide alignment approach and spectral data validation.
PMID 19752006 · PMC2808259 · Molecular & cellular proteomics : MCP · 2010 · 6 claims · 5 setups
IDEAL-Q predicts the elution time of peptides unidentified in a given LC-MS/MS run (but identified in others) using a computation-efficient linear regression plus fragmental refining function, avoiding costly whole-dataset pattern recognition
-
Full-text index only
Multiplexed genetic analysis using an expanded genetic alphabet.
PMID 15319316 · PMC1592527 · Clinical chemistry · 2004 · 7 claims · 6 setups
MultiCode PLx is a three-step platform (PCR, target-specific extension, liquid chip decoding) performed in a single reaction vessel and completed in ~3 h
-
Full-text index only
Systematic identification of pseudogenes through whole genome expression evidence profiling.
PMID 16945953 · PMC1636364 · Nucleic acids research · 2006 · 8 claims · 8 setups
Developed a novel bioinformatics method that identifies pseudogenes by profiling whole-genome transcript and protein expression evidence
-
Full-text index only
FGFR3 protein expression and its relationship to mutation status and prognostic variables in bladder cancer.
PMID 17668422 · PMC2443273 · The Journal of pathology · 2007 · 8 claims · 4 setups
FGFR3 mutations occur in 42% of primary urothelial carcinomas and are significantly associated with low tumour grade and stage
-
Full-text index only
Genetic variation in an individual human exome.
PMID 18704161 · PMC2493042 · PLoS genetics · 2008 · 8 claims · 7 setups
The ~12,500 nonsilent coding variants in the HuRef exome can be reduced ~8-fold to a set of ~1,600 variants most likely to affect protein function.
-
Has reproduction · 64
Blood Transcriptome Analysis of Septic Patients Reveals a Long Non-Coding Alu-RNA in the Complement C5a Receptor 1 Gene.
PMID 35447887 · PMC9027897 · Non-coding RNA · 2022 · 6 claims · 7 setups
A computational pipeline intersecting immune gene coordinates with Alu element coordinates can identify candidate Alu-lncRNAs
-
Has reproduction · 93
Characterization of protein isoform diversity in human umbilical vein endothelial cells via long-read proteogenomics.
PMID 36457147 · PMC9721438 · RNA biology · 2022 · 8 claims · 7 setups
Long-read RNA-seq detected 53,863 transcript isoforms from 10,426 genes in HUVECs, of which 22,195 were novel