Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
ProMiR II: a web server for the probabilistic prediction of clustered, nonclustered, conserved and nonconserved microRNAs.
PMID 16845048 · PMC1538778 · Nucleic acids research · 2006 · 6 claims · 4 setups
ProMiR II improves on the original ProMiR by integrating free energy, G/C ratio, conservation score and entropy for more controllable miRNA prediction
-
Full-text index only
Polymorphix: a sequence polymorphism database.
PMID 15608242 · PMC540030 · Nucleic acids research · 2005 · 8 claims · 5 setups
Polymorphix is an ACNUC-structured database that organizes EMBL/GenBank sequences into within-species homologous sequence families using similarity and bibliographic criteria, with alignments, outgroups and phylogenetic trees provided.
-
Has reproduction · 100
Intratumoral heterogeneity in microsatellite instability status at single-cell resolution.
PMID 41767255 · PMC12936829 · iScience · 2026 · 8 claims · 7 setups
A novel computational (Snakemake) pipeline quantifies intratumoral heterogeneity in MSI status at single-cell resolution
-
Full-text index only
Serum protein profile in systemic-onset juvenile idiopathic arthritis differentiates response versus nonresponse to therapy.
PMID 15987476 · PMC1175022 · Arthritis research & therapy · 2005 · 8 claims · 8 setups
SELDI-TOF MS can differentiate serum protein profiles of active versus well-controlled SJIA
-
Full-text index only
Improvements to GALA and dbERGE II: databases featuring genomic sequence alignment, annotation and experimental results.
PMID 15608239 · PMC539999 · Nucleic acids research · 2005 · 8 claims · 8 setups
GALA is now a set of interlinked relational databases covering five vertebrate species: human, chimpanzee, mouse, rat and chicken.
-
Full-text index only
Multiplex sequencing of paired-end ditags (MS-PET): a strategy for the ultra-high-throughput analysis of transcriptomes and genomes.
PMID 16840528 · PMC1524903 · Nucleic acids research · 2006 · 7 claims · 5 setups
MS-PET, which dimerizes PETs prior to 454 multiplex sequencing, achieves an approximate 100-fold efficiency increase over standard Sanger-based PET analysis
-
Full-text index only
EPD in its twentieth year: towards complete promoter coverage of selected model organisms.
PMID 16381980 · PMC1347508 · Nucleic acids research · 2006 · 7 claims · 4 setups
EPD is an annotated, non-redundant collection of experimentally defined eukaryotic POL II promoters accessed via genome position pointers.
-
Full-text index only
Systems biology approach for mapping the response of human urothelial cells to infection by Enterococcus faecalis.
PMID 18047719 · PMC2099488 · BMC bioinformatics · 2007 · 8 claims · 5 setups
Deconvoluting gene expression variance into technical (Gaussian, ~6.5% relative SD) and biological components identifies hypervariable (HV) genes that reflect true biological response to infection without requiring replicates
-
Full-text index only
DAVID Bioinformatics Resources: expanded annotation database and novel algorithms to better extract biology from large gene lists.
PMID 17576678 · PMC1933169 · Nucleic acids research · 2007 · 8 claims · 4 setups
The DAVID Gene Concept uses a single-linkage method to agglomerate tens of millions of gene/protein identifiers from NCBI, PIR, UniProt and other resources into unified DAVID genes.
-
Full-text index only
The TIGR Gene Indices: clustering and assembling EST and known genes and integration with eukaryotic genomes.
PMID 15608288 · PMC540018 · Nucleic acids research · 2005 · 8 claims · 8 setups
The TIGR Gene Indices (TGI) are a collection of 77 species-specific databases that cluster and assemble EST and known gene sequences into tentative consensus (TC) sequences to identify and characterize expressed transcripts.
-
Has reproduction · 57
Genome-wide kinetic properties of transcriptional bursting in mouse embryonic stem cells.
PMID 32596448 · PMC7299619 · Science advances · 2020 · 8 claims · 8 setups
Transcriptional bursting kinetics is regulated by a combination of promoter- and gene body-binding proteins, including the polycomb repressive complex 2 (PRC2) and transcription elongation factors
-
Has reproduction · 90
Transcriptomic data meta-analysis reveals common and injury model specific gene expression changes in the regenerating zebrafish heart.
PMID 37012284 · PMC10070245 · Scientific reports · 2023 · 7 claims · 8 setups
Batch correction using sequencing platform as the correcting variable (via Combat-Seq) removes technical variability so that samples cluster by injury condition rather than dataset origin.
-
Has reproduction · 83
Hierarchical classification-based pan-cancer methylation analysis to classify primary cancer.
PMID 38066424 · PMC10709847 · BMC bioinformatics · 2023 · 8 claims · 8 setups
CHCT, a two-tier hierarchical classification tool built from methylation data, accurately classifies primary cancer type across 30 cancer types.
-
Has reproduction
miRge3.0: a comprehensive microRNA and tRF sequencing analysis pipeline.
PMID 34308351 · PMC8294687 · NAR genomics and bioinformatics · 2021 · 8 claims · 6 setups
miRge3.0 is a Python 3-based small RNA-seq and tRF analysis pipeline that improves on miRge2.0 (which was Python 2.7-based)
-
Has reproduction · 75
Graph-Based Approaches Significantly Improve the Recovery of Antibiotic Resistance Genes From Complex Metagenomic Datasets.
PMID 34690959 · PMC8528159 · Frontiers in microbiology · 2021 · 8 claims · 6 setups
GraphAMR, a Nextflow pipeline that aligns AMR profile HMMs (or AA sequences) to metagenomic assembly graphs via PathRacer, then dereplicates and annotates hits, recovers more and more complete AMR genes than contig-based or read-based methods.
-
Has reproduction · 95
transXpress: a Snakemake pipeline for streamlined de novo transcriptome assembly and annotation.
PMID 37016291 · PMC10074830 · BMC bioinformatics · 2023 · 6 claims · 7 setups
transXpress is a Snakemake pipeline that streamlines de novo transcriptome assembly, quantification, and annotation for non-model organisms
-
Has reproduction · 50
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization.
PMID 41266599 · PMC12635123 · Communications biology · 2025 · 8 claims · 8 setups
BiRNA-BERT uses adaptive dual-tokenization that dynamically selects nucleotide-level (NUC) or byte-pair encoding (BPE) tokens based on input sequence length
-
Full-text index only
ECgene: genome annotation for alternative splicing.
PMID 15608289 · PMC540072 · Nucleic acids research · 2005 · 8 claims · 5 setups
ECgene combines genome-based EST clustering with a graph-theoretic transcript assembly procedure to predict gene models including alternative splicing events.
-
Has reproduction · 83
Macrel: antimicrobial peptide screening in genomes and metagenomes.
PMID 33384902 · PMC7751412 · PeerJ · 2020 · 8 claims · 8 setups
Macrel introduces a novel set of 22 peptide features (6 local, 16 global), including a new Free Energy Transition (FET) feature group, for AMP and hemolytic activity classification
-
Full-text index only
Improvements to cardiovascular gene ontology.
PMID 19046747 · PMC2706316 · Atherosclerosis · 2009 · 8 claims · 8 setups
Gene Ontology (GO) provides a controlled vocabulary that links current functional knowledge of genes to high-throughput genomic and proteomic datasets, aiding data interpretation.