Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 24
MiGPC: a comprehensive catalog of enzybiotics from environmental metagenomes.
PMID 41888223 · PMC13172421 · Scientific reports · 2026 · 8 claims · 8 setups
MiGPC is the first genome-resolved metagenomic gene and protein catalog specifically targeted to enzybiotics
-
Full-text index only
Processing and population genetic analysis of multigenic datasets with ProSeq3 software.
PMID 19797407 · PMC2778335 · Bioinformatics (Oxford, England) · 2009 · 8 claims · 7 setups
ProSeq3 is a program with a graphic user interface that simplifies preparation and basic population genetic analysis of multigenic DNA polymorphism datasets
-
Has reproduction · 87
Enhanced Generalizability of RNA Secondary Structure Prediction via Convolutional Block Attention Network and Ensemble Learning.
PMID 40871599 · PMC12388828 · Molecules (Basel, Switzerland) · 2025 · 8 claims · 8 setups
TrioFold integrates base-pairing clues from thermodynamic- and DL-based methods via ensemble learning and a convolutional block attention mechanism to enhance RSS prediction generalizability.
-
Full-text index only
Genotyping of genetically monomorphic bacteria: DNA sequencing in Mycobacterium tuberculosis highlights the limitations of current methodologies.
PMID 19915672 · PMC2772813 · PloS one · 2009 · 8 claims · 8 setups
MLSA of 89 genes across 108 global MTBC strains yields a single, highly robust phylogeny with virtually no homoplasy, congruent across parsimony, NJ, ML, and Bayesian methods.
-
Full-text index only
Grammar-based distance in progressive multiple sequence alignment.
PMID 18616828 · PMC2478692 · BMC bioinformatics · 2008 · 7 claims · 3 setups
A grammar-based (LZ complexity) distance metric can be used to determine the order in which sequences are progressively pairwise aligned
-
Has reproduction · 58
MZPAQ: a FASTQ data compression tool.
PMID 31171931 · PMC6547476 · Source code for biology and medicine · 2019 · 8 claims · 4 setups
MZPAQ, a hybrid of MFCompress and ZPAQ, achieves the highest compression ratio compared to all evaluated state-of-the-art and general-purpose tools on all benchmark datasets.
-
Has reproduction · 84
An accurate method for identifying recent recombinants from unaligned sequences.
PMID 35025988 · PMC8963311 · Bioinformatics (Oxford, England) · 2022 · 8 claims · 4 setups
A novel algorithm combining the JHMM (Zilversmit et al. 2013) mosaic representation with a distance-based triple comparison can identify recombinant sequences and their parents from unaligned, gene-length sequences without a reference panel.
-
Has reproduction · 44
Detecting DNA modifications from SMRT sequencing data by modeling sequence context dependence of polymerase kinetic.
PMID 23516341 · PMC3597545 · PLoS computational biology · 2013 · 8 claims · 7 setups
Local sequence context strongly determines position-specific polymerase kinetic rate: roughly 80% of IPD variation is explained by a 10 bp context (7 bases upstream, 2 bases downstream of the incorporation site), saturating at 7 bases upstream.
-
Full-text index only
Evolutionary sequence analysis of complete eukaryote genomes.
PMID 15762985 · PMC1274250 · BMC bioinformatics · 2005 · 8 claims · 6 setups
A conservative genome-comparison method (MIA) identifies panorthologs — strict single-copy 1:1 orthologs containing only species divergences, no paralogy — to minimize errors from gene duplication in evolutionary sequence analysis.
-
Full-text index only
Applications for protein sequence-function evolution data: mRNA/protein expression analysis and coding SNP scoring tools.
PMID 16912992 · PMC1538848 · Nucleic acids research · 2006 · 7 claims · 8 setups
PANTHER HMMs built from family/subfamily multiple sequence alignments can classify novel protein sequences into functional groups based on statistically significant HMM match scores
-
Full-text index only
MODBASE, a database of annotated comparative protein structure models and associated resources.
PMID 18948282 · PMC2686492 · Nucleic acids research · 2009 · 8 claims · 8 setups
MODBASE contains 5,152,695 reliable comparative protein structure models for 1,593,209 unique protein sequences.
-
Full-text index only
Simultaneous analysis of all SNPs in genome-wide and re-sequencing association studies.
PMID 18654633 · PMC2464715 · PLoS genetics · 2008 · 8 claims · 5 setups
A Bayesian-inspired penalised maximum likelihood stochastic search method can simultaneously analyse all SNPs (up to 500K) from a GWA study in a few hours on a desktop workstation
-
Full-text index only
Phylogeographic reconstruction of a bacterial species with high levels of lateral gene transfer.
PMID 19922616 · PMC2784454 · BMC biology · 2009 · 8 claims · 7 setups
The ratio of homologous recombination to mutation in B. pseudomallei is over two times higher than in Streptococcus pneumoniae, the highest yet reported in bacteria.
-
Full-text index only
Sequence variation in G-protein-coupled receptors: analysis of single nucleotide polymorphisms.
PMID 15784611 · PMC1069129 · Nucleic acids research · 2005 · 7 claims · 8 setups
Position-specific phylogenetic features describing evolutionary conservation at a site (e.g. SIFT score, normalized site entropy, residue frequency change) are the best individual discriminators of disease-causing versus neutral GPCR mutations.
-
Has reproduction · 65
FusionQ: a novel approach for gene fusion detection and quantification from paired-end RNA-Seq.
PMID 23768108 · PMC3691734 · BMC bioinformatics · 2013 · 8 claims · 8 setups
FusionQ is a novel tool that detects gene fusions, constructs chimerical transcript structures, and estimates their abundances from paired-end RNA-Seq data.
-
Full-text index only
Genome-wide in silico identification and analysis of cis natural antisense transcripts (cis-NATs) in ten species.
PMID 16849434 · PMC1524920 · Nucleic acids research · 2006 · 8 claims · 7 setups
A fast integrative in silico pipeline combining UniGene mRNA/EST mapping to GoldenPath genomes with CDS, poly(A) signal, poly(A) tail and splicing site evidence can reliably identify cis-NATs genome-wide across multiple species
-
Full-text index only
Inconsistencies in Neanderthal genomic DNA sequences.
PMID 17937503 · PMC2014787 · PLoS genetics · 2007 · 8 claims · 6 setups
The Noonan et al. and Green et al. Neanderthal nuclear DNA datasets yield mutually inconsistent estimates of population split time and Neanderthal admixture proportion when analyzed with the same method
-
Full-text index only
Cataloging coding sequence variations in human genome databases.
PMID 18974781 · PMC2570488 · PloS one · 2008 · 8 claims · 7 setups
A significant proportion of CVs overlap between HGMD and dbSNP (4.36% of HGMD CVs registered in dbSNP; 8.11% of dbSNP CVs registered in HGMD), warranting caution when interpreting phenotypic relevance of concurrent CVs.
-
Full-text index only
Comparative phosphoproteomics reveals evolutionary and functional conservation of phosphorylation across eukaryotes.
PMID 18828897 · PMC2760871 · Genome biology · 2008 · 8 claims · 8 setups
The overlap between phosphoproteomes of six eukaryotes (human, mouse, fly, yeast, plant, zebrafish) is significantly greater than expected by chance.
-
Full-text index only
pSTIING: a 'systems' approach towards integrating signalling pathways, interaction and transcriptional regulatory networks in inflammation and cancer.
PMID 16381926 · PMC1347407 · Nucleic acids research · 2006 · 8 claims · 3 setups
pSTIING is a publicly accessible web-based knowledgebase integrating protein-protein, protein-lipid, protein-small molecule interactions, transcriptional regulatory associations, ligand-receptor-cell type information, and signal transduction modules, with a focus on inflammation, cell migration and cancer.