Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Has reproduction · 87
A target enrichment method for gathering phylogenetic information from hundreds of loci: An example from the Compositae.
PMID 25202605 · PMC4103609 · Applications in plant sciences · 2014 · 8 claims · 8 setups
A custom sequence capture probe set (9678 baits targeting 1061 orthologous genes) was designed to enrich COS loci across the Compositae.
-
Has reproduction · 88
Constructing an APOBEC-related gene signature with predictive value in the overall survival and therapeutic sensitivity in lung adenocarcinoma.
PMID 37954334 · PMC10637964 · Heliyon · 2023 · 8 claims · 8 setups
APOBEC family genes are differentially expressed across cancer types and tissues, with APOBEC3B being the most aberrantly upregulated gene in most cancers including LUAD
-
Full-text index only
nf-core/proteinfamilies: a scalable pipeline for the generation of protein families.
PMID 41563008 · PMC12950615 · GigaScience · 2026 · 8 claims · 3 setups
nf-core/proteinfamilies is a scalable, parametrizable, open-source Nextflow pipeline that generates new protein families or assigns sequences to existing families using profile HMMs and MSAs
-
Full-text index only
Clustering by neurocognition for fine mapping of the schizophrenia susceptibility loci on chromosome 6p.
PMID 19694819 · PMC4286260 · Genes, brain, and behavior · 2009 · 6 claims · 6 setups
A family-based clustering strategy using neurocognitive test scores (CPT, WCST) can identify more homogeneous subgroups of schizophrenia families for genetic association analysis
-
Full-text index only
Structural evolution of the protein kinase-like superfamily.
PMID 16244704 · PMC1261164 · PLoS computational biology · 2005 · 8 claims · 5 setups
All kinases in the superfamily share a 'universal core' domain consisting only of the regions required for ATP binding and the phosphotransfer reaction.
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Has reproduction · 83
Gene-expression patterns in peripheral blood classify familial breast cancer susceptibility.
PMID 26538066 · PMC4634735 · BMC medical genomics · 2015 · 8 claims · 7 setups
A multigene expression biomarker from PBMCs accurately classifies familial breast cancer (FBC) status
-
Full-text index only
Genome- and Transcriptome-Wide Characterization of AP2/ERF Transcription Factor Superfamily Reveals Their Relevance in Stylosanthes scabra Vogel Under Water Deficit Stress.
PMID 41515103 · PMC12787715 · Plants (Basel, Switzerland) · 2026 · 8 claims · 8 setups
295 AP2/ERF transcription factor genes were identified and classified in the S. scabra genome
-
Full-text index only
Cancer-wide in silico analyses using differentially expressed genes demonstrate the functions and clinical relevance of JAG, DLL, and NOTCH.
PMID 39074091 · PMC11285958 · PloS one · 2024 · 7 claims · 8 setups
JAG, DLL, and NOTCH family gene/protein expression varies diversely across 15 cancer types relative to normal tissue, sometimes discordant between mRNA and protein levels.
-
Full-text index only
Genome wide survey of G protein-coupled receptors in Tetraodon nigroviridis.
PMID 16022726 · PMC1187884 · BMC evolutionary biology · 2005 · 8 claims · 8 setups
466 Tetraodon GPCRs (Tnig-GPCRs) were identified genome-wide, of which 457 had not been previously reported
-
Full-text index only
Development of a split-toxin CRISPR screening platform to systematically identify regulators of human myoblast fusion.
PMID 41540035 · PMC12808753 · Nature communications · 2026 · 8 claims · 7 setups
A CRISPR screening platform combining human myoblast models, a custom muscle-targeted gRNA library (MyoCRISPR-KO Lib), and a split-toxin selection system enables quantitative enrichment of fusion-defective myocytes.
-
Has reproduction · 88
Comprehensive benchmarking of large language models for RNA secondary structure prediction.
PMID 40205851 · PMC11982019 · Briefings in bioinformatics · 2025 · 7 claims · 4 setups
Existing RNA-LLMs had not previously been evaluated for secondary structure prediction in a unified, fair experimental setup with the same datasets and prediction model.
-
Has reproduction · 87
Enhanced Generalizability of RNA Secondary Structure Prediction via Convolutional Block Attention Network and Ensemble Learning.
PMID 40871599 · PMC12388828 · Molecules (Basel, Switzerland) · 2025 · 8 claims · 8 setups
TrioFold integrates base-pairing clues from thermodynamic- and DL-based methods via ensemble learning and a convolutional block attention mechanism to enhance RSS prediction generalizability.
-
Has reproduction · 92
Telomere-to-telomere reference genome for Panax ginseng highlights the evolution of saponin biosynthesis.
PMID 38883331 · PMC11179851 · Horticulture research · 2024 · 8 claims · 8 setups
A telomere-to-telomere reference genome of P. ginseng was assembled (3.45 Gb, 24 chromosomes, 77266 protein-coding genes)
-
Full-text index only
Coiled-coil protein composition of 22 proteomes--differences and common themes in subcellular infrastructure and traffic control.
PMID 16288662 · PMC1322226 · BMC evolutionary biology · 2005 · 7 claims · 5 setups
Proteins with extended coiled-coil domains (>250 amino acids) are largely absent from bacterial genomes but present in archaea and eukaryotes.
-
Full-text index only
Construction and analysis of tag single nucleotide polymorphism maps for six human-mouse orthologous candidate genes in type 1 diabetes.
PMID 15720714 · PMC551616 · BMC genetics · 2005 · 7 claims · 5 setups
None of the six candidate gene regions showed evidence of association with type 1 diabetes (all multi-locus/single-locus test P values > 0.2)
-
Has reproduction · 42
The electrostatic profile of consecutive Cβ atoms applied to protein structure quality assessment.
PMID 25506420 · PMC4257144 · F1000Research · 2013 · 8 claims · 8 setups
The EPD between Cβ atoms of consecutive residues provides unique signatures of amino acid pair types and can discriminate native from decoy protein structures.
-
Full-text index only
Local combinational variables: an approach used in DNA-binding helix-turn-helix motif prediction with sequence information.
PMID 19651875 · PMC2761287 · Nucleic acids research · 2009 · 8 claims · 7 setups
The LCV approach predicts HTH motifs with 93.29% accuracy, 93.93% sensitivity and 92.66% specificity using only primary sequence information
-
Full-text index only
NGSTroubleFinder: a tool for detection and quantification of contamination and kinship across human NGS data.
PMID 41608734 · PMC12838523 · NAR genomics and bioinformatics · 2026 · 8 claims · 8 setups
NGSTroubleFinder detects cross-sample contamination, sample swaps, kinship, and sex mismatches from BAM/CRAM files without requiring additional variant-calling steps