Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 67
Satellitome Analysis and Transposable Elements Comparison in Geographically Distant Populations of Spodoptera frugiperda.
PMID 35455012 · PMC9026859 · Life (Basel, Switzerland) · 2022 · 8 claims · 5 setups
Most transposable elements are commonly shared across all eight geographically distant S. frugiperda samples, except Maverick and PIF/Harbinger elements which show divergent repeat copies
-
Full-text index only
Clustering by neurocognition for fine mapping of the schizophrenia susceptibility loci on chromosome 6p.
PMID 19694819 · PMC4286260 · Genes, brain, and behavior · 2009 · 6 claims · 6 setups
A family-based clustering strategy using neurocognitive test scores (CPT, WCST) can identify more homogeneous subgroups of schizophrenia families for genetic association analysis
-
Has reproduction · 50
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization.
PMID 41266599 · PMC12635123 · Communications biology · 2025 · 8 claims · 8 setups
BiRNA-BERT uses adaptive dual-tokenization that dynamically selects nucleotide-level (NUC) or byte-pair encoding (BPE) tokens based on input sequence length
-
Full-text index only
nf-core/proteinfamilies: a scalable pipeline for the generation of protein families.
PMID 41563008 · PMC12950615 · GigaScience · 2026 · 8 claims · 3 setups
nf-core/proteinfamilies is a scalable, parametrizable, open-source Nextflow pipeline that generates new protein families or assigns sequences to existing families using profile HMMs and MSAs
-
Full-text index only
Polymorphix: a sequence polymorphism database.
PMID 15608242 · PMC540030 · Nucleic acids research · 2005 · 8 claims · 5 setups
Polymorphix is an ACNUC-structured database that organizes EMBL/GenBank sequences into within-species homologous sequence families using similarity and bibliographic criteria, with alignments, outgroups and phylogenetic trees provided.
-
Full-text index only
The flexible pocketome engine for structural chemogenomics.
PMID 19727619 · PMC2975493 · Methods in molecular biology (Clifton, N.J.) · 2009 · 8 claims · 8 setups
A comprehensive structural Pocketome combined with ensemble docking enables de novo, structure-based prediction of ligand binding poses and activities for new proteins and new chemical scaffolds.
-
Full-text index only
Coiled-coil protein composition of 22 proteomes--differences and common themes in subcellular infrastructure and traffic control.
PMID 16288662 · PMC1322226 · BMC evolutionary biology · 2005 · 7 claims · 5 setups
Proteins with extended coiled-coil domains (>250 amino acids) are largely absent from bacterial genomes but present in archaea and eukaryotes.
-
Full-text index only
More breast cancer genes?
PMID 11305950 · PMC138680 · Breast cancer research : BCR · 2001 · 8 claims · 7 setups
A new high-risk breast cancer gene termed BRCAX may exist on chromosome 13q, identified via CGH and linkage analysis in Nordic families
-
Full-text index only
Identification of multiple independent horizontal gene transfers into poxviruses using a comparative genomics approach.
PMID 18304319 · PMC2268676 · BMC evolutionary biology · 2008 · 8 claims · 4 setups
Comparative synteny conservation around a horizontally transferred gene (HTgene) can distinguish single versus multiple independent HGT events even without a robust phylogenetic tree.
-
Has reproduction · 84
Evolutionary Genomics of Sex-Related Chromosomes at the Base of the Green Lineage.
PMID 34599324 · PMC8557840 · Genome biology and evolution · 2021 · 8 claims · 6 setups
The divergence of the MT+ and MT- alleles predates speciation events within the Ostreococcus genus
-
Full-text index only
Much ado about nothing: modeling amino acid replacement with predicted protein structures.
PMID 42036821 · PMC13171170 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 7 setups
AFSM was constructed from over 660,000 structural alignments across ~21,000 proteins (297 InterPro families), following the BLOSUM log-odds methodology.
-
Has reproduction · 68
Complete Genome Sequencing of Lactobacillus plantarum ZLP001, a Potential Probiotic That Enhances Intestinal Epithelial Barrier Function and Defense Against Pathogens in Pigs.
PMID 30542296 · PMC6277807 · Frontiers in physiology · 2018 · 8 claims · 8 setups
The complete genome of L. plantarum ZLP001 was sequenced and compared with 18 other L. plantarum genomes to identify genomic features underlying its gut-health-promoting probiotic properties.
-
Has reproduction · 90
Comparative Genomics Provides Insight into the Function of Broad-Host Range Sponge Symbionts.
PMID 34519538 · PMC8546597 · mBio · 2021 · 8 claims · 8 setups
Eleven new genomes were added to the Tethybacterales order and a novel family (Polydorabacteraceae) was identified
-
Full-text index only
The global landscape of sequence diversity.
PMID 17996061 · PMC2258180 · Genome biology · 2007 · 7 claims · 5 setups
Eukaryotic sequence datasets show substantially greater genetic diversity (higher sequence/gene family discovery rates) than bacterial datasets, likely related to differences in modes of genetic inheritance.
-
Full-text index only
Reconstructing Indian population history.
PMID 19779445 · PMC2842210 · Nature · 2009 · 8 claims · 8 setups
Most Indian populations descend from a mixture of two ancient, genetically divergent populations: ANI (close to Middle Easterners, Central Asians, Europeans) and ASI (as distinct from ANI and East Asians as those are from each other).
-
Full-text index only
TEPEAK: A novel method for identifying and characterizing polymorphic transposable elements in non-model species populations.
PMID 41494038 · PMC12788660 · PLoS computational biology · 2026 · 8 claims · 6 setups
TEPEAK identifies and characterizes polymorphic TEs in populations without any prior TE sequence or loci information, using only a chromosome-level reference assembly.
-
Full-text index only
Retentive Network promotes efficient RNA language modeling of long sequences.
PMID 41814064 · PMC13111708 · Communications biology · 2026 · 8 claims · 6 setups
RNAret, a RetNet-based RNA language model with O(n) complexity, achieves training parallelism and low computational overhead while processing long RNA sequences
-
Has reproduction · 93
A comparative study on recombination activity in cattle.
PMID 41942849 · PMC13067647 · Genetics, selection, evolution : GSE · 2026 · 8 claims · 8 setups
Genotype data with high systematic missingness across breeds and arrays can be streamlined and analysed with three complementary recombination-estimation approaches (HMM-based LINKPHASE3, deterministic hsphase, likelihood-based hsrecombi)
-
Has reproduction · 87
Enhanced Generalizability of RNA Secondary Structure Prediction via Convolutional Block Attention Network and Ensemble Learning.
PMID 40871599 · PMC12388828 · Molecules (Basel, Switzerland) · 2025 · 8 claims · 8 setups
TrioFold integrates base-pairing clues from thermodynamic- and DL-based methods via ensemble learning and a convolutional block attention mechanism to enhance RSS prediction generalizability.
-
Full-text index only
The DAVID Gene Functional Classification Tool: a novel biological module-centric algorithm to functionally analyze large gene lists.
PMID 17784955 · PMC2375021 · Genome biology · 2007 · 8 claims · 6 setups
Gene-gene functional similarity can be measured using kappa statistics applied to a binary gene-annotation-term matrix built from 14 annotation categories.