Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 95
Mouse-Geneformer: A deep learning model for mouse single-cell transcriptome and its cross-species utility.
PMID 40106407 · PMC11964219 · PLoS genetics · 2025 · 7 claims · 6 setups
Mouse-Geneformer, a Transformer Encoder model pre-trained via masked-token self-supervised learning on mouse-Genecorpus-20M, was successfully constructed following the original human Geneformer architecture.
-
Full-text index only
Separating selection from mutation in antibody language models.
PMID 41944291 · PMC13056363 · eLife · 2026 · 8 claims · 6 setups
Masked antibody language models such as AbLang2 are biased by nucleotide-level mutation processes (germline memorization, codon table, SHM rate variation)
-
Full-text index only
Advancing codon language modeling with synonymous codon constrained masking.
PMID 41736545 · PMC12956333 · Nucleic acids research · 2026 · 8 claims · 7 setups
SynCodonLM introduces synonymous codon-constrained masking, restricting masked-codon prediction to only synonymous codon options via logit masking before softmax
-
Full-text index only
scLong: a billion-parameter foundation model for capturing long-range gene context in single-cell transcriptomics.
PMID 41639087 · PMC12982784 · Nature communications · 2026 · 7 claims · 4 setups
scLong performs self-attention across all ~27,874 human genes, including lowly expressed ones, to capture long-range gene dependencies missed by models restricted to highly expressed gene subsets
-
Full-text index only
Retentive Network promotes efficient RNA language modeling of long sequences.
PMID 41814064 · PMC13111708 · Communications biology · 2026 · 8 claims · 6 setups
RNAret, a RetNet-based RNA language model with O(n) complexity, achieves training parallelism and low computational overhead while processing long RNA sequences
-
Has reproduction · 88
Comprehensive benchmarking of large language models for RNA secondary structure prediction.
PMID 40205851 · PMC11982019 · Briefings in bioinformatics · 2025 · 7 claims · 4 setups
Existing RNA-LLMs had not previously been evaluated for secondary structure prediction in a unified, fair experimental setup with the same datasets and prediction model.
-
Full-text index only
Pegasys: software for executing and integrating analyses of biological sequences.
PMID 15096276 · PMC406494 · BMC bioinformatics · 2004 · 8 claims · 7 setups
Pegasys is a flexible, modular, customizable software system for executing and integrating heterogeneous biological sequence analysis tools
-
Full-text index only
Using ESTs to improve the accuracy of de novo gene prediction.
PMID 16817966 · PMC1534067 · BMC bioinformatics · 2006 · 8 claims · 8 setups
TWINSCAN_EST combines EST alignments with TWINSCAN via a trainable 'ESTseq' representation and improves exact gene structure prediction accuracy on the whole C. elegans genome
-
Full-text index only
CLAMP: predicting specific protein-mediated chromatin loops in diverse species with a chromatin accessibility language model.
PMID 41555433 · PMC12903630 · Genome biology · 2026 · 8 claims · 8 setups
CLAMP, a chromatin-accessibility language model, predicts protein-mediated chromatin loops across 10 species, 18 proteins, and 24 cell types with superior performance versus existing methods.
-
Has reproduction · 50
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization.
PMID 41266599 · PMC12635123 · Communications biology · 2025 · 8 claims · 8 setups
BiRNA-BERT uses adaptive dual-tokenization that dynamically selects nucleotide-level (NUC) or byte-pair encoding (BPE) tokens based on input sequence length
-
Full-text index only
Eukan: a fully automated nuclear genome annotation pipeline for less studied and divergent eukaryotes.
PMID 41567515 · PMC12817076 · NAR genomics and bioinformatics · 2026 · 8 claims · 7 setups
Eukan automatically leverages RNA-Seq coverage to inform generalized Hidden Markov Model gene prediction and intron lengths to inform protein sequence alignments
-
Full-text index only
Predicting failure rate of PCR in large genomes.
PMID 18492719 · PMC2441781 · Nucleic acids research · 2008 · 7 claims · 8 setups
The number of predicted primer-binding sites in genomic DNA is the most important factor determining PCR failure.
-
Full-text index only
The genome of the reef-building coral Porites harrisoni from the southern Persian/Arabian Gulf.
PMID 41929243 · PMC13040556 · GigaByte (Hong Kong, China) · 2026 · 8 claims · 8 setups
P. harrisoni from the southern PAG has a genome assembly of 626.7 Mb across 1,883 contigs with contig N50 of 807.4 kb
-
Full-text index only
The Vertebrate Genome Annotation (Vega) database.
PMID 15608237 · PMC540089 · Nucleic acids research · 2005 · 8 claims · 8 setups
Vega is a community database for browsing manual annotation of finished vertebrate genome sequences, based on an extended Ensembl-style schema.
-
Full-text index only
Highly contiguous chromosome-level assembly of the rock goby (Gobius paganellus) genome.
PMID 41611730 · PMC12957465 · Scientific data · 2026 · 8 claims · 8 setups
Chromosome-level genome assembly of Gobius paganellus spans 813 Mb with >99.9% of sequence anchored to 23 pseudochromosomes
-
Full-text index only
First genome assemblies of Neotropical Thoracobombus bumblebees Bombus pauloensis and Bombus pullatus.
PMID 41436027 · PMC12958814 · G3 (Bethesda, Md.) · 2026 · 7 claims · 8 setups
This study produced the first genome assemblies of Neotropical Bombus (Thoracobombus) species, B. pauloensis and B. pullatus
-
Full-text index only
An Improved Chromosome-Level Genome Assembly and Comprehensive Annotation of the Model Ascidian Ciona savignyi.
PMID 41792167 · PMC13087180 · Scientific data · 2026 · 8 claims · 8 setups
An improved chromosome-level genome assembly of C. savignyi was generated using Illumina short reads, ONT long reads, and Hi-C data.
-
Has reproduction · 92
Chromosome-scale genome sequencing, assembly and annotation of six genomes from subfamily Leishmaniinae.
PMID 34489462 · PMC8421402 · Scientific data · 2021 · 8 claims · 8 setups
Chromosome-scale genomes of six Leishmaniinae species (five L. (Mundinia) species and one Porcisia species) were sequenced, assembled and annotated, providing genome, proteome, transcriptome and GFF outputs for taxa previously lacking public reference genomes