Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Protein coding potential of retroviruses and other transposable elements in vertebrate genomes.
PMID 15716312 · PMC549403 · Nucleic acids research · 2005 · 8 claims · 5 setups
About 1000 genes across four vertebrate gene sets analyzed contain at least one RETRA marker protein domain
-
Full-text index only
Separating selection from mutation in antibody language models.
PMID 41944291 · PMC13056363 · eLife · 2026 · 8 claims · 6 setups
Masked antibody language models such as AbLang2 are biased by nucleotide-level mutation processes (germline memorization, codon table, SHM rate variation)
-
Full-text index only
Predicting failure rate of PCR in large genomes.
PMID 18492719 · PMC2441781 · Nucleic acids research · 2008 · 7 claims · 8 setups
The number of predicted primer-binding sites in genomic DNA is the most important factor determining PCR failure.
-
Full-text index only
pmid-41592570
PMID 41592570 · PMC13069865 · 8 claims · 7 setups
ChromBERT is pre-trained via masked reconstruction on the Cistrome-Human-6K dataset (6,391 cistromes, 991 transcription regulators) to learn genome-wide interaction syntax of transcription regulators
-
Has reproduction · 95
Mouse-Geneformer: A deep learning model for mouse single-cell transcriptome and its cross-species utility.
PMID 40106407 · PMC11964219 · PLoS genetics · 2025 · 7 claims · 6 setups
Mouse-Geneformer, a Transformer Encoder model pre-trained via masked-token self-supervised learning on mouse-Genecorpus-20M, was successfully constructed following the original human Geneformer architecture.
-
Full-text index only
Evaluating the performance of ancient DNA genetic relatedness estimation methods using high-fidelity pedigree simulations.
PMID 41796349 · PMC13081257 · Genome biology · 2026 · 8 claims · 5 setups
BADGER, an automated snakemake pipeline, was developed to simulate high-fidelity pedigrees and raw ancient DNA sequence data for benchmarking genetic relatedness methods
-
Full-text index only
The EH1 motif in metazoan transcription factors.
PMID 16309560 · PMC1310626 · BMC genomics · 2005 · 8 claims · 5 setups
There is a statistically significant association between EH1hox motif HMM score and transcription factor function across human, Drosophila and C. elegans proteomes.
-
Full-text index only
Pegasys: software for executing and integrating analyses of biological sequences.
PMID 15096276 · PMC406494 · BMC bioinformatics · 2004 · 8 claims · 7 setups
Pegasys is a flexible, modular, customizable software system for executing and integrating heterogeneous biological sequence analysis tools
-
Has reproduction · 92
Chromosome-scale genome sequencing, assembly and annotation of six genomes from subfamily Leishmaniinae.
PMID 34489462 · PMC8421402 · Scientific data · 2021 · 8 claims · 8 setups
Chromosome-scale genomes of six Leishmaniinae species (five L. (Mundinia) species and one Porcisia species) were sequenced, assembled and annotated, providing genome, proteome, transcriptome and GFF outputs for taxa previously lacking public reference genomes
-
Full-text index only
The Vertebrate Genome Annotation (Vega) database.
PMID 15608237 · PMC540089 · Nucleic acids research · 2005 · 8 claims · 8 setups
Vega is a community database for browsing manual annotation of finished vertebrate genome sequences, based on an extended Ensembl-style schema.
-
Full-text index only
The Chromosome-Scale Genome Assembly of the Redlip Blenny, Ophioblennius macclurei (Blenniidae).
PMID 41378738 · PMC12758960 · Genome biology and evolution · 2026 · 8 claims · 12 setups
A chromosome-scale genome assembly of O. macclurei was generated (529.6 Mb, scaffold N50 23.7 Mb, GC 43.49%) using ONT long reads, Illumina short reads, and Hi-C scaffolding.