Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
A large-scale proteomic analysis of human embryonic stem cells.
PMID 18162134 · PMC2211323 · BMC genomics · 2007 · 8 claims · 5 setups
Two large-scale western blot systems (PowerBlot, Kinexus) identify over 600 proteins expressed in undifferentiated hESCs across 18 functional classes
-
Full-text index only
ORFer--retrieval of protein sequences and open reading frames from GenBank and storage into relational databases or text files.
PMID 12493080 · PMC139979 · BMC bioinformatics · 2002 · 6 claims · 6 setups
ORFer retrieves protein and nucleic acid sequences and annotations from NCBI GenBank using the XML sequence format
-
Full-text index only
High-accuracy proteome maps of human body fluids.
PMID 17140426 · PMC1794581 · Genome biology · 2006 · 8 claims · 5 setups
Large-scale, high-accuracy MS analyses of tear fluid, urine, and seminal plasma provide high-quality datasets useful for biomarker discovery
-
Full-text index only
Predicting failure rate of PCR in large genomes.
PMID 18492719 · PMC2441781 · Nucleic acids research · 2008 · 7 claims · 8 setups
The number of predicted primer-binding sites in genomic DNA is the most important factor determining PCR failure.
-
Full-text index only
The VIZIER project: preparedness against pathogenic RNA viruses.
PMID 18083241 · PMC7114271 · Antiviral research · 2008 · 8 claims · 6 setups
Almost all newly emerging human pathogenic viruses are RNA viruses, largely because their error-prone RNA-dependent RNA polymerases and zoonotic reservoirs allow rapid adaptation to new hosts.
-
Full-text index only
Genome sequences and great expectations.
PMID 11178275 · PMC150431 · Genome biology · 2001 · 8 claims · 3 setups
Function is known or can be predicted for an average of 62% of proteins across 31 analyzed genomes.
-
Full-text index only
Complex genetic diseases: controversy over the Croesus code.
PMID 11532206 · PMC138948 · Genome biology · 2001 · 8 claims · 3 setups
The common disease/common variant hypothesis is predicted by population genetic theory (founder population dynamics, mutation-drift-selection balance) and supported by empirical examples such as APOE*E4.
-
Full-text index only
Human genome research in China.
PMID 15168679 · PMC7079922 · Journal of molecular medicine (Berlin, Germany) · 2004 · 8 claims · 8 setups
China completed its assigned 1% share of the international Human Genome Project sequencing effort and contributed ~10% of the HapMap effort
-
Has reproduction · 84
Fractional ridge regression: a fast, interpretable reparameterization of ridge regression.
PMID 33252656 · PMC7702219 · GigaScience · 2020 · 7 claims · 2 setups
Ridge regression can be reparameterized in terms of the fraction γ between the L2-norms of the regularized and unregularized coefficient solutions (fractional ridge regression, FRR).
-
Full-text index only
Characterization of large rearrangements in autosomal dominant polycystic kidney disease and the PKD1/TSC2 contiguous gene syndrome.
PMID 18818683 · PMC2756756 · Kidney international · 2008 · 8 claims · 8 setups
Developed an MLPA assay with PKD1 exon 1-33 probes designed at single base-pair mismatches with the six PKD1 pseudogenes to achieve locus specificity
-
Has reproduction · 88
Comprehensive benchmarking of large language models for RNA secondary structure prediction.
PMID 40205851 · PMC11982019 · Briefings in bioinformatics · 2025 · 7 claims · 4 setups
Existing RNA-LLMs had not previously been evaluated for secondary structure prediction in a unified, fair experimental setup with the same datasets and prediction model.
-
Has reproduction · 49
oPOSSUM-3: advanced analysis of regulatory motif over-representation across genes or ChIP-Seq datasets.
PMID 22973536 · PMC3429929 · G3 (Bethesda, Md.) · 2012 · 8 claims · 6 setups
oPOSSUM-3 is a web-accessible system that identifies over-represented TFBS and TFBS families in DNA sequences of co-expressed genes or in sequences from high-throughput methods such as ChIP-Seq.
-
Full-text index only
Systems biology approaches for the study of multiple sclerosis.
PMID 18505469 · PMC3865652 · Journal of cellular and molecular medicine · 2008 · 8 claims · 8 setups
The MHC locus on chromosome 6p21 is the strongest genetic region linked to MS susceptibility.
-
Full-text index only
Gene function in the mammalian genome, courtesy of the mouse.
PMID 12537544 · PMC151280 · Genome biology · 2003 · 8 claims · 8 setups
Mosaicism of Mus musculus domesticus and Mus musculus musculus haplotypes exists across the inbred laboratory mouse genome, and genome-wide haplotype mapping can enhance positional cloning
-
Full-text index only
Similarities and differences in genome-wide expression data of six organisms.
PMID 14737187 · PMC300882 · PLoS biology · 2004 · 8 claims · 8 setups
Coexpression of functionally related genes is frequently conserved across evolutionarily distant organisms
-
Full-text index only
High-throughput genomic technology in research and clinical management of breast cancer. Evolving landscape of genetic epidemiological studies.
PMID 16834767 · PMC1557740 · Breast cancer research : BCR · 2006 · 8 claims · 5 setups
Candidate polymorphism-based genetic epidemiological studies have yielded little success in identifying low-penetrance breast cancer susceptibility genes, largely due to poor genomic coverage and inadequate statistical power.
-
Full-text index only
G2Cdb: the Genes to Cognition database.
PMID 18984621 · PMC2686544 · Nucleic acids research · 2009 · 7 claims · 7 setups
G2Cdb integrates experimentally validated synapse proteome datasets with mouse/human genomic annotation, phenotype, and human disease data in a gene-centric database.
-
Full-text index only
FGF: a web tool for Fishing Gene Family in a whole genome database.
PMID 17584790 · PMC1933194 · Nucleic acids research · 2007 · 6 claims · 3 setups
FGF efficiently searches and identifies gene families in whole-genome databases and outputs visual phylogenetic trees annotated with gene structure, chromosome position, duplication fate, and selective pressure (Ka/Ks)
-
Has reproduction · 65
Wireless sensor network design with reliable and long network lifetime.
PMID 41981015 · PMC13083951 · Scientific reports · 2026 · 6 claims · 3 setups
All four fundamental WSN design problems and network reliability can be addressed together in an integrated set of mixed-integer mathematical models.
-
Has reproduction · 84
Pharokka: a fast scalable bacteriophage annotation tool.
PMID 36453861 · PMC9805569 · Bioinformatics (Oxford, England) · 2023 · 8 claims · 5 setups
Pharokka is a one-line, fast, scalable bacteriophage annotation tool producing standards-compliant outputs, installable via a two-line bioconda command