Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Direct inference of SNP heterozygosity rates and resolution of LOH detection.
PMID 18052545 · PMC2098867 · PLoS computational biology · 2007 · 6 claims · 7 setups
A large proportion of SNPs in dbSNP have high-variance HET rate estimates, limiting their reliability for LOH study design.
-
Has reproduction · 71
Protein structure quality assessment based on the distance profiles of consecutive backbone Cα atoms.
PMID 24555103 · PMC3892923 · F1000Research · 2013 · 8 claims · 8 setups
The distance between consecutive backbone Cα atoms in high-quality structures is normally distributed with mean 3.8 Å and standard deviation 0.04 Å, justifying a reference state in which all consecutive Cα atoms are 3.8 Å apart.
-
Has reproduction · 42
The electrostatic profile of consecutive Cβ atoms applied to protein structure quality assessment.
PMID 25506420 · PMC4257144 · F1000Research · 2013 · 8 claims · 8 setups
The EPD between Cβ atoms of consecutive residues provides unique signatures of amino acid pair types and can discriminate native from decoy protein structures.
-
Has reproduction · 83
Accurate prediction of metagenome-assembled genome completeness by MAGISTA, a random forest model built on alignment-free intra-bin statistics.
PMID 35248155 · PMC8898458 · Environmental microbiome · 2022 · 7 claims · 7 setups
MAGISTA, a random forest model built on alignment-free intra-bin distance-distribution statistics, can estimate MAG completeness and purity without relying on reference marker genes.
-
Has reproduction · 67
Generative and integrative modeling for transcriptomics with formalin fixed paraffin embedded material.
PMID 41029822 · PMC12486589 · Journal of translational medicine · 2025 · 8 claims · 5 setups
fRNA-seq transcript counts are best fit by the negative binomial distribution, with little evidence supporting zero-inflated extensions
-
Full-text index only
Calculating expected DNA remnants from ancient founding events in human population genetics.
PMID 18928554 · PMC2588638 · BMC genetics · 2008 · 8 claims · 3 setups
Genetic parameters (native/migrant population size, mutation rate, generations since admixture) strongly determine the final frequency of migrant alleles detectable today.
-
Full-text index only
BTW: a web server for Boltzmann time warping of gene expression time series.
PMID 16845055 · PMC1538860 · Nucleic acids research · 2006 · 5 claims · 4 setups
Symmetric time warping distance is more flexible than Euclidean distance or correlation coefficient for identifying genes with similar temporal expression profiles, especially across sequences of different length.
-
Full-text index only
Using comparative genomics to reorder the human genome sequence into a virtual sheep genome.
PMID 17663790 · PMC2323240 · Genome biology · 2007 · 8 claims · 6 setups
A sheep BAC library (CHORI-243) with ~13.5-fold genome coverage was constructed and end-sequenced.
-
Full-text index only
Simple models of genomic variation in human SNP density.
PMID 17553150 · PMC1919371 · BMC genomics · 2007 · 6 claims · 4 setups
Hierarchical Poisson model B, which allows both the mutation-rate proxy (Beta-distributed Λ) and the ARG-size proxy (Gamma-distributed T) to vary, fits the observed SNP density distribution significantly better than models with only one or neither varying.
-
Has reproduction · 82
Reusable building blocks in biological systems.
PMID 30958230 · PMC6303794 · Journal of the Royal Society, Interface · 2018 · 8 claims · 4 setups
Biological systems can be decomposed into phenotypic building blocks (PBBs) via k-maximally reusable decompositions (k-MRD) that maximize average reusability across conditions.
-
Has reproduction · 50
Exploiting convergent phenotypes to derive a pan-cancer cisplatin response gene expression signature.
PMID 37076665 · PMC10115855 · NPJ precision oncology · 2023 · 8 claims · 8 setups
A convergent-phenotype-based seed gene/co-expression method can extract consensus gene expression signatures predictive of response to chemotherapeutic drugs in the GDSC database
-
Full-text index only
A space-efficient and accurate method for mapping and aligning cDNA sequences onto genomic sequence.
PMID 18344523 · PMC2377433 · Nucleic acids research · 2008 · 7 claims · 6 setups
Spaln maps and aligns large cDNA sequence sets onto whole mammalian genomes using substantially less memory than comparable existing tools
-
Full-text index only
Variation in genetic admixture and population structure among Latinos: the Los Angeles Latino eye study (LALES).
PMID 19903357 · PMC3087512 · BMC genetics · 2009 · 7 claims · 6 setups
LALES Latinos show strong evidence of recent population admixture, primarily from Native American and European ancestries with smaller Asian and African contributions.
-
Full-text index only
A statistical model to identify differentially expressed proteins in 2D PAGE gels.
PMID 19763172 · PMC2734266 · PLoS computational biology · 2009 · 7 claims · 5 setups
A mixture likelihood model incorporating both detected and non-detected proteins has higher statistical power to detect differential expression than standard approaches like the Student's t-test.
-
Has reproduction · 44
Detecting DNA modifications from SMRT sequencing data by modeling sequence context dependence of polymerase kinetic.
PMID 23516341 · PMC3597545 · PLoS computational biology · 2013 · 8 claims · 7 setups
Local sequence context strongly determines position-specific polymerase kinetic rate: roughly 80% of IPD variation is explained by a 10 bp context (7 bases upstream, 2 bases downstream of the incorporation site), saturating at 7 bases upstream.
-
Full-text index only
Synonymous substitution rates predict HIV disease progression as a result of underlying replication dynamics.
PMID 17305421 · PMC1797821 · PLoS computational biology · 2007 · 8 claims · 8 setups
The synonymous substitution rate (dS) of HIV env is strongly correlated with disease progression parameters (progression time, CD4+ decline rate, viral load increase rate), unlike the nonsynonymous rate (dN).
-
Has reproduction · 72
Analysis of the genome of the New Zealand giant collembolan (Holacanthella duospinosa) sheds light on hexapod evolution.
PMID 29041914 · PMC5644144 · BMC genomics · 2017 · 8 claims · 8 setups
A high-quality ~375 Mbp draft genome and transcriptome of Holacanthella duospinosa was assembled and annotated, providing a genomic resource for hexapod evolution.
-
Has reproduction · 100
Betacoronavirus-specific alternate splicing.
PMID 35074468 · PMC8782732 · Genomics · 2022 · 8 claims · 8 setups
Genes showing differential alternative splicing in SARS-CoV-2 have a similar functional profile to those in SARS-CoV and MERS, affecting a diverse set of genes and biological functions related to virus biology.
-
Full-text index only
Evolutionary cores of domain co-occurrence networks.
PMID 15788102 · PMC1079808 · BMC evolutionary biology · 2005 · 8 claims · 4 setups
The innermost (globally central) cores of protein domain co-occurrence networks gradually grow in size with increasing evolutionary/developmental complexity of the organism.
-
Has reproduction · 48
Rbfox2 controls autoregulation in RNA-binding protein networks.
PMID 24637117 · PMC3967051 · Genes & development · 2014 · 8 claims · 8 setups
Rbfox2 cross-regulates AS-NMD events within RNA-binding protein genes to alter their expression, tuning autoregulatory splicing networks and placing Rbfox2 at a critical node of a multilayer regulatory network.