Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
The distribution of SNPs in human gene regulatory regions.
PMID 16209714 · PMC1260019 · BMC genomics · 2005 · 8 claims · 6 setups
SNPs occur with higher density closer to the transcriptional start site within gene promoter regions than in further upstream regions
-
Has reproduction · 67
Generative and integrative modeling for transcriptomics with formalin fixed paraffin embedded material.
PMID 41029822 · PMC12486589 · Journal of translational medicine · 2025 · 8 claims · 5 setups
fRNA-seq transcript counts are best fit by the negative binomial distribution, with little evidence supporting zero-inflated extensions
-
Full-text index only
Identifying the important HIV-1 recombination breakpoints.
PMID 18787691 · PMC2522274 · PLoS computational biology · 2008 · 8 claims · 3 setups
Local sequence identity between co-packaged parental RNAs strongly influences the probability of strand-transfer/breakpoint location, with fewer breakpoints occurring near mismatches
-
Has reproduction · 65
FusionQ: a novel approach for gene fusion detection and quantification from paired-end RNA-Seq.
PMID 23768108 · PMC3691734 · BMC bioinformatics · 2013 · 8 claims · 8 setups
FusionQ is a novel tool that detects gene fusions, constructs chimerical transcript structures, and estimates their abundances from paired-end RNA-Seq data.
-
Has reproduction · 86
RNASEQR--a streamlined and accurate RNA-seq sequence analysis program.
PMID 22199257 · PMC3315322 · Nucleic acids research · 2012 · 8 claims · 7 setups
RNASEQR is a new RNA-seq mapper/aligner that combines a BWT-based (Bowtie) transcriptomic/genomic alignment with hash-based BLAT local alignment in three sequential steps: transcriptome mapping, novel exon detection, and anchor-and-align novel splice junction identification.
-
Has reproduction · 71
Protein structure quality assessment based on the distance profiles of consecutive backbone Cα atoms.
PMID 24555103 · PMC3892923 · F1000Research · 2013 · 8 claims · 8 setups
The distance between consecutive backbone Cα atoms in high-quality structures is normally distributed with mean 3.8 Å and standard deviation 0.04 Å, justifying a reference state in which all consecutive Cα atoms are 3.8 Å apart.
-
Full-text index only
Duplication count distributions in DNA sequences.
PMID 19256873 · PMC3121164 · Physical review. E, Statistical, nonlinear, and soft matter physics · 2008 · 8 claims · 8 setups
Duplication count distributions N(c) for complex 40-mers show power-law-like decay for c roughly 3 to 50 (or higher) across human, C. elegans, A. thaliana, and D. melanogaster genomes.
-
Has reproduction · 68
Bayesian transcriptome assembly.
PMID 25367074 · PMC4397945 · Genome biology · 2014 · 8 claims · 8 setups
Bayesembler, a probabilistic transcriptome assembler built on a Bayesian model of the RNA sequencing process with Gibbs sampling over expressed candidates, abundances and read assignments, is introduced.
-
Has reproduction · 71
Spatial organization shapes the turnover of a bacterial transcriptome.
PMID 27198188 · PMC4874777 · eLife · 2016 · 7 claims · 6 setups
The E. coli transcriptome is spatially organized genome-wide: mRNAs encoding inner-membrane proteins are enriched at the membrane, while mRNAs encoding cytoplasmic, periplasmic and outer-membrane proteins are distributed throughout the cytoplasm.
-
Full-text index only
The DAVID Gene Functional Classification Tool: a novel biological module-centric algorithm to functionally analyze large gene lists.
PMID 17784955 · PMC2375021 · Genome biology · 2007 · 8 claims · 6 setups
Gene-gene functional similarity can be measured using kappa statistics applied to a binary gene-annotation-term matrix built from 14 annotation categories.
-
Has reproduction · 60
A comparative analysis of blastoid models through single-cell transcriptomics.
PMID 39524369 · PMC11543915 · iScience · 2024 · 8 claims · 7 setups
EPSC-derived blastoids are transcriptomically distinct from nPSC-derived blastoids, with nPSC-blastoids clustering closer to natural blastocysts.
-
Full-text index only
A computational screen for type I polyketide synthases in metagenomics shotgun data.
PMID 18953415 · PMC2568958 · PloS one · 2008 · 8 claims · 6 setups
Combining HMM domain searches with maximum-likelihood phylogenetic trees can discriminate true PKS I sequences from evolutionarily related but functionally different enzymes (e.g., FAS I) in metagenomic data.
-
Has reproduction · 65
High-throughput sequencing SELEX for the determination of DNA-binding protein specificities in vitro.
PMID 35776646 · PMC9243297 · STAR protocols · 2022 · 8 claims · 8 setups
HT-SELEX enables unbiased, in vitro determination of preferred DNA target motifs for DNA-binding proteins by iterative selection and PCR amplification of bound oligonucleotides
-
Has reproduction · 86
LMAS: evaluating metagenomic short de novo assembly methods through defined communities.
PMID 36576131 · PMC9795473 · GigaScience · 2022 · 8 claims · 5 setups
LMAS (Last Metagenomic Assembler Standing) is a flexible, Nextflow-based, Docker-containerized automated workflow for benchmarking de novo metagenomic assemblers against defined mock communities, producing an interactive HTML report.
-
Full-text index only
Applications for protein sequence-function evolution data: mRNA/protein expression analysis and coding SNP scoring tools.
PMID 16912992 · PMC1538848 · Nucleic acids research · 2006 · 7 claims · 8 setups
PANTHER HMMs built from family/subfamily multiple sequence alignments can classify novel protein sequences into functional groups based on statistically significant HMM match scores
-
Has reproduction · 58
A comparative study of techniques for differential expression analysis on RNA-Seq data.
PMID 25119138 · PMC4132098 · PloS one · 2014 · 8 claims · 8 setups
edgeR performs slightly better than DESeq and Cuffdiff2 in terms of the ability to uncover true positives.
-
Full-text index only
Sequence variation in G-protein-coupled receptors: analysis of single nucleotide polymorphisms.
PMID 15784611 · PMC1069129 · Nucleic acids research · 2005 · 7 claims · 8 setups
Position-specific phylogenetic features describing evolutionary conservation at a site (e.g. SIFT score, normalized site entropy, residue frequency change) are the best individual discriminators of disease-causing versus neutral GPCR mutations.
-
Has reproduction · 42
The electrostatic profile of consecutive Cβ atoms applied to protein structure quality assessment.
PMID 25506420 · PMC4257144 · F1000Research · 2013 · 8 claims · 8 setups
The EPD between Cβ atoms of consecutive residues provides unique signatures of amino acid pair types and can discriminate native from decoy protein structures.