Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Full-text index only
InParanoid 7: new algorithms and tools for eukaryotic orthology analysis.
PMID 19892828 · PMC2808972 · Nucleic acids research · 2010 · 8 claims · 7 setups
InParanoid 7 expands the database by an order of magnitude to 100 species, 1.3 million proteins, and 42.7 million pairwise ortholog groups.
-
Full-text index only
Benchmarking ortholog identification methods using functional genomics data.
PMID 16613613 · PMC1557999 · Genome biology · 2006 · 8 claims · 7 setups
InParanoid is the best overall ortholog identification method for identifying functionally equivalent proteins when sensitivity and selectivity are combined into an overall score.
-
Has reproduction · 63
RummaGEO: Automatic mining of human and mouse gene sets from GEO.
PMID 39569206 · PMC11573963 · Patterns (New York, N.Y.) · 2024 · 8 claims · 7 setups
RummaGEO is a gene expression signature search engine built from automatically mined human and mouse RNA-seq perturbation studies in GEO
-
Full-text index only
Genome-wide analysis of human disease alleles reveals that their locations are correlated in paralogous proteins.
PMID 18989397 · PMC2565504 · PLoS computational biology · 2008 · 7 claims · 5 setups
The locations of sequence variants are correlated between paralogous human proteins more than expected by chance.
-
Full-text index only
Coverage and characteristics of the Affymetrix GeneChip Human Mapping 100K SNP set.
PMID 16680197 · PMC1456318 · PLoS genetics · 2006 · 7 claims · 7 setups
SNPs in the Affymetrix 100K set are undersampled from coding regions (both synonymous and nonsynonymous) and oversampled from regions outside genes, relative to HapMap SNPs
-
Has reproduction · 59
Comparison between short-term stress and long-term adaptive responses reveal common paths to molecular adaptation.
PMID 35243257 · PMC8873613 · iScience · 2022 · 8 claims · 7 setups
Short-term stress and long-term adaptations share common metabolic pathways
-
Full-text index only
Comparative genomics.
PMID 14624258 · PMC261895 · PLoS biology · 2003 · 8 claims · 7 setups
Conserved DNA between species tends to encode shared functional features, while divergent DNA underlies species differences
-
Full-text index only
Identification and analysis of co-occurrence networks with NetCutter.
PMID 18781200 · PMC2526157 · PloS one · 2008 · 8 claims · 4 setups
Random sampling from a complete permutation set of the bipartite graph permits co-occurrence analysis with optimal stringency, and the edge-swapping (ES) model closely approximates this and is the preferred null-model among six tested.
-
Full-text index only
AutoCSA, an algorithm for high throughput DNA sequence variant detection in cancer genomes.
PMID 17485433 · PMC5947781 · Bioinformatics (Oxford, England) · 2007 · 7 claims · 2 setups
AutoCSA is an automated algorithm, extended from the CSA protocol, that detects DNA sequence variants in cancer genomes with minimal manual intervention
-
Has reproduction · 95
Pathway-targeting gene matrix for Drosophila gene set enrichment analysis.
PMID 34710184 · PMC8553153 · PloS one · 2021 · 8 claims · 4 setups
Gene matrix files for GSEA are largely unavailable for Drosophila, limiting pathway-level enrichment analysis in this model organism
-
Has reproduction · 69
A comparison across non-model animals suggests an optimal sequencing depth for de novo transcriptome assembly.
PMID 23496952 · PMC3655071 · BMC genomics · 2013 · 8 claims · 8 setups
Representative de novo transcriptome assemblies are generated with as few as ~20 million reads for single-tissue samples and ~30 million reads for whole animals at the mRNA-coverage level.
-
Full-text index only
Codon usage comparison of novel genes in clinical isolates of Haemophilus influenzae.
PMID 15983137 · PMC1160521 · Nucleic acids research · 2005 · 8 claims · 4 setups
A codon usage similarity statistic (ε, based on squared/absolute differences of codon frequencies with an optimized amino acid usage factor) was developed to compare ORFs against a set of 80 reference genomes.
-
Full-text index only
Inference of transcriptional regulation using gene expression data from the bovine and human genomes.
PMID 17683551 · PMC1978505 · BMC genomics · 2007 · 7 claims · 8 setups
Using human reference promoter sequences is a useful approach for studying gene expression regulation in species with limited or non-existing genomic sequence, such as cattle.
-
Full-text index only
Applying proteomics to the diagnosis and treatment of ALS and related diseases.
PMID 19670321 · PMC2836583 · Muscle & nerve · 2009 · 8 claims · 8 setups
Protein-based biomarkers for ALS/MND require further verification and large-scale validation/qualification studies, including disease mimics, before clinical use
-
Full-text index only
Importance sampling for the infinite sites model.
PMID 18976228 · PMC2832804 · Statistical applications in genetics and molecular biology · 2008 · 7 claims · 2 setups
A new importance sampling proposal distribution for the ISM, derived from a new result on exact sampling from a single segregating site, generally shows greater efficiency than the GT and SD proposals.
-
Has reproduction · 90
Transcriptomic data meta-analysis reveals common and injury model specific gene expression changes in the regenerating zebrafish heart.
PMID 37012284 · PMC10070245 · Scientific reports · 2023 · 7 claims · 8 setups
Batch correction using sequencing platform as the correcting variable (via Combat-Seq) removes technical variability so that samples cluster by injury condition rather than dataset origin.
-
Has reproduction · 87
Enhanced Generalizability of RNA Secondary Structure Prediction via Convolutional Block Attention Network and Ensemble Learning.
PMID 40871599 · PMC12388828 · Molecules (Basel, Switzerland) · 2025 · 8 claims · 8 setups
TrioFold integrates base-pairing clues from thermodynamic- and DL-based methods via ensemble learning and a convolutional block attention mechanism to enhance RSS prediction generalizability.
-
Has reproduction · 60
Core transcriptional signatures of phase change in the migratory locust.
PMID 31292921 · PMC6881432 · Protein & cell · 2019 · 8 claims · 7 setups
PhaseCore genes defined by AC-PCA contribution to phase differentiation predict phase status with >87.5% accuracy
-
Full-text index only
Paircomp, FamilyRelationsII and Cartwheel: tools for interspecific sequence comparison.
PMID 15790396 · PMC1087472 · BMC bioinformatics · 2005 · 8 claims · 7 setups
Paircomp, FamilyRelationsII, and Cartwheel together form an integrated system for comparing, viewing, and managing analyses of BAC-sized (~100 kb) genomic sequence pairs.