Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
CLUES A Comprehensive Workflow for Integrating Geospatial Data in Biomedical Research.
PMID 42128886 · PMC13172076 · Nature communications · 2026 · 8 claims · 5 setups
CLUES is an open-source, end-to-end workflow that automates selection, download, harmonization, and linkage of open-access geospatial environmental data to individual-level biomedical data without requiring geospatial expertise.
-
Full-text index only
Detecting unannotated splicing events in short-read RNA-seq with SAMI, a UMI-aware Nextflow pipeline.
PMID 42166739 · PMC13242923 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 5 setups
SAMI is a UMI-aware, Singularity-contained Nextflow pipeline that detects splicing events diverging from transcript annotations directly from raw FASTQ files.
-
Full-text index only
PolyGenie: a reproducible Nextflow pipeline for phenome-wide association studies using polygenic risk scores.
PMID 42272542 · PMC13247587 · NAR genomics and bioinformatics · 2026 · 7 claims · 6 setups
PolyGenie is an open-source Nextflow pipeline that takes precomputed PRS and phenotype data as input and performs scalable PheWAS analysis across binary and continuous outcomes
-
Full-text index only
A note on generalized Genome Scan Meta-Analysis statistics.
PMID 15717930 · PMC551600 · BMC bioinformatics · 2005 · 7 claims · 3 setups
An Edgeworth series approximation to the null distribution of the weighted GSMA statistic provides a more accurate representation than the normal approximation, especially in the tails
-
Has reproduction · 83
Accurate prediction of metagenome-assembled genome completeness by MAGISTA, a random forest model built on alignment-free intra-bin statistics.
PMID 35248155 · PMC8898458 · Environmental microbiome · 2022 · 7 claims · 7 setups
MAGISTA, a random forest model built on alignment-free intra-bin distance-distribution statistics, can estimate MAG completeness and purity without relying on reference marker genes.
-
Full-text index only
Ensembl 2006.
PMID 16381931 · PMC1347495 · Nucleic acids research · 2006 · 8 claims · 5 setups
Ensembl now provides annotation for 19 genomes, up from 4 the previous year, including new mammalian (Rhesus macaque, Opossum), chordate (Ciona intestinalis), and yeast genomes.
-
Has reproduction · 84
Fractional ridge regression: a fast, interpretable reparameterization of ridge regression.
PMID 33252656 · PMC7702219 · GigaScience · 2020 · 8 claims · 2 setups
Ridge regression can be reparameterized in terms of γ, the ratio between the L2-norms of the regularized and unregularized (OLS) coefficient solutions, defining 'fractional ridge regression' (FRR)
-
Has reproduction · 94
Topological signatures in regulatory network enable phenotypic heterogeneity in small cell lung cancer.
PMID 33729159 · PMC8012062 · eLife · 2021 · 8 claims · 7 setups
The SCLC regulatory network is multistable and its steady states map onto four experimentally observed phenotypes (ASCL1high/NEUROD1low, ASCL1low/NEUROD1high, ASCL1high/NEUROD1high, ASCL1low/NEUROD1low)
-
Has reproduction · 79
Genome-wide prediction of DNase I hypersensitivity using gene expression.
PMID 29051481 · PMC5715040 · Nature communications · 2017 · 8 claims · 5 setups
Gene expression can, to a large extent, predict genome-wide DNase I hypersensitivity (chromatin accessibility)
-
Has reproduction · 71
RNAmountAlign: Efficient software for local, global, semiglobal pairwise and multiple RNA sequence/structure alignment.
PMID 31978147 · PMC6980424 · PloS one · 2020 · 7 claims · 6 setups
RNAmountAlign performs pairwise local, global, and semiglobal (query search) alignment and progressive multiple alignment (global and local) using incremental ensemble mountain height, running in O(n^3) time and O(n^2) space for two sequences of length n
-
Has reproduction · 96
GC-biased gene conversion conceals the prediction of the nearly neutral theory in avian genomes.
PMID 30616647 · PMC6322265 · Genome biology · 2019 · 8 claims · 6 setups
gBGC conceals the correlation between life-history traits and dN/dS in birds; accounting for it reveals correlations consistent with nearly neutral theory
-
Has reproduction · 49
oPOSSUM-3: advanced analysis of regulatory motif over-representation across genes or ChIP-Seq datasets.
PMID 22973536 · PMC3429929 · G3 (Bethesda, Md.) · 2012 · 8 claims · 6 setups
oPOSSUM-3 is a web-accessible system that identifies over-represented TFBS and TFBS families in DNA sequences of co-expressed genes or in sequences from high-throughput methods such as ChIP-Seq.
-
Full-text index only
Distinctive pattern of sequence polymorphism in the NS3 protein of hepatitis C virus type 1b reflects conflicting evolutionary pressures.
PMID 18632963 · PMC2577380 · The Journal of general virology · 2008 · 7 claims · 6 setups
NS3 shows less evidence of purifying selection acting on its CTL epitopes than the other 9 HCV proteins, while outside the CTL epitopes NS3 is more conserved than the other proteins.
-
Has reproduction · 50
Polymorphism identification and improved genome annotation of Brassica rapa through Deep RNA sequencing.
PMID 25122667 · PMC4232532 · G3 (Bethesda, Md.) · 2014 · 8 claims · 8 setups
330,995 SNPs were identified in transcribed regions between B. rapa genotypes R500 and IMB211, at an average frequency of one SNP per 200 bases.
-
Full-text index only
Teasing apart the joint effect of demography and natural selection in the birth of a contact zone.
PMID 36093739 · PMC9828440 · The New phytologist · 2022 · 8 claims · 8 setups
Natural selection contributed to the establishment and maintenance of the Scandinavian contact zone between the NFE and CSE genetic clusters.
-
Full-text index only
DupyliCate: mining, classifying, and characterizing gene duplications.
PMID 42209743 · PMC13219399 · Scientific reports · 2026 · 8 claims · 8 setups
DupyliCate is a Python tool for identifying and classifying gene duplication arrays, using BUSCO-based species-specific thresholds and offering integrated expression divergence and Ka/Ks analysis.
-
Full-text index only
htSNPer1.0: software for haplotype block partition and htSNPs selection.
PMID 15740612 · PMC1274247 · BMC bioinformatics · 2005 · 6 claims · 1 setups
The GBB algorithm finds the globally optimal minimal htSNP set with far less computing time than exhaustive/enumeration search.
-
Full-text index only
Visualization-based discovery and analysis of genomic aberrations in microarray data.
PMID 15953389 · PMC1181623 · BMC bioinformatics · 2005 · 8 claims · 7 setups
ChARMView integrates dynamic visualization with automated statistical analysis (EM-based breakpoint detection, one-sample sign test, permutation mean test) to discover chromosomal aberrations from array CGH and gene expression data
-
Full-text index only
The use of edge-betweenness clustering to investigate biological function in protein interaction networks.
PMID 15740614 · PMC555937 · BMC bioinformatics · 2005 · 8 claims · 7 setups
Edge-Betweenness clustering separates protein interaction graphs into subgraphs whose GO term distributions show significant correlations, revealing biologically meaningful functional modules.
-
Full-text index only
Identifying drug effects via pathway alterations using an integer linear programming optimization formulation on phosphoproteomic data.
PMID 19997482 · PMC2776985 · PLoS computational biology · 2009 · 7 claims · 4 setups
An ILP formulation of the Boolean pathway optimization problem fits phosphoproteomic data faster and more efficiently than the previously used genetic algorithm (GA) approach.