Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 67
binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets.
PMID 36239393 · PMC9677464 · Briefings in bioinformatics · 2022 · 8 claims · 8 setups
binny outperforms or is highly competitive with commonly used and state-of-the-art binning methods (MetaBAT2, MaxBin2, CONCOCT, VAMB, SemiBin, MetaDecoder)
-
Has reproduction · 67
Leveraging RNA-seq deconvolution to improve complex in vitro model characterization.
PMID 40701251 · PMC12391696 · The Journal of biological chemistry · 2025 · 8 claims · 6 setups
RNA-seq deconvolution can predict cell type proportions from bulk RNA-seq using scRNA-seq references, offering a useful characterization tool for CIVMs where single-cell methods are impractical
-
Full-text index only
Uncovering information on expression of natural antisense transcripts in Affymetrix MOE430 datasets.
PMID 17598913 · PMC1929078 · BMC genomics · 2007 · 8 claims · 4 setups
Standard Affymetrix expression GeneChips (MOE430, HG-U133) contain probe sets that detect natural antisense transcripts (NATs)
-
Full-text index only
Functional annotation and identification of candidate disease genes by computational analysis of normal tissue gene expression data.
PMID 18560577 · PMC2409962 · PloS one · 2008 · 7 claims · 5 setups
Ranked Coexpression Groups (RCG) built from k=6 nearest coexpressed genes, combined with a majority-rule functional characterization, integrate multiple datasets/coexpression measures to generate high-confidence functional annotation predictions
-
Full-text index only
A non-parametric meta-analysis approach for combining independent microarray datasets: application using two microarray datasets pertaining to chronic allograft nephropathy.
PMID 18302764 · PMC2276496 · BMC genomics · 2008 · 8 claims · 6 setups
A novel non-parametric meta-analysis approach for combining independent microarray datasets is presented, requiring no distributional assumptions and being logically intuitive.
-
Full-text index only
GOLD.db: genomics of lipid-associated disorders database.
PMID 15588328 · PMC544894 · BMC genomics · 2004 · 8 claims · 4 setups
GOLD.db integrates annotated pathways, gene/protein reference information, and curated gene expression datasets for lipid-associated disorders research
-
Has reproduction · 83
Analyzing biomarker discovery: Estimating the reproducibility of biomarker sets.
PMID 35901020 · PMC9333302 · PloS one · 2022 · 7 claims · 3 setups
A Reproducibility Score, RS(D,BD), defined as the average Jaccard overlap between biomarker sets found by the same discovery process on comparable datasets from the same distribution, quantifies biomarker reproducibility on a 0-1 scale
-
Full-text index only
Visualization of three-way comparisons of omics data.
PMID 17335588 · PMC1831488 · BMC bioinformatics · 2007 · 7 claims · 3 setups
A novel HSB (hue, saturation, brightness) color-coding scheme can represent three-way comparisons of corresponding datapoints from three datasets.
-
Has reproduction · 90
Developmental hematopoietic stem cell variation explains clonal hematopoiesis later in life.
PMID 39592593 · PMC11599844 · Nature communications · 2024 · 8 claims · 2 setups
Weak selection conferred by HSC variation created before birth can reliably yield clonal hematopoiesis later in life, demonstrated via shared prenatal circulation of monozygotic (MZ) twins.
-
Has reproduction · 81
SEMdag: Fast learning of Directed Acyclic Graphs via node or layer ordering.
PMID 39775401 · PMC11709272 · PloS one · 2025 · 8 claims · 5 setups
SEMdag() is a two-step order-based algorithm for fast learning of high-dimensional linear SEMs, using knowledge-based (KB) or data-driven bottom-up (BU) node/layer ordering followed by penalized (L1) DAG estimation
-
Has reproduction · 97
CellFishing.jl: an ultrafast and scalable cell search method for single-cell RNA sequencing.
PMID 30744683 · PMC6371477 · Genome biology · 2019 · 8 claims · 8 setups
CellFishing.jl searches prebuilt databases for cells with similar expression patterns with high accuracy and throughput using locality-sensitive hashing.
-
Has reproduction · 60
Core transcriptional signatures of phase change in the migratory locust.
PMID 31292921 · PMC6881432 · Protein & cell · 2019 · 8 claims · 7 setups
PhaseCore genes defined by AC-PCA contribution to phase differentiation predict phase status with >87.5% accuracy
-
Has reproduction · 68
Systematic identification of ACE2 expression modulators reveals cardiomyopathy as a risk factor for mortality in COVID-19 patients.
PMID 35012625 · PMC8743438 · Genome biology · 2022 · 7 claims · 8 setups
GENEVA is a semi-automated, study-design-agnostic framework that mines large-scale public RNA-seq data to identify conditions modulating a gene of interest's expression
-
Has reproduction · 82
Reusable building blocks in biological systems.
PMID 30958230 · PMC6303794 · Journal of the Royal Society, Interface · 2018 · 8 claims · 4 setups
Biological systems can be decomposed into phenotypic building blocks (PBBs) via k-maximally reusable decompositions (k-MRD) that maximize average reusability across conditions.
-
Has reproduction · 79
Interpretable prediction models for widespread m6A RNA modification across cell lines and tissues.
PMID 37995291 · PMC10697738 · Bioinformatics (Oxford, England) · 2023 · 7 claims · 6 setups
CLSM6A, a CNN-based model set, predicts single-nucleotide-resolution m6A RNA modification sites across eight cell lines and three tissues in H. sapiens
-
Has reproduction · 94
Hierarchical cell-type identifier accurately distinguishes immune-cell subtypes enabling precise profiling of tissue microenvironment with single-cell RNA-sequencing.
PMID 36681937 · PMC10025442 · Briefings in bioinformatics · 2023 · 8 claims · 8 setups
HiCAT is a hierarchical, marker-based cell-type identifier that uses gene set analysis (GSA) scoring with markers structured in a three-level taxonomy tree (major-type, minor-type, subset)
-
Has reproduction · 66
A global database for modeling tumor-immune cell communication.
PMID 37438390 · PMC10338499 · Scientific data · 2023 · 7 claims · 6 setups
TICCom integrates 739 experimentally-validated or manually-curated TIC interactions collected from more than 3,000 literatures
-
Has reproduction · 71
Single-Cell Transcriptomic Landscape of Right-Sided Colon Cancer Reveals Cellular and Molecular Features of Metastatic Potential.
PMID 41898210 · PMC13024220 · Biomedicines · 2026 · 8 claims · 8 setups
Liver metastatic potential in RCC is marked by stem-like tumor states, metabolic plasticity, and microenvironmental remodeling.
-
Has reproduction · 87
Enhanced Generalizability of RNA Secondary Structure Prediction via Convolutional Block Attention Network and Ensemble Learning.
PMID 40871599 · PMC12388828 · Molecules (Basel, Switzerland) · 2025 · 8 claims · 8 setups
TrioFold integrates base-pairing clues from thermodynamic- and DL-based methods via ensemble learning and a convolutional block attention mechanism to enhance RSS prediction generalizability.
-
Has reproduction · 85
ScLRTC: imputation for single-cell RNA-seq data via low-rank tensor completion.
PMID 34844559 · PMC8628418 · BMC genomics · 2021 · 8 claims · 8 setups
scLRTC imputes dropout entries closest to the original expression values on simulated datasets, outperforming other state-of-the-art methods by SSE and PCC.