Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
AutoCSA, an algorithm for high throughput DNA sequence variant detection in cancer genomes.
PMID 17485433 · PMC5947781 · Bioinformatics (Oxford, England) · 2007 · 7 claims · 2 setups
AutoCSA is an automated algorithm, extended from the CSA protocol, that detects DNA sequence variants in cancer genomes with minimal manual intervention
-
Has reproduction · 81
SEMdag: Fast learning of Directed Acyclic Graphs via node or layer ordering.
PMID 39775401 · PMC11709272 · PloS one · 2025 · 8 claims · 5 setups
SEMdag() is a two-step order-based algorithm for fast learning of high-dimensional linear SEMs, using knowledge-based (KB) or data-driven bottom-up (BU) node/layer ordering followed by penalized (L1) DAG estimation
-
Full-text index only
Systems biology in human health and disease.
PMID 17893698 · PMC2013921 · Molecular systems biology · 2007 · 8 claims · 5 setups
High-throughput quantitative proteomics of signaling networks (e.g., HER2 overexpression) can be correlated with biological responses like proliferation and migration to better understand pathways deregulated in cancer.
-
Full-text index only
From microarrays to genome duplications.
PMID 12914655 · PMC193639 · Genome biology · 2003 · 8 claims · 8 setups
Gene3D shows that most genes across sequenced genomes can be assigned to known structural domain families, many of which are shared across kingdoms of life
-
Has reproduction · 88
Human methylome variation across Infinium 450K data on the Gene Expression Omnibus.
PMID 33937763 · PMC8061458 · NAR genomics and bioinformatics · 2021 · 8 claims · 8 setups
Among annotated HM450K GEO samples, about two-thirds were from blood, one-quarter from brain, and about one-third were from cancer patients.
-
Has reproduction · 79
The relationship between PLOD1 expression level and glioma prognosis investigated using public databases.
PMID 34040895 · PMC8127981 · PeerJ · 2021 · 8 claims · 8 setups
PLOD1 mRNA expression is significantly higher in glioma tissue than in normal brain tissue
-
Has reproduction · 96
Mammary cell gene expression atlas links epithelial cell remodeling events to breast carcinogenesis.
PMID 34079055 · PMC8172904 · Communications biology · 2021 · 8 claims · 8 setups
Integration of five mouse scRNAseq datasets reveals a trifurcating lineage trajectory originating from embryonic mammary stem cells (MaSCs) that differentiates into three epithelial lineages (Basal, L-Alv, L-Hor) via unipotent progenitor clusters
-
Has reproduction · 61
miEAA 2.0: integrating multi-species microRNA enrichment analysis and workflow management systems.
PMID 32374865 · PMC7319446 · Nucleic acids research · 2020 · 8 claims · 5 setups
miEAA 2.0 extends miRNA enrichment analysis to ten species (previously only Homo sapiens), accepting both precursor and mature miRNA input.
-
Has reproduction · 50
SMAC, a computational system to link literature, biomedical and expression data.
PMID 31324861 · PMC6642118 · Scientific reports · 2019 · 8 claims · 8 setups
SMAC is a tool that extracts, prioritises, integrates and analyses biomedical and molecular data according to user-defined terms
-
Has reproduction · 68
Systematic identification of ACE2 expression modulators reveals cardiomyopathy as a risk factor for mortality in COVID-19 patients.
PMID 35012625 · PMC8743438 · Genome biology · 2022 · 8 claims · 9 setups
GENEVA (Gene Expression Variance Analysis) is a semi-automated framework that mines large-scale public RNA-seq datasets to identify conditions associated with a gene's expression variance
-
Has reproduction · 85
Deciphering the Immune Microenvironment at the Forefront of Tumor Aggressiveness by Constructing a Regulatory Network with Single-Cell and Spatial Transcriptomic Data.
PMID 38254989 · PMC10815467 · Genes · 2024 · 7 claims · 8 setups
High expression of transcription factors FOXA1 and EZH2 in malignant cells at the invasive front plays a key role in driving tumor progression
-
Has reproduction · 66
A global database for modeling tumor-immune cell communication.
PMID 37438390 · PMC10338499 · Scientific data · 2023 · 7 claims · 6 setups
TICCom integrates 739 experimentally-validated or manually-curated TIC interactions collected from more than 3,000 literatures
-
Has reproduction · 71
Comprehensive analysis of a novel RNA modifications-related model in the prognostic characterization, immune landscape and drug therapy of bladder cancer.
PMID 37124622 · PMC10131083 · Frontiers in genetics · 2023 · 8 claims · 8 setups
Two distinct RNA modification patterns exist among BCa samples with radically varying clinical outcomes and biological characteristics
-
Has reproduction · 67
Generative and integrative modeling for transcriptomics with formalin fixed paraffin embedded material.
PMID 41029822 · PMC12486589 · Journal of translational medicine · 2025 · 8 claims · 6 setups
The negative binomial distribution best fits fRNA-seq transcript counts, with little evidence supporting zero-inflated extensions
-
Has reproduction
Artificial Intelligence Approach in Machine Learning-Based Modeling and Networking of the Coronavirus Pathogenesis Pathway.
PMID 40699865 · PMC12191508 · Current issues in molecular biology · 2025 · 8 claims · 8 setups
The coronavirus pathogenesis pathway is activated in SARS-CoV-2-infected iPSC-derived cardiac cells and in SARS-CoV/SARS-CoV-2-infected LUAD cells
-
Has reproduction · 97
Determination of complete chromosomal haplotypes by bulk DNA sequencing.
PMID 33957932 · PMC8101039 · Genome biology · 2021 · 8 claims · 8 setups
A hierarchical computational strategy that first builds high-confidence local haplotype blocks from long-range/linked-read linkage and then concatenates them into whole-chromosome haplotypes using Hi-C contacts
-
Has reproduction · 87
Ultra-deep sequencing data from a liquid biopsy proficiency study demonstrating analytic validity.
PMID 35418127 · PMC9008010 · Scientific data · 2022 · 6 claims · 5 setups
This dataset is the most comprehensive public-facing dataset of ultra-deep ctDNA sequencing data generated to date
-
Has reproduction · 95
nf-rnaSeqCount: A Nextflow pipeline for obtaining raw read counts from RNA-seq data.
PMID 35574063 · PMC9097006 · South African computer journal = Suid-Afrikaanse rekenaartydskrif · 2021 · 7 claims · 5 setups
nf-rnaSeqCount is a portable, reproducible Nextflow pipeline that maps RNA-seq reads to a reference genome and quantifies gene abundance for differential expression analysis
-
Has reproduction · 50
Implementing the reuse of public DIA proteomics datasets: from the PRIDE database to Expression Atlas.
PMID 35701420 · PMC9197839 · Scientific data · 2022 · 8 claims · 8 setups
Introduced an open, containerised, Nextflow-orchestrated reanalysis pipeline for public SWATH-MS/DIA datasets covering metadata curation, data analysis, statistics, and Expression Atlas integration.
-
Has reproduction · 50
TOSCA: an automated Tumor Only Somatic CAlling workflow for somatic mutation detection without matched normal samples.
PMID 36699358 · PMC9710689 · Bioinformatics advances · 2022 · 6 claims · 4 setups
TOSCA is the first automated, modular open-source tumor-only somatic calling workflow for whole-exome and targeted panel sequencing, covering raw reads through variant classification.