Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 95
Increased prevalence of hybrid epithelial/mesenchymal state and enhanced phenotypic heterogeneity in basal breast cancer.
PMID 38974967 · PMC11225361 · iScience · 2024 · 7 claims · 7 setups
Luminal breast cancer gene expression signature is closely/positively associated with an epithelial signature
-
Has reproduction · 24
MiGPC: a comprehensive catalog of enzybiotics from environmental metagenomes.
PMID 41888223 · PMC13172421 · Scientific reports · 2026 · 8 claims · 8 setups
MiGPC is the first genome-resolved metagenomic gene and protein catalog specifically targeted to enzybiotics
-
Has reproduction · 80
VGEA: an RNA viral assembly toolkit.
PMID 34567846 · PMC8428259 · PeerJ · 2021 · 8 claims · 5 setups
VGEA is a Snakemake workflow that chains existing tools (fastp, BWA, SAMtools, IVA, shiver, SeqKit, QUAST, MultiQC) into an all-in-one RNA viral genome assembly pipeline
-
Has reproduction · 49
EDGE COVID-19: a web platform to generate submission-ready genomes from SARS-CoV-2 sequencing efforts.
PMID 35561186 · PMC9113274 · Bioinformatics (Oxford, England) · 2022 · 7 claims · 5 setups
EDGE COVID-19 (EC-19) is a web-based platform that automates QC, reference-based variant/consensus calling, lineage determination, and submission of SARS-CoV-2 genomes and metadata to GenBank, GISAID and INSDC for both Illumina and ONT data.
-
Has reproduction · 50
Implementing the reuse of public DIA proteomics datasets: from the PRIDE database to Expression Atlas.
PMID 35701420 · PMC9197839 · Scientific data · 2022 · 8 claims · 5 setups
An open, containerised, Nextflow-orchestrated reanalysis pipeline combining metadata annotation, SWATH-MS analysis, statistical analysis, and Expression Atlas integration was developed for public DIA data.
-
Has reproduction · 64
metaGEM: reconstruction of genome scale metabolic models directly from metagenomes.
PMID 34614189 · PMC8643649 · Nucleic acids research · 2021 · 8 claims · 8 setups
metaGEM enables end-to-end reconstruction of FBA-ready GEMs directly from metagenomes without relying on reference genomes
-
Has reproduction · 65
FusionQ: a novel approach for gene fusion detection and quantification from paired-end RNA-Seq.
PMID 23768108 · PMC3691734 · BMC bioinformatics · 2013 · 8 claims · 8 setups
FusionQ is a novel tool that detects gene fusions, constructs chimerical transcript structures, and estimates their abundances from paired-end RNA-Seq data.
-
Has reproduction · 43
Compression of structured high-throughput sequencing data.
PMID 24260313 · PMC3832420 · PloS one · 2013 · 8 claims · 7 setups
Leveraging an explicit data schema (separate field encoding, field modeling, template compression, domain modeling) enables stronger compression of HTS alignment data than general-purpose compression of serialized bytes.
-
Full-text index only
Comparative phosphoproteomics reveals evolutionary and functional conservation of phosphorylation across eukaryotes.
PMID 18828897 · PMC2760871 · Genome biology · 2008 · 8 claims · 8 setups
The overlap between phosphoproteomes of six eukaryotes (human, mouse, fly, yeast, plant, zebrafish) is significantly greater than expected by chance.
-
Has reproduction · 64
Characterizing Neutrophil Subtypes in Cancer Using scRNA Sequencing Demonstrates the Importance of IL1β/CXCR2 Axis in Generation of Metastasis-specific Neutrophils.
PMID 38358352 · PMC10903300 · Cancer research communications · 2024 · 8 claims · 7 setups
Two main neutrophil subtypes exist in primary tumors: an activated subtype sharing transcriptomic signatures with healthy neutrophils, and a tumor-specific subtype.
-
Has reproduction · 26
Integrating transcriptomic datasets across neurological disease identifies unique myeloid subpopulations driving disease-specific signatures.
PMID 36527260 · PMC10952672 · Glia · 2023 · 6 claims · 3 setups
The bulk microglial and monocyte transcriptomic program is highly contingent on the disease environment, challenging the notion of a universal microglial disease signature
-
Full-text index only
Analyses and comparison of accuracy of different genotype imputation methods.
PMID 18958166 · PMC2569208 · PloS one · 2008 · 8 claims · 3 setups
Stronger LD produces higher imputation accuracy rates for all five methods
-
Full-text index only
Size matters: just how big is BIG?: Quantifying realistic sample size requirements for human genome epidemiology.
PMID 18676414 · PMC2639365 · International journal of epidemiology · 2009 · 7 claims · 2 setups
Conventional power calculations for case-control studies disregard analytic complexity (e.g. clinical assessment errors, unmeasured aetiological determinants) and can seriously underestimate true sample size requirements
-
Has reproduction · 78
Detecting tipping points of complex diseases by network information entropy.
PMID 38960408 · PMC11221888 · Briefings in bioinformatics · 2024 · 8 claims · 4 setups
NIEE can detect critical states or tipping points in diverse data types, including bulk and single-sample expression data
-
Has reproduction · 67
Comparison of Metagenomics and Metatranscriptomics Tools: A Guide to Making the Right Choice.
PMID 36553546 · PMC9777648 · Genes · 2022 · 8 claims · 1 setups
16S rRNA gene sequencing enables taxonomic identification of bacteria/archaea via hypervariable regions without amplifying human DNA, but is limited by short-read biases (GC bias, sequencing errors) and poor species-level resolution
-
Has reproduction · 80
SLDMS: A Tool for Calculating the Overlapping Regions of Sequences.
PMID 35046988 · PMC8761809 · Frontiers in plant science · 2021 · 8 claims · 5 setups
SLDMS is a novel method for computing overlapping regions of sequencing reads using suffix array (SA), longest common prefix (LCP) array, document array (DA), and a monotonic stack.
-
Has reproduction · 87
A robust data scaling algorithm to improve classification accuracies in biomedical data.
PMID 27612635 · PMC5016890 · BMC bioinformatics · 2016 · 8 claims · 2 setups
Models trained on data scaled by the GL algorithm outperform models trained on data scaled by the Min-max or Z-score algorithms across 16 binary classification tasks, measured by AUROC and percentage of correct classification
-
Has reproduction · 80
PanglaoDB: a web server for exploration of mouse and human single-cell RNA sequencing data.
PMID 30951143 · PMC6450036 · Database : the journal of biological databases and curation · 2019 · 7 claims · 7 setups
PanglaoDB is a web server providing pre-processed and pre-computed analyses of >1054 single-cell experiments (>4 million cells) from mouse and human across many tissues and platforms.
-
Has reproduction · 91
A reference profile-free deconvolution method to infer cancer cell-intrinsic subtypes and tumor-type-specific stromal profiles.
PMID 32111252 · PMC7049190 · Genome medicine · 2020 · 8 claims · 8 setups
DeClust is a reference-profile-free deconvolution method that incorporates molecular subtyping directly into the deconvolution process, outputting cohort-level cancer subtype and stromal reference profiles rather than per-individual profiles
-
Has reproduction · 96
Mammary cell gene expression atlas links epithelial cell remodeling events to breast carcinogenesis.
PMID 34079055 · PMC8172904 · Communications biology · 2021 · 8 claims · 8 setups
Integration of five mouse scRNAseq datasets reveals a trifurcating lineage trajectory originating from embryonic mammary stem cells (MaSCs) that differentiates into three epithelial lineages (Basal, L-Alv, L-Hor) via unipotent progenitor clusters