Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
The human phylome.
PMID 17567924 · PMC2394744 · Genome biology · 2007 · 6 claims · 5 setups
Reconstruction of the human phylome: evolutionary trees for all human proteins and their homologs among 39 fully sequenced eukaryotic genomes, using a pipeline combining alignment trimming, NJ, ML (PhyML) and Bayesian (MrBayes) methods.
-
Has reproduction
Geometric Reliability of Super-Resolution Reconstructed Images from Clinical Fetal MRI in the Second Trimester.
PMID 37284977 · PMC10406722 · Neuroinformatics · 2023 · 7 claims · 8 setups
NiftyMIC and MIALSRTK provide reliable SR reconstructed volumes suitable for biometric assessments
-
Full-text index only
CLEAN: CLustering Enrichment ANalysis.
PMID 19640299 · PMC2734555 · BMC bioinformatics · 2009 · 8 claims · 4 setups
The gene-specific CLEAN score improves reproducibility of cluster analysis conclusions across independent datasets compared to the traditional cluster-wide score (cwCLEAN).
-
Has reproduction · 97
CellFishing.jl: an ultrafast and scalable cell search method for single-cell RNA sequencing.
PMID 30744683 · PMC6371477 · Genome biology · 2019 · 8 claims · 5 setups
CellFishing.jl achieves accuracy comparable to state-of-the-art software (scmap-cell) but is markedly faster
-
Full-text index only
PPC: an algorithm for accurate estimation of SNP allele frequencies in small equimolar pools of DNA using data from high density microarrays.
PMID 16199750 · PMC1240117 · Nucleic acids research · 2005 · 7 claims · 6 setups
The PPC algorithm, which applies a probe-pair-specific second-degree polynomial correction, increases the accuracy of allele frequency estimates from pooled DNA compared with previously described algorithms
-
Full-text index only
Computer-aided identification of polymorphism sets diagnostic for groups of bacterial and viral genetic variants.
PMID 17672919 · PMC1973086 · BMC bioinformatics · 2007 · 6 claims · 8 setups
The Not-N algorithm, incorporated into the Minimum SNPs program, identifies small marker sets diagnostic for user-defined subgroups of genetic variants with 0% false negatives
-
Full-text index only
Swarm intelligence based wavelet coefficient feature selection for mass spectral classification: an application to proteomics data.
PMID 19733729 · PMC2748225 · Analytica chimica acta · 2009 · 8 claims · 4 setups
ACA-based wavelet coefficient feature selection can achieve up to 100% classification accuracy on training, validating, and independent testing sets using only 5 selected features.
-
Full-text index only
iS2C2: a cointelligent platform for mechanistic discovery of disease cellular crosstalk.
PMID 42108258 · PMC13158306 · Signal transduction and targeted therapy · 2026 · 8 claims · 5 setups
iS2C2 integrates the S2C2 cell-cell communication algorithm with LLMs to generate biologically interpretable hypotheses from scRNA-seq and spatial transcriptomics data
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Full-text index only
IDEAL-Q, an automated tool for label-free quantitation analysis using an efficient peptide alignment approach and spectral data validation.
PMID 19752006 · PMC2808259 · Molecular & cellular proteomics : MCP · 2010 · 6 claims · 5 setups
IDEAL-Q predicts the elution time of peptides unidentified in a given LC-MS/MS run (but identified in others) using a computation-efficient linear regression plus fragmental refining function, avoiding costly whole-dataset pattern recognition
-
Has reproduction · 53
Gene Dosage Analysis on the Single-Cell Transcriptomes Linking Cotranslational Protein Targeting to Metastatic Triple-Negative Breast Cancer.
PMID 34577617 · PMC8472593 · Pharmaceuticals (Basel, Switzerland) · 2021 · 8 claims · 7 setups
A computational framework mapping single-cell Z-score expression to matched patient-level CNV data can identify recurrent, cross-cell-type copy-number-driven expression (dosage) events.
-
Full-text index only
Integration of text- and data-mining using ontologies successfully selects disease gene candidates.
PMID 15767279 · PMC1065256 · Nucleic acids research · 2005 · 7 claims · 6 setups
Integrating eVOC anatomical ontology-based text-mining of PubMed abstracts with data-mining of gene expression annotation successfully selects and prioritizes candidate disease genes
-
Full-text index only
CorGen--measuring and generating long-range correlations for DNA sequence analysis.
PMID 16845099 · PMC1538783 · Nucleic acids research · 2006 · 8 claims · 3 setups
CorGen is a web server that measures long-range correlations in DNA sequences and generates random sequences with the same (or user-specified) correlation and composition parameters
-
Has reproduction · 83
Hierarchical classification-based pan-cancer methylation analysis to classify primary cancer.
PMID 38066424 · PMC10709847 · BMC bioinformatics · 2023 · 8 claims · 8 setups
CHCT, a two-tier hierarchical classification tool built from methylation data, accurately classifies primary cancer type across 30 cancer types.
-
Has reproduction · 76
Tracing human genetic histories and natural selection with precise local ancestry inference.
PMID 40379651 · PMC12084304 · Nature communications · 2025 · 7 claims · 7 setups
Orchestra, a two-stage LAI method combining a recombination-distance base layer with a deep learning (convolutional + attention) smoothing module, outperforms RFmix, FLARE and Gnomix in precision and recall across simulated admixture generations.
-
Full-text index only
Clustering of phosphorylation site recognition motifs can be exploited to predict the targets of cyclin-dependent kinase.
PMID 17316440 · PMC1852407 · Genome biology · 2007 · 8 claims · 6 setups
CDK consensus motifs are frequently clustered (closely spaced) in known CDK substrate proteins rather than uniformly distributed
-
Full-text index only
Discovery of molecular subtypes in leiomyosarcoma through integrative molecular profiling.
PMID 19901961 · PMC2820592 · Oncogene · 2010 · 8 claims · 6 setups
Unsupervised gene expression clustering identifies 3 reproducible molecular subtypes of LMS (Group I/muscle-enriched, Group II, Group III)
-
Full-text index only
Full-length 16S rRNA nanopore sequencing enables species resolution of Fusobacterium associated with colorectal cancer.
PMID 41963777 · PMC13078227 · Gut microbes · 2026 · 8 claims · 7 setups
Full-length 16S rRNA ONT sequencing combined with custom demultiplexing (nanoMux) enables robust species-level discrimination within the Fusobacterium genus