Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 87
Analysis of Tumor-Infiltrating T-Cell Transcriptomes Reveal a Unique Genetic Signature across Different Types of Cancer.
PMID 36232369 · PMC9569723 · International journal of molecular sciences · 2022 · 8 claims · 8 setups
Common genes shared across five cancer types differ from those found in nonmalignant tissue-resident T-cells for each subset (CD4-T, CD8-T, Treg)
-
Has reproduction · 90
Assessment of genotyping array performance for genome-wide association studies and imputation in African cattle.
PMID 36057548 · PMC9441065 · Genetics, selection, evolution : GSE · 2022 · 7 claims · 6 setups
Commercially available bovine arrays are ineffective at capturing variants segregating among African indicine animals, with only 6% of high-LD (r2>0.8) variants captured by the best arrays versus 17% in African taurine and 25% in European taurine.
-
Full-text index only
GeneTide--Terra Incognita Discovery Endeavor: a new transcriptome focused member of the GeneCards/GeneNote suite of databases.
PMID 15608261 · PMC540076 · Nucleic acids research · 2005 · 8 claims · 7 setups
GeneTide integrates UniGene, DoTS, AceView, BLAT/GeneLoc genomic alignment, and GeneAnnot probe-set data into a unified Consensus/Uniqueness/Score scheme to associate ESTs with GeneCards genes
-
Full-text index only
Independent component analysis reveals new and biologically significant structures in micro array data.
PMID 16762055 · PMC1557674 · BMC bioinformatics · 2006 · 7 claims · 8 setups
ICA applied to three microarray datasets reveals many biologically significant components, including low-ranking ones not obvious by rank alone
-
Full-text index only
CRSD: a comprehensive web server for composite regulatory signature discovery.
PMID 16845073 · PMC1538777 · Nucleic acids research · 2006 · 7 claims · 5 setups
CRSD is a comprehensive web server integrating six large-scale databases (UniGene, mature microRNAs, putative promoter, TRANSFAC, pathway, GO) plus two newly constructed genome-wide databases (MRS and TRS) for composite regulatory signature discovery
-
Full-text index only
Ontological visualization of protein-protein interactions.
PMID 15707487 · PMC550656 · BMC bioinformatics · 2005 · 8 claims · 8 setups
Aggregating independently made GO 'protein binding' (IPI) annotations reveals larger, previously undescribed mouse protein-protein interaction networks
-
Full-text index only
SNAP: predict effect of non-synonymous polymorphisms on function.
PMID 17526529 · PMC1920242 · Nucleic acids research · 2007 · 7 claims · 8 setups
SNAP, a neural network-based method using sequence-derived information, predicts whether a non-synonymous SNP is neutral or non-neutral for protein function
-
Has reproduction
D2H2: diabetes data and hypothesis hub.
PMID 38107655 · PMC10723036 · Bioinformatics advances · 2023 · 6 claims · 5 setups
D2H2 is a web portal hosting hundreds of curated, uniformly reprocessed diabetes-relevant transcriptomics datasets from GEO with per-study visualization, differential expression, and single-gene queries.
-
Full-text index only
TPRpred: a tool for prediction of TPR-, PPR- and SEL1-like repeats from protein sequences.
PMID 17199898 · PMC1774580 · BMC bioinformatics · 2007 · 7 claims · 8 setups
TPRpred detects divergent/remote-homolog TPR repeat units that existing resources (Pfam, SMART, REP) fail to detect
-
Full-text index only
Efficacy assessment of SNP sets for genome-wide disease association studies.
PMID 17726055 · PMC2034459 · Nucleic acids research · 2007 · 6 claims · 4 setups
τ, derived from Shannon entropy and swept radius ɛ, approximates the relative sample size efficiency of a marker set for mapping a causal variant at a given map position compared to a maximally polymorphic SNP
-
Has reproduction · 74
An open RNA-Seq data analysis pipeline tutorial with an example of reprocessing data from a recent Zika virus study.
PMID 27583132 · PMC4972086 · F1000Research · 2016 · 6 claims · 6 setups
An open-source, reproducible RNA-seq pipeline delivered as an IPython notebook and Docker image can process raw RNA-seq data into interactive PCA/HC plots, enrichment results, and small-molecule predictions with minimal setup overhead
-
Has reproduction · 65
Cancer-predicting transcriptomic and epigenetic signatures revealed for ulcerative colitis in patient-derived epithelial organoids.
PMID 29983891 · PMC6033374 · Oncotarget · 2018 · 8 claims · 6 setups
UC patient-derived epithelial organoids histologically phenocopy primary UC tissue, while non-IBD organoids resemble healthy colonic epithelium.
-
Has reproduction · 76
GeneSetCart: assembling, augmenting, combining, visualizing, and analyzing gene sets.
PMID 40208796 · PMC11984350 · GigaScience · 2025 · 8 claims · 8 setups
GeneSetCart is a web-based platform that lets users assemble, augment, combine, visualize, and analyze gene sets from multiple sources in one place
-
Has reproduction · 100
Lipopolysaccharide distinctively alters human microglia transcriptomes to resemble microglia from Alzheimer's disease mouse models.
PMID 36254682 · PMC9612871 · Disease models & mechanisms · 2022 · 8 claims · 8 setups
iPSC-microglia show a shared core transcriptional response to ATPγS and to LPS+IFN-γ, suggesting a convergent mechanism of action
-
Has reproduction · 49
oPOSSUM-3: advanced analysis of regulatory motif over-representation across genes or ChIP-Seq datasets.
PMID 22973536 · PMC3429929 · G3 (Bethesda, Md.) · 2012 · 8 claims · 6 setups
oPOSSUM-3 is a web-accessible system that identifies over-represented TFBS and TFBS families in DNA sequences of co-expressed genes or in sequences from high-throughput methods such as ChIP-Seq.
-
Has reproduction · 63
RummaGEO: Automatic mining of human and mouse gene sets from GEO.
PMID 39569206 · PMC11573963 · Patterns (New York, N.Y.) · 2024 · 8 claims · 7 setups
RummaGEO is a gene expression signature search engine built from automatically mined human and mouse RNA-seq perturbation studies in GEO
-
Full-text index only
Predicting failure rate of PCR in large genomes.
PMID 18492719 · PMC2441781 · Nucleic acids research · 2008 · 7 claims · 8 setups
The number of predicted primer-binding sites in genomic DNA is the most important factor determining PCR failure.
-
Full-text index only
Cancer-specific high-throughput annotation of somatic mutations: computational prediction of driver missense mutations.
PMID 19654296 · PMC2763410 · Cancer research · 2009 · 7 claims · 7 setups
CHASM, a Random Forest-based computational method, was developed to identify and prioritize missense mutations likely to be functional drivers of tumor cell proliferation.
-
Full-text index only
The global landscape of sequence diversity.
PMID 17996061 · PMC2258180 · Genome biology · 2007 · 7 claims · 5 setups
Eukaryotic sequence datasets show substantially greater genetic diversity (higher sequence/gene family discovery rates) than bacterial datasets, likely related to differences in modes of genetic inheritance.
-
Has reproduction · 71
Spatial organization shapes the turnover of a bacterial transcriptome.
PMID 27198188 · PMC4874777 · eLife · 2016 · 7 claims · 6 setups
The E. coli transcriptome is spatially organized genome-wide: mRNAs encoding inner-membrane proteins are enriched at the membrane, while mRNAs encoding cytoplasmic, periplasmic and outer-membrane proteins are distributed throughout the cytoplasm.