Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Ensembl 2006.
PMID 16381931 · PMC1347495 · Nucleic acids research · 2006 · 8 claims · 5 setups
Ensembl now provides annotation for 19 genomes, up from 4 the previous year, including new mammalian (Rhesus macaque, Opossum), chordate (Ciona intestinalis), and yeast genomes.
-
Full-text index only
Ensembl 2007.
PMID 17148474 · PMC1761443 · Nucleic acids research · 2007 · 8 claims · 7 setups
Ensembl added 18 new chordate genomes this year, increasing total genomes available from 15 to 33, the largest yearly increase to date.
-
Has reproduction · 83
Accurate prediction of metagenome-assembled genome completeness by MAGISTA, a random forest model built on alignment-free intra-bin statistics.
PMID 35248155 · PMC8898458 · Environmental microbiome · 2022 · 7 claims · 7 setups
MAGISTA, a random forest model built on alignment-free intra-bin distance-distribution statistics, can estimate MAG completeness and purity without relying on reference marker genes.
-
Full-text index only
A SNP-centric database for the investigation of the human genome.
PMID 15046636 · PMC395999 · BMC bioinformatics · 2004 · 8 claims · 3 setups
SNPper is a web-based, integrated SNP database combining dbSNP, the Human Genome sequence (Goldenpath), LocusLink, GeneOntology, and SWISS-PROT data with querying, visualization, and export tools.
-
Full-text index only
DAVID Knowledgebase: a gene-centered database integrating heterogeneous gene annotation resources to facilitate high-throughput gene functional analysis.
PMID 17980028 · PMC2186358 · BMC bioinformatics · 2007 · 7 claims · 3 setups
The DAVID Gene Concept, a single-linkage algorithm, merges gene clusters from Entrez Gene, UniRef100, and PIR-NREF100 that share protein IDs and species into unified DAVID gene clusters, improving cross-referencing between NCBI and UniProt systems
-
Has reproduction · 86
Plasmid transmission dynamics and evolution of partner quality in a natural population of Rhizobium leguminosarum.
PMID 41212030 · PMC12691615 · mBio · 2025 · 8 claims · 8 setups
Of the four most frequent plasmid types, types II and III have more stable size, larger core genomes, and track the chromosomal phylogeny (more vertical transmission), while types I and IV (pSym) vary in size and gene content with phylogenies consistent with frequent horizontal transmission.
-
Has reproduction · 24
MiGPC: a comprehensive catalog of enzybiotics from environmental metagenomes.
PMID 41888223 · PMC13172421 · Scientific reports · 2026 · 8 claims · 8 setups
MiGPC is the first genome-resolved metagenomic gene and protein catalog specifically targeted to enzybiotics
-
Has reproduction · 82
SMAGEXP: a galaxy tool suite for transcriptomics data meta-analysis.
PMID 30698691 · PMC6354025 · GigaScience · 2019 · 8 claims · 5 setups
SMAGEXP integrates the metaMA and metaRNASeq R packages into Galaxy to provide a unified tool suite for transcriptomics meta-analysis.
-
Full-text index only
An online database for brain disease research.
PMID 16594998 · PMC1489945 · BMC genomics · 2006 · 7 claims · 5 setups
SMRIDB is a comprehensive web-based database integrating gene expression data and clinical metadata to aid understanding of the genetic effects of brain disease (bipolar disorder, schizophrenia, depression)
-
Full-text index only
A re-annotation pipeline for Illumina BeadArrays: improving the interpretation of gene expression data.
PMID 19923232 · PMC2817484 · Nucleic acids research · 2010 · 8 claims · 7 setups
A Perl-based pipeline that BLASTs/BLATs Illumina probe sequences against genomes and transcript databases (RefSeq, UCSC Known Genes, UniGene/GenBank, Ensembl) can classify probes by quality grade (Perfect/Good/Bad/No match) and is applicable across 8 BeadArray platforms and other array types
-
Has reproduction · 67
Unraveling the timeline of gene expression: A pseudotemporal trajectory analysis of single-cell RNA sequencing data.
PMID 37994351 · PMC10663991 · F1000Research · 2023 · 7 claims · 7 setups
A reproducible R-based workflow combines Seurat (QC, clustering, integration), monocle3 (trajectory inference), and edgeR (pseudo-bulk time course analysis) to perform single-cell pseudotemporal time course analysis.
-
Has reproduction · 100
Exploring Gene Expression Patterns in Alzheimer's Disease Using a Human Microarray Data Meta-Analysis.
PMID 41744654 · PMC12938635 · Biology · 2026 · 6 claims · 7 setups
AD brains show a distinct transcriptomic profile with up-regulation of immune/inflammation genes and down-regulation of synapse/neuronal-signaling genes
-
Full-text index only
A clinical genetic method to identify mechanisms by which pain causes depression and anxiety.
PMID 16623937 · PMC1488826 · Molecular pain · 2006 · 8 claims · 4 setups
Pain-triggered depression/anxiety are mediated by distinct spino-parabrachial-hypothalamic-amygdalar neurochemical pathways compared to mood disorders independent of pain
-
Full-text index only
Neonatal salivary analysis reveals global developmental gene expression changes in the premature infant.
PMID 19959617 · PMC2853178 · Clinical chemistry · 2010 · 7 claims · 6 setups
Salivary genomic microarray analysis reveals global developmental gene expression changes in premature infants over postnatal age
-
Has reproduction · 89
Evaluating sequence data quality from the Swift Accel-Amplicon CFTR Panel.
PMID 31913291 · PMC6949293 · Scientific data · 2020 · 6 claims · 7 setups
The Accel-Amplicon CFTR panel generates sequencing data with high coverage depth and near 100% on-target reads.
-
Has reproduction · 37
A Bayesian approach to accurate and robust signature detection on LINCS L1000 data.
PMID 32003771 · PMC7203754 · Bioinformatics (Oxford, England) · 2020 · 6 claims · 4 setups
A novel Bayesian-based peak deconvolution algorithm gives unbiased likelihood estimations for peak locations and characterizes peaks with probability-based z-scores.
-
Has reproduction · 50
Exploiting convergent phenotypes to derive a pan-cancer cisplatin response gene expression signature.
PMID 37076665 · PMC10115855 · NPJ precision oncology · 2023 · 8 claims · 8 setups
A convergent-phenotype-based seed gene/co-expression method can extract consensus gene expression signatures predictive of response to chemotherapeutic drugs in the GDSC database
-
Has reproduction · 51
SGCP: a spectral self-learning method for clustering genes in co-expression networks.
PMID 38956463 · PMC11221046 · BMC bioinformatics · 2024 · 7 claims · 4 setups
SGCP, a spectral self-learning method, yields gene co-expression modules with higher GO enrichment than WGCNA, CoExpNets, and CEMiTool across 12 real gene expression datasets.
-
Has reproduction · 51
Cell type-specific eQTL analysis of COVID-19 based on single-cell transcriptomic data.
PMID 41064594 · PMC12501775 · NAR genomics and bioinformatics · 2025 · 8 claims · 8 setups
Single-cell eQTL analysis across eight immune cell types identified 2593 genes whose expression is significantly associated with common genetic polymorphisms, with most genes showing cell type-specific effects
-
Has reproduction · 61
Unraveling the role of bacteria with heritable versus non-heritable relative abundance in the gut on boar semen quality.
PMID 41199168 · PMC12590650 · Genetics, selection, evolution : GSE · 2025 · 6 claims · 7 setups
39 heritable and 91 non-heritable bacterial genera were identified in the boar gut based on heritability of relative abundance