Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Bases and spaces: resources on the web for accessing the draft human genome.
PMID 11178254 · PMC138875 · Genome biology · 2000 · 8 claims · 8 setups
By combining currently available genomic databases and mapping resources (GenBank/Entrez, UniGene, RH maps, BAC fingerprint maps, Ensembl, NIX), it is possible to devise strategies that fully exploit the fragmentary draft human genome sequence.
-
Full-text index only
NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.
PMID 15608248 · PMC539979 · Nucleic acids research · 2005 · 7 claims · 5 setups
RefSeq provides a curated, non-redundant, explicitly linked collection of genomic, transcript and protein sequences spanning prokaryotes, eukaryotes and viruses.
-
Full-text index only
The Functional RNA Database 3.0: databases to support mining and annotation of functional RNAs.
PMID 18948287 · PMC2686472 · Nucleic acids research · 2009 · 8 claims · 5 setups
fRNAdb 3.0 is a completely rebuilt sequence database hosting a much larger collection of known/predicted non-coding RNA sequences with improved search functionality
-
Full-text index only
MBGD update 2010: toward a comprehensive resource for exploring microbial genome diversity.
PMID 19906735 · PMC2808943 · Nucleic acids research · 2010 · 8 claims · 6 setups
MBGD allows users to create ortholog groups using a specified subgroup of organisms, distinguishing it from other comparative genomics resources
-
Has reproduction · 90
pysradb: A Python package to query next-generation sequencing metadata and data from NCBI Sequence Read Archive.
PMID 31114675 · PMC6505635 · F1000Research · 2019 · 6 claims · 7 setups
pysradb provides a simple, user-friendly command-line interface for querying metadata and downloading datasets from SRA without requiring knowledge of a programming language.
-
Has reproduction · 83
Current status of use of high throughput nucleotide sequencing in rheumatology.
PMID 33408124 · PMC7789458 · RMD open · 2021 · 8 claims · 4 setups
RNA-Seq is the most represented HTS assay in rheumatology research (n=457, 65%), used for biomarker identification in blood or synovial tissue
-
Has reproduction
All of gene expression (AOE): An integrated index for public gene expression databases.
PMID 31978081 · PMC6980531 · PloS one · 2020 · 8 claims · 5 setups
AOE integrates publicly available gene expression data from GEO, ArrayExpress, and GEA into a single searchable index.
-
Has reproduction · 53
PulmonDB: a curated lung disease gene expression database.
PMID 31949184 · PMC6965635 · Scientific reports · 2020 · 6 claims · 6 setups
PulmonDB is a curated, web-based gene expression database and R package integrating microarray and RNA-seq data for COPD and IPF with manually curated controlled-vocabulary annotation.
-
Has reproduction · 59
Downregulation of Splicing Factor PTBP1 Curtails FBXO5 Expression to Promote Cellular Senescence in Lung Adenocarcinoma.
PMID 39057099 · PMC11276454 · Current issues in molecular biology · 2024 · 8 claims · 8 setups
PTBP1 is significantly upregulated across multiple cancer types including LUAD, and higher PTBP1 levels are associated with worse LUAD patient survival
-
Has reproduction · 85
Exploring microproteins from various model organisms using the mip-mining database.
PMID 37919660 · PMC10623795 · BMC genomics · 2023 · 5 claims · 4 setups
Mip-mining is a database of 336 curated RNA-seq datasets from 8626 samples across nine species, built specifically to explore microprotein functions under stress and disease conditions
-
Has reproduction · 77
Representing and querying disease networks using graph databases.
PMID 27462371 · PMC4960687 · BioData mining · 2016 · 7 claims · 8 setups
Graph databases are well suited for representing biological information that is highly connected, semi-structured, and unpredictable.
-
Has reproduction · 78
Machine learning and free energy clustering reveal PAH protein binding linked to AD risk.
PMID 41953002 · PMC13053772 · iScience · 2026 · 7 claims · 8 setups
An integrated framework of bioinformatics, machine learning, and ΔG clustering can prioritize PAHs for AD-associated neurotoxicity.
-
Has reproduction · 89
Improved eukaryotic detection compatible with large-scale automated analysis of metagenomes.
PMID 37032329 · PMC10084625 · Microbiome · 2023 · 8 claims · 7 setups
MAPQ ≥30 filtering improves precision but substantially reduces recall, especially for unrepresented/divergent eukaryotic taxa
-
Full-text index only
Comparison of complete nuclear receptor sets from the human, Caenorhabditis elegans and Drosophila genomes.
PMID 11532213 · PMC55326 · Genome biology · 2001 · 7 claims · 5 setups
The human genome contains fewer than 50 functional nuclear receptors, far fewer than C. elegans and about twice as many as Drosophila
-
Full-text index only
An integrated database of genes responsive to the Myc oncogenic transcription factor: identification of direct genomic targets.
PMID 14519204 · PMC328458 · Genome biology · 2003 · 8 claims · 6 setups
The Myc Target Gene database integrates literature evidence to prioritize candidate Myc-responsive genes and cluster them into functional groups
-
Full-text index only
Molecular phylogeny of the kelch-repeat superfamily reveals an expansion of BTB/kelch proteins in animals.
PMID 13678422 · PMC222960 · BMC bioinformatics · 2003 · 8 claims · 8 setups
The human genome encodes at least 71 kelch-repeat proteins
-
Full-text index only
Analysis of the glutathione S-transferase (GST) gene family.
PMID 15607001 · PMC3500200 · Human genomics · 2004 · 8 claims · 4 setups
The complete human GST gene family comprises 16 genes in six subfamilies: alpha (GSTA), mu (GSTM), omega (GSTO), pi (GSTP), theta (GSTT) and zeta (GSTZ).
-
Full-text index only
The revolution of the biology of the genome.
PMID 15040884 · PMC7091781 · Cell research · 2004 · 8 claims · 6 setups
Polyploidization and gene duplication are the major mechanisms increasing eukaryotic genome size.
-
Full-text index only
A genome-wide survey demonstrates widespread non-linear mRNA in expressed sequences from multiple species.
PMID 16237125 · PMC1258171 · Nucleic acids research · 2005 · 8 claims · 6 setups
A genome-wide computational survey identifies 245 genes in mammals (264 across six species) that produce RREO events in expressed sequences
-
Full-text index only
Human Lsg1 defines a family of essential GTPases that correlates with the evolution of compartmentalization.
PMID 16209721 · PMC1262696 · BMC biology · 2005 · 8 claims · 9 setups
hLsg1 is the human orthologue of yeast Lsg1p and defines a family of circularly permuted GTPases named YRG (YlqF Related GTPases)