Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 60
TRAPID 2.0: a web application for taxonomic and functional analysis of de novo transcriptomes.
PMID 34197621 · PMC8464036 · Nucleic acids research · 2021 · 8 claims · 8 setups
TRAPID 2.0 is a web application performing global characterization of de novo transcriptomes via structural, functional, and taxonomic annotation in an initial processing phase, followed by an exploratory phase of downstream analyses.
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes
-
Full-text index only
A general definition and nomenclature for alternative splicing events.
PMID 18688268 · PMC2467475 · PLoS computational biology · 2008 · 6 claims · 4 setups
Existing AS nomenclatures (Malko et al.'s 5-letter strings, Nagasaki et al.'s bit matrices, and the ASD/ATD/AEdb system) are redundant, ambiguous, or incapable of representing complex or large splicing variations.
-
Full-text index only
Needles in the haystack: identifying individuals present in pooled genomic data.
PMID 19798441 · PMC2747273 · PLoS genetics · 2009 · 8 claims · 7 setups
The distribution of T for null samples (individuals not in F or G) deviates strongly from the assumed standard normal, in both location and width.
-
Has reproduction · 87
CoINcIDE: A framework for discovery of patient subtypes across multiple datasets.
PMID 26961683 · PMC4784276 · Genome medicine · 2016 · 8 claims · 4 setups
CoINcIDE is a novel framework for discovering patient subtypes across multiple datasets that requires no between-dataset transformations (e.g., batch correction)
-
Full-text index only
Coverage of whole proteome by structural genomics observed through protein homology modeling database.
PMID 17146617 · PMC1769342 · Journal of structural and functional genomics · 2006 · 8 claims · 7 setups
FAMSBASE, a homology-modeling database of whole-genome ORFs, currently covers about 50% of predicted ORFs (368,724 of 734,193) across 276 genomes with modeled 3D structures.
-
Full-text index only
From genomics to chemical genomics: new developments in KEGG.
PMID 16381885 · PMC1347464 · Nucleic acids research · 2006 · 8 claims · 5 setups
KEGG BRITE has been formally added as a fourth main KEGG database to establish a logical foundation for functional interpretation and pathway reconstruction.
-
Full-text index only
Transposable elements are driving rapid adaptation of Enterococcus faecium.
PMID 42020750 · PMC13216065 · Nature · 2026 · 8 claims · 8 setups
E. faecium has the highest IS density among ESKAPEE pathogens, dominated by replicative ISL3 family elements
-
Has reproduction · 82
Whole-genome analysis of a multidrug-resistant Klebsiella michiganensis environmental isolate from an orthopedic ward in Mwanza, Tanzania reveals IncF-family plasmid replicon signatures associated with resistance determinants.
PMID 41957580 · PMC13173886 · BMC genomics · 2026 · 8 claims · 8 setups
An isolate phenotypically identified as K. oxytoca was reclassified by genome-based taxonomy (GTDB-Tk, ANI, BLAST) as Klebsiella michiganensis
-
Has reproduction · 63
Comparative transcriptome analysis of tomato (Solanum lycopersicum) in response to exogenous abscisic acid.
PMID 24289302 · PMC4046761 · BMC genomics · 2013 · 8 claims · 7 setups
Exogenous ABA alters the expression of a majority (54.73%) of expressed tomato leaf transcripts, with 2,787 significantly differentially expressed genes, predominantly up-regulated.
-
Full-text index only
Deducing topology of protein-protein interaction networks from experimentally measured sub-networks.
PMID 18598366 · PMC2474618 · BMC bioinformatics · 2008 · 7 claims · 6 setups
Experimentally measured protein-protein interaction sub-networks are not random samples of their parent networks.
-
Has reproduction · 61
Comprehensive transcriptome study to develop molecular resources of the copepod Calanus sinicus for their potential ecological applications.
PMID 24982883 · PMC4055022 · BioMed research international · 2014 · 8 claims · 8 setups
Illumina RNA-Seq with Trinity de novo assembly produced a C. sinicus transcriptome of 69,751 contigs (average 928.8 bp, N50 1,127 bp) from 58.9 million reads.
-
Full-text index only
Metagenomic study of the oral microbiota by Illumina high-throughput sequencing.
PMID 19796657 · PMC3568755 · Journal of microbiological methods · 2009 · 8 claims · 6 setups
The 16S rRNA V5 hypervariable region, amplified as a short ~82-base segment, provides reliable taxonomic identification of oral bacteria against public databases like HOMD.
-
Full-text index only
An AI-Enabled Single-Cell Transcriptomic Analysis Pipeline for Gene Signature Discovery in Natural Killer Cells Linked to Remission Outcomes in Chronic Myeloid Leukemia.
PMID 41972591 · PMC13072394 · Biology · 2026 · 8 claims · 7 setups
GAFA integrates latent-space representation, pseudotime trajectory modeling, GRN inference, and machine learning-based gene panel discovery into a single coherent pipeline, unlike existing workflows that treat these steps independently.
-
Full-text index only
scSurv: a deep generative model for single-cell survival analysis.
PMID 41429574 · PMC12797213 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 6 setups
scSurv combines a Cox proportional hazards model with a deep generative model (VAE) of single-cell transcriptomes to estimate individual cellular contributions to clinical outcomes
-
Full-text index only
Information extraction from full text scientific articles: where are the keywords?
PMID 12775220 · PMC166134 · BMC bioinformatics · 2003 · 8 claims · 5 setups
The keyword content of the five article sections (A, I, M, R, D) is heterogeneous, i.e., different sections carry different kinds of information.
-
Has reproduction · 76
Bayesian prediction of microbial oxygen requirement.
PMID 26913185 · PMC4743139 · F1000Research · 2013 · 7 claims · 8 setups
A naive Bayesian classifier based on presence/absence of class-associated Pfam-A domains can distinguish three oxygen requirement classes (aerobe, anaerobe, facultative anaerobe) from genome sequence, unlike prior studies that only made pairwise distinctions.
-
Has reproduction · 75
A step forward for Shiga toxin-producing Escherichia coli identification and characterization in raw milk using long-read metagenomics.
PMID 36748417 · PMC9836091 · Microbial genomics · 2022 · 8 claims · 6 setups
Long-read metagenomics enables isolation-independent identification and characterization of eae-positive STEC directly from raw milk.
-
Has reproduction · 30
Minimal metabolic pathway structure is consistent with associated biomolecular interactions.
PMID 24987116 · PMC4299494 · Molecular systems biology · 2014 · 8 claims · 8 setups
MinSpan, a mixed-integer linear optimization algorithm, computes the shortest, linearly independent pathways (sparsest basis of the null space of the stoichiometric matrix S) for genome-scale metabolic networks, which convex approaches (extreme pathways, elementary flux modes) cannot do at genome scale.
-
Full-text index only
metaFun: An analysis pipeline for metagenomic big data with fast and unified functional searches.
PMID 41530917 · PMC12818822 · Gut microbes · 2026 · 8 claims · 8 setups
metaFun is an open-source, end-to-end Nextflow/Apptainer pipeline integrating quality control, taxonomic profiling, functional profiling, de novo assembly, binning, genome assessment, comparative genomics, network analysis, and strain-level microdiversity analysis into a unified framework