Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 77
Representing and querying disease networks using graph databases.
PMID 27462371 · PMC4960687 · BioData mining · 2016 · 7 claims · 8 setups
Graph databases are well suited for representing biological information that is highly connected, semi-structured, and unpredictable.
-
Has reproduction · 75
Graph-Based Approaches Significantly Improve the Recovery of Antibiotic Resistance Genes From Complex Metagenomic Datasets.
PMID 34690959 · PMC8528159 · Frontiers in microbiology · 2021 · 8 claims · 6 setups
GraphAMR, a Nextflow pipeline that aligns AMR profile HMMs (or AA sequences) to metagenomic assembly graphs via PathRacer, then dereplicates and annotates hits, recovers more and more complete AMR genes than contig-based or read-based methods.
-
Full-text index only
Prediction by graph theoretic measures of structural effects in proteins arising from non-synonymous single nucleotide polymorphisms.
PMID 18654622 · PMC2447880 · PLoS computational biology · 2008 · 8 claims · 5 setups
Bongo identifies mutations causing local and global structural effects with a remarkably low false positive rate
-
Has reproduction · 89
Graph Random Forest: A Graph Embedded Algorithm for Identifying Highly Connected Important Features.
PMID 37509188 · PMC10377046 · Biomolecules · 2023 · 8 claims · 3 setups
Graph Random Forest (GRF) embeds graph/network information directly into the decision-tree building process by splitting on features in the k-hop neighborhood of a data-driven head-splitting node.
-
Has reproduction · 95
Sparse and skew hashing of K-mers.
PMID 35758794 · PMC9235479 · Bioinformatics (Oxford, England) · 2022 · 7 claims · 4 setups
Exploiting sparsity and skewed distribution of k-mer minimizers with minimal perfect hashing substantially improves the space/time trade-off of a k-mer dictionary compared to best-known solutions
-
Has reproduction · 86
LMAS: evaluating metagenomic short de novo assembly methods through defined communities.
PMID 36576131 · PMC9795473 · GigaScience · 2022 · 8 claims · 5 setups
LMAS (Last Metagenomic Assembler Standing) is a flexible, Nextflow-based, Docker-containerized automated workflow for benchmarking de novo metagenomic assemblers against defined mock communities, producing an interactive HTML report.
-
Full-text index only
Oligomeric protein structure networks: insights into protein-protein interactions.
PMID 16336694 · PMC1326230 · BMC bioinformatics · 2005 · 8 claims · 6 setups
Interface amino acid clusters identified at Imin=6% correlate well with residues losing accessible surface area (δASA) upon oligomerization
-
Full-text index only
The use of edge-betweenness clustering to investigate biological function in protein interaction networks.
PMID 15740614 · PMC555937 · BMC bioinformatics · 2005 · 8 claims · 7 setups
Edge-Betweenness clustering separates protein interaction graphs into subgraphs whose GO term distributions show significant correlations, revealing biologically meaningful functional modules.
-
Full-text index only
Identifying repeat domains in large genomes.
PMID 16507140 · PMC1431705 · Genome biology · 2006 · 7 claims · 5 setups
A repeat domain graph, built using a modified A-Bruijn graph framework, decomposes a repeat library into shared repeat domains and reveals the mosaic structure of repeat families.
-
Full-text index only
The signal in the genomes.
PMID 16683016 · PMC1447653 · PLoS computational biology · 2006 · 7 claims · 3 setups
A high breakpoint reuse rate in the output of rearrangement algorithms indicates loss of historical signal, not good evidence for genomic fragile regions
-
Full-text index only
A new method for 2D gel spot alignment: application to the analysis of large sample sets in clinical proteomics.
PMID 18957120 · PMC2628390 · BMC bioinformatics · 2008 · 8 claims · 2 setups
Sili2DGel represents recursive gel matching results as a weighted undirected graph and identifies SAP by finding cliques and pseudocliques (dense subgraphs) after edge-weight filtering, strength-metric-based graph reduction, and cluster refinement.
-
Full-text index only
ECgene: an alternative splicing database update.
PMID 17132829 · PMC1716719 · Nucleic acids research · 2007 · 8 claims · 5 setups
ECgene provides functional annotation (domain, GO, expression pattern) for alternatively spliced genes
-
Has reproduction · 45
Accurate sequence variant genotyping in cattle using variation-aware genome graphs.
PMID 31092189 · PMC6521551 · Genetics, selection, evolution : GSE · 2019 · 8 claims · 7 setups
Graphtyper outperformed GATK and SAMtools in genotype concordance, non-reference sensitivity, and non-reference discrepancy compared to microarray genotypes
-
Has reproduction · 45
Identifying and classifying trait linked polymorphisms in non-reference species by walking coloured de bruijn graphs.
PMID 23536903 · PMC3607606 · PloS one · 2013 · 8 claims · 9 setups
Bubbleparse detects sequence variants directly from NGS reads without a reference genome, using the coloured de Bruijn graph implementation of Cortex plus a new depth-first bubble-finding module.
-
Full-text index only
Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space.
PMID 16481312 · PMC1373602 · Nucleic acids research · 2006 · 8 claims · 7 setups
The number of protein families continues to expand steadily as more genomes are sequenced, showing no sign of saturation.
-
Full-text index only
Identification and analysis of co-occurrence networks with NetCutter.
PMID 18781200 · PMC2526157 · PloS one · 2008 · 8 claims · 4 setups
Random sampling from a complete permutation set of the bipartite graph permits co-occurrence analysis with optimal stringency, and the edge-swapping (ES) model closely approximates this and is the preferred null-model among six tested.
-
Has reproduction · 81
SEMdag: Fast learning of Directed Acyclic Graphs via node or layer ordering.
PMID 39775401 · PMC11709272 · PloS one · 2025 · 8 claims · 5 setups
SEMdag() is a two-step order-based algorithm for fast learning of high-dimensional linear SEMs, using knowledge-based (KB) or data-driven bottom-up (BU) node/layer ordering followed by penalized (L1) DAG estimation
-
Full-text index only
An analysis of human microRNA and disease associations.
PMID 18923704 · PMC2559869 · PloS one · 2008 · 8 claims · 8 setups
MicroRNAs tend to show similar dysfunctional evidence (both up- or both down-regulated) for diseases within the same disease cluster, and different dysfunctional evidence between different disease clusters.
-
Has reproduction · 76
Topologically inferring pathway activity toward precise cancer classification via integrating genomic and metabolomic data: prostate cancer as a case.
PMID 26286638 · PMC4541321 · Scientific reports · 2015 · 6 claims · 4 setups
DRW-GM integrates gene expression and metabolomic profiles via directed random walk on a global gene–metabolite pathway graph to weight genes by topological importance and infer reproducible pathway activities
-
Full-text index only
Reconstructing the genomic architecture of mammalian ancestors using multispecies comparative maps.
PMID 15601531 · PMC3525001 · Human genomics · 2003 · 8 claims · 4 setups
The MGR algorithm applied to human, mouse, cat and cattle comparative maps can impute an ancestral mammalian genome composed of conserved segments.