Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
PolySearch: a web-based text mining system for extracting relationships between human diseases, genes, mutations, drugs and metabolites.
PMID 18487273 · PMC2447794 · Nucleic acids research · 2008 · 8 claims · 7 setups
PolySearch supports more than 50 different classes of queries against nearly a dozen types of text, abstract, or bioinformatic databases
-
Full-text index only
Filtering high-throughput protein-protein interaction data using a combination of genomic features.
PMID 15833142 · PMC1127019 · BMC bioinformatics · 2005 · 8 claims · 8 setups
A combination of three genomic features (interacting Pfam domains, GO annotations, sequence homology) using naive Bayesian networks predicts true protein-protein interactions with high sensitivity and good specificity.
-
Full-text index only
Detection of alternative splicing: deep sequencing or deep learning?
PMID 41520225 · PMC12790623 · Briefings in bioinformatics · 2026 · 8 claims · 8 setups
Sequence-based deep learning tools (AlphaGenome, SpliceAI, DeepSplice) show potential for initial hypothesis development and as additional filters in standard RNA-seq pipelines, especially when sequencing depth is limited.
-
Full-text index only
Cancer genome standards for long-read sequencing using cancer cell line mixtures.
PMID 41934171 · PMC13137868 · GigaScience · 2026 · 8 claims · 6 setups
Long-read variant calling tools achieve recall rates comparable to short-read gold standards
-
Has reproduction · 78
A network-based model of Aspergillus fumigatus elucidates regulators of development and defensive natural products of an opportunistic pathogen.
PMID 41505094 · PMC12781895 · Nucleic acids research · 2026 · 7 claims · 6 setups
MERLIN-P-TFA network inference on 18 curated public RNA-seq datasets produced a genome-wide GRN resource for A. fumigatus called GRAsp.
-
Full-text index only
Detection of germline BRCA1 mutations by Multiple-Dye Cleavase Fragment Length Polymorphism (MD-CFLP) method.
PMID 11556835 · PMC2375072 · British journal of cancer · 2001 · 6 claims · 4 setups
MD-CFLP can detect DNA sequence alterations (single-base substitutions, small insertions/deletions) in BRCA1 fragments longer than 1 kb
-
Full-text index only
Genome annotation errors in pathway databases due to semantic ambiguity in partial EC numbers.
PMID 16034025 · PMC1179732 · Nucleic acids research · 2005 · 7 claims · 4 setups
Partial EC numbers are semantically ambiguous, and databases that assign a gene to all reactions sharing the same partial EC number make a faulty inference, causing systematic misannotation.
-
Full-text index only
High-throughput chromatin information enables accurate tissue-specific prediction of transcription factor binding sites.
PMID 18988630 · PMC2662491 · Nucleic acids research · 2009 · 8 claims · 8 setups
Incorporating H3K4me3 chromatin modification estimates greatly improves the accuracy of in silico prediction of in vivo TF binding for a wide range of TFs in human and mouse
-
Full-text index only
Genome-wide prioritization of disease genes and identification of disease-disease associations from an integrated human functional linkage network.
PMID 19728866 · PMC2768980 · Genome biology · 2009 · 6 claims · 6 setups
Integrating 16 genomic features (32 sub-features) via a naïve Bayes classifier produces a genome-scale FLN of 21,657 human genes and 22,388,609 weighted links that outperforms any individual data source for inferring functional linkages.
-
Full-text index only
DoBSeqWF: a framework for sensitive detection of individual genetic variation in pooled sequencing data.
PMID 41704565 · PMC12907731 · NAR genomics and bioinformatics · 2026 · 7 claims · 5 setups
DoBSeqWF, a Nextflow-based pipeline, processes pooled DoBSeq sequencing data through alignment, variant calling, machine-learning-based filtering, and variant pinpointing/assignment to individuals.
-
Full-text index only
Integrated multi-omic atlas reveals the hierarchy of spatiotemporal regulatory networks of mouse gastrulation.
PMID 41526381 · PMC12902073 · Nature communications · 2026 · 8 claims · 8 setups
BioCRE, a novel bi-orientation regression algorithm, more accurately links genes to candidate cis-regulatory elements (CREs) than existing tools Signac and ArchR
-
Has reproduction · 80
DMN-seq enriches DNA hypomethylated regions for biomarker discovery using 5-methylcytosine glycosylase.
PMID 41673887 · PMC13097799 · Genome biology · 2026 · 8 claims · 9 setups
DMN-seq (DMN+) uses DME to nick DNA specifically at 5mC sites, enabling 5mC detection at single-base resolution via selective adaptor ligation
-
Has reproduction · 85
High performance imputation of structural and single nucleotide variants using low-coverage whole genome sequencing.
PMID 40155798 · PMC11951665 · Genetics, selection, evolution : GSE · 2025 · 7 claims · 6 setups
SNVs are imputed with high accuracy and recall across all tested WGS depths (1-4x), including in samples external to the reference panel.
-
Has reproduction · 84
Foster thy young: enhanced prediction of orphan genes in assembled genomes.
PMID 34928390 · PMC9023268 · Nucleic acids research · 2022 · 8 claims · 6 setups
Each of the five tested gene prediction pipelines under-predicts orphan genes, as few as 11% detected under one scenario
-
Full-text index only
Discovery of protein-protein interactions using a combination of linguistic, statistical and graphical information.
PMID 15941473 · PMC1164402 · BMC bioinformatics · 2005 · 8 claims · 5 setups
A combined linguistic+statistical+rule-based method achieves precision 0.61 and recall 0.97 (f=0.74) detecting yeast protein-protein interactions across 12,300 Medline abstracts.
-
Full-text index only
Information-based methods for predicting gene function from systematic gene knock-downs.
PMID 18959798 · PMC2596148 · BMC bioinformatics · 2008 · 8 claims · 4 setups
Information-based metrics, which incorporate a phenotype's genomic frequency, outperform non-information-based metrics for detecting gene-gene functional similarity from phenotypic knock-down profiles.
-
Has reproduction · 58
A comparative study of techniques for differential expression analysis on RNA-Seq data.
PMID 25119138 · PMC4132098 · PloS one · 2014 · 8 claims · 8 setups
edgeR performs slightly better than DESeq and Cuffdiff2 in terms of the ability to uncover true positives.
-
Has reproduction · 83
Accurate prediction of metagenome-assembled genome completeness by MAGISTA, a random forest model built on alignment-free intra-bin statistics.
PMID 35248155 · PMC8898458 · Environmental microbiome · 2022 · 7 claims · 7 setups
MAGISTA, a random forest model built on alignment-free intra-bin distance-distribution statistics, can estimate MAG completeness and purity without relying on reference marker genes.
-
Has reproduction · 50
MEDUSA: A Pipeline for Sensitive Taxonomic Classification and Flexible Functional Annotation of Metagenomic Shotgun Sequences.
PMID 35330728 · PMC8940201 · Frontiers in genetics · 2022 · 7 claims · 6 setups
MEDUSA correctly identifies more species than MEGAN 6 CE, especially less abundant species.
-
Full-text index only
Allelic drop-out may occur with a primer binding site polymorphism for the commonly used RFLP assay for the -1131T>C polymorphism of the Apolipoprotein AV gene.
PMID 16670016 · PMC1513378 · Lipids in health and disease · 2006 · 8 claims · 6 setups
A -987C>T polymorphism located 4bp from the 3' end of the MseI RFLP forward primer causes allelic drop-out, producing incorrect -1131T>C genotypes.