Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Genome-wide microRNA profiling in human fetal nervous tissues by oligonucleotide microarray.
PMID 16983573 · PMC1705512 · Child's nervous system : ChNS : official journal of the International Society for Pediatric Neurosurgery · 2006 · 8 claims · 5 setups
72-83% of assayed miRNAs are expressed across human fetal organs, with G24w cerebrum showing the most miRNAs expressed
-
Has reproduction · 76
nf-core/circrna: a portable workflow for the quantification, miRNA target prediction and differential expression analysis of circular RNAs.
PMID 36694127 · PMC9875403 · BMC bioinformatics · 2023 · 8 claims · 4 setups
Existing circRNA workflows are limited: none delineate circRNA-miRNA interactions and only one performs differential expression analysis, requiring users to supplement missing analysis types with in-house expertise
-
Full-text index only
Recent additions and improvements to the Onto-Tools.
PMID 15980579 · PMC1160233 · Nucleic acids research · 2005 · 7 claims · 3 setups
The Onto-Tools back-end database was redesigned around the Entrez Gene data model after NCBI phased out LocusLink in February 2005.
-
Full-text index only
SNP-RFLPing: restriction enzyme mining for SNPs in genomes.
PMID 16503968 · PMC1386656 · BMC genomics · 2006 · 8 claims · 2 setups
SNP-RFLPing accepts three flexible input types (dbSNP rs#/ss# IDs, HUGO gene name/Entrez gene ID, or free-form SNP-in-sequence including IUPAC or [dNTP1/dNTP2] formats) for human, rat, and mouse genomes
-
Full-text index only
InParanoid 6: eukaryotic ortholog clusters with inparalogs.
PMID 18055500 · PMC2238924 · Nucleic acids research · 2008 · 8 claims · 3 setups
InParanoid 6 is an updated eukaryotic ortholog database covering 35 species (34 eukaryotes plus E. coli as outgroup), providing pairwise ortholog clusters with inparalogs for all species pairs.
-
Has reproduction · 56
Analysis of subcellular transcriptomes by RNA proximity labeling with Halo-seq.
PMID 34875090 · PMC8887463 · Nucleic acids research · 2022 · 6 claims · 8 setups
Halo-seq pairs a light-activatable Halo-DBF ligand with Click chemistry to label and purify spatially defined RNA populations in living cells with high spatial specificity (~100 nm radius)
-
Has reproduction · 50
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization.
PMID 41266599 · PMC12635123 · Communications biology · 2025 · 8 claims · 8 setups
BiRNA-BERT uses adaptive dual-tokenization that dynamically selects nucleotide-level (NUC) or byte-pair encoding (BPE) tokens based on input sequence length
-
Full-text index only
Increased DNA microarray hybridization specificity using sscDNA targets.
PMID 15847692 · PMC1090574 · BMC genomics · 2005 · 7 claims · 5 setups
A single round of ribo-SPIA amplification produces sufficient sscDNA for microarray hybridization from as little as 5 ng of starting total RNA
-
Full-text index only
SNPmasker: automatic masking of SNPs and repeats across eukaryotic genomes.
PMID 16845091 · PMC1538889 · Nucleic acids research · 2006 · 8 claims · 4 setups
SNPmasker is a web service combining SNP masking and repeat masking, supporting both coordinate-defined and homology-search-defined input regions, a combination not offered by prior tools
-
Has reproduction · 71
Sustainable data analysis with Snakemake.
PMID 34035898 · PMC8114187 · F1000Research · 2021 · 8 claims · 4 setups
Reproducibility alone is insufficient for sustainable data analysis; transparency and adaptability are equally important additional properties.
-
Has reproduction · 42
KAGE: fast alignment-free graph-based genotyping of SNPs and short indels.
PMID 36195962 · PMC9531401 · Genome biology · 2022 · 7 claims · 7 setups
KAGE combines population-based kmer count modeling with single-variant prior adjustment into an alignment-free genotyper that matches the accuracy of the best existing alignment-free genotypers while being an order of magnitude faster.
-
Has reproduction · 69
High-resolution transcriptome and genome-wide dynamics of RNA polymerase and NusA in Mycobacterium tuberculosis.
PMID 23222129 · PMC3553938 · Nucleic acids research · 2013 · 8 claims · 7 setups
NusA interacts with RNAP ubiquitously throughout the M. tuberculosis chromosome and its ChIP-seq profile mirrors RNAP distribution in both exponential and stationary phase, despite NusA not binding DNA directly.
-
Has reproduction · 80
Bisulfite sequencing of chromatin immunoprecipitated DNA (BisChIP-seq) directly informs methylation status of histone-modified DNA.
PMID 22466171 · PMC3371705 · Genome research · 2012 · 8 claims · 8 setups
BisChIP-seq — bisulfite sequencing of chromatin immunoprecipitated DNA — enables direct genome-wide, base-resolution interrogation of DNA methylation on histone-modified DNA molecules
-
Full-text index only
Seeded Bayesian Networks: constructing genetic networks from microarray data.
PMID 18601736 · PMC2474592 · BMC systems biology · 2008 · 8 claims · 4 setups
Seeding Bayesian Network analysis with prior networks derived from literature and/or PPI data improves recovery of known gene-gene interactions compared to BN analysis without a seed
-
Has reproduction · 80
SLDMS: A Tool for Calculating the Overlapping Regions of Sequences.
PMID 35046988 · PMC8761809 · Frontiers in plant science · 2021 · 8 claims · 5 setups
SLDMS is a novel method for computing overlapping regions of sequencing reads using suffix array (SA), longest common prefix (LCP) array, document array (DA), and a monotonic stack.
-
Has reproduction · 50
MoDLE: high-performance stochastic modeling of DNA loop extrusion interactions.
PMID 36451166 · PMC9710047 · Genome biology · 2022 · 7 claims · 6 setups
MoDLE is a high-performance stochastic model that simulates DNA-DNA contacts from loop extrusion genome-wide in minutes using less than 1 GB of RAM
-
Full-text index only
Boosting accuracy of automated classification of fluorescence microscope images for location proteomics.
PMID 15207009 · PMC449699 · BMC bioinformatics · 2004 · 8 claims · 8 setups
New classifiers (SVMs, ensembles) and new wavelet-derived (Gabor, Daubechies) features improve recognition of protein subcellular location patterns over the previous neural network approach
-
Full-text index only
GeneKeyDB: a lightweight, gene-centric, relational database to support data mining environments.
PMID 15790402 · PMC1274265 · BMC bioinformatics · 2005 · 8 claims · 6 setups
GeneKeyDB is a lightweight, gene-centric relational database that supports data mining and integration with computational analysis tools.
-
Full-text index only
Silhouette scores for assessment of SNP genotype clusters.
PMID 15760469 · PMC555759 · BMC genomics · 2005 · 7 claims · 5 setups
Silhouette scores provide a relevant, objective numeric measure of SNP genotype cluster quality, condensing tightness and separation into a single value from -1.0 to 1.0.
-
Full-text index only
Exhaustive prediction of disease susceptibility to coding base changes in the human genome.
PMID 18793467 · PMC2537574 · BMC bioinformatics · 2008 · 8 claims · 7 setups
Inter-species conservation is the strongest single predictor of disease-associated coding mutations among the factors tested.