Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 95
nf-rnaSeqCount: A Nextflow pipeline for obtaining raw read counts from RNA-seq data.
PMID 35574063 · PMC9097006 · South African computer journal = Suid-Afrikaanse rekenaartydskrif · 2021 · 7 claims · 5 setups
nf-rnaSeqCount is a portable, reproducible Nextflow pipeline that maps RNA-seq reads to a reference genome and quantifies gene abundance for differential expression analysis
-
Full-text index only
Predicting the phenotypic effects of non-synonymous single nucleotide polymorphisms based on support vector machines.
PMID 18005451 · PMC2216041 · BMC bioinformatics · 2007 · 8 claims · 5 setups
Parepro, an SVM-based method integrating three attribute sets (RD, MI, IE) derived from evolutionary and residue-property information, predicts whether an nsSNP is deleterious or neutral.
-
Has reproduction · 90
Cell Specific eQTL Analysis without Sorting Cells.
PMID 25955312 · PMC4425538 · PLoS genetics · 2015 · 7 claims · 4 setups
A genotype x predicted-cell-count interaction (GxE) meta-analysis across whole blood datasets can detect neutrophil-specific cis-eQTLs without cell sorting
-
Has reproduction · 42
KAGE: fast alignment-free graph-based genotyping of SNPs and short indels.
PMID 36195962 · PMC9531401 · Genome biology · 2022 · 7 claims · 7 setups
KAGE combines population-based kmer count modeling with single-variant prior adjustment into an alignment-free genotyper that matches the accuracy of the best existing alignment-free genotypers while being an order of magnitude faster.
-
Full-text index only
Boosting accuracy of automated classification of fluorescence microscope images for location proteomics.
PMID 15207009 · PMC449699 · BMC bioinformatics · 2004 · 8 claims · 8 setups
New classifiers (SVMs, ensembles) and new wavelet-derived (Gabor, Daubechies) features improve recognition of protein subcellular location patterns over the previous neural network approach
-
Full-text index only
A novel deep learning-driven framework for improving lncRNA comprehensive annotation with LncADeep 2.0.
PMID 41923359 · PMC13090826 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 8 setups
LncADeep 2.0 outperforms LncADeep and other existing tools for lncRNA identification on both GENCODE annotated transcripts and independent RNA-seq data
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.
-
Full-text index only
Conserved elements with potential to form polymorphic G-quadruplex structures in the first intron of human genes.
PMID 18187510 · PMC2275096 · Nucleic acids research · 2008 · 8 claims · 6 setups
G-richness downstream of the TSS is strand-biased, concentrated on the nontemplate strand, with a peak at +200 to +300 bp
-
Full-text index only
Local combinational variables: an approach used in DNA-binding helix-turn-helix motif prediction with sequence information.
PMID 19651875 · PMC2761287 · Nucleic acids research · 2009 · 8 claims · 7 setups
The LCV approach predicts HTH motifs with 93.29% accuracy, 93.93% sensitivity and 92.66% specificity using only primary sequence information
-
Full-text index only
Dynamic and Ongoing De Novo L1 Retrotransposition Contributes to Genome Plasticity and Intrapatient Heterogeneity in Ovarian Cancer.
PMID 41223332 · PMC13055634 · Cancer research · 2026 · 8 claims · 5 setups
HGSC tumors show high inter-patient heterogeneity in total de novo L1 insertion burden.
-
Full-text index only
FGF: a web tool for Fishing Gene Family in a whole genome database.
PMID 17584790 · PMC1933194 · Nucleic acids research · 2007 · 6 claims · 3 setups
FGF efficiently searches and identifies gene families in whole-genome databases and outputs visual phylogenetic trees annotated with gene structure, chromosome position, duplication fate, and selective pressure (Ka/Ks)
-
Has reproduction · 50
Integrated drug resistance and leukemic stemness gene-expression scores predict outcomes in large cohort of over 3500 AML patients from 10 trials.
PMID 39090192 · PMC11294346 · NPJ precision oncology · 2024 · 7 claims · 6 setups
A 5-gene ADE-Resistance Score (ADE-RS5), derived via LASSO regression from 67 pharmacologically relevant genes, predicts MRD positivity, EFS and OS in pediatric AML.
-
Has reproduction · 38
RNA-Seq transcriptome profiling of upland cotton (Gossypium hirsutum L.) root tissue under water-deficit stress.
PMID 24324815 · PMC3855774 · PloS one · 2013 · 8 claims · 8 setups
A total of 1,530 transcripts were differentially expressed between well-watered and water-deficit stressed field-grown upland cotton root tissues (913 up-regulated, 617 down-regulated).
-
Full-text index only
StrainMake: reproducible hybrid metagenomics with MAG recovery and strain-level resolution.
PMID 42097292 · PMC13188985 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 5 setups
StrainMake is a Snakemake-based, Conda-managed workflow for de novo metagenomic analysis from short, long, or hybrid sequencing data.
-
Full-text index only
scDenorm: a denormalization tool for integrating single-cell transcriptomics data.
PMID 41915012 · PMC13142155 · GigaScience · 2026 · 8 claims · 7 setups
Inconsistent delta-method normalization across datasets introduces biases (e.g., B-cell separation) that persist even after integration with Harmony, scanorama, or BBKNN.
-
Full-text index only
AICellType: a large language model-based platform for accurate cell type annotation.
PMID 42001469 · PMC13092268 · Briefings in bioinformatics · 2026 · 8 claims · 8 setups
Claude 3.5 Sonnet achieved the best overall performance among 79 benchmarked LLMs for cell type annotation, balancing accuracy, robustness, speed, and cost-efficiency
-
Full-text index only
Bulk RNA-seq datasets analysis integration identifies robust drought-responsive genes and functional networks in Eucalyptus grandis.
PMID 42038403 · PMC13106539 · Frontiers in bioinformatics · 2026 · 7 claims · 7 setups
Meta-analysis integration of three independent RNA-seq drought studies identifies 472 robust differentially expressed genes (274 up, 198 down) that remain significant across the full meta-analysis and all leave-one-out Jackknife iterations
-
Full-text index only
The PeptideAtlas project.
PMID 16381952 · PMC1347403 · Nucleic acids research · 2006 · 8 claims · 5 setups
PeptideAtlas provides an automated repository that identifies peptides by MS/MS, statistically validates identifications, and maps them to eukaryotic genomes to enable data exchange and integration with genomic data.
-
Has reproduction · 71
Assessment tool based on fatty acid metabolic signatures for predicting the prognosis and treatment response in bladder cancer.
PMID 38076064 · PMC10703629 · Heliyon · 2023 · 8 claims · 8 setups
Consensus clustering of prognosis-related fatty acid metabolism genes (FAMGs) identifies three molecular subtypes of BLCA (FAMC1, FAMC2, FAMC3) with distinct prognoses and tumor microenvironments
-
Has reproduction · 78
Single duplex DNA sequencing with CODEC detects mutations with high sensitivity.
PMID 37106072 · PMC10181940 · Nature genetics · 2023 · 8 claims · 8 setups
CODEC concatenates both strands of an original DNA duplex into a single NGS read pair via an adapter quadruplex and strand-displacing extension, enabling single-duplex resolution