Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Information extraction from full text scientific articles: where are the keywords?
PMID 12775220 · PMC166134 · BMC bioinformatics · 2003 · 8 claims · 5 setups
The keyword content of the five article sections (A, I, M, R, D) is heterogeneous, i.e., different sections carry different kinds of information.
-
Has reproduction · 64
Celline: a flexible tool for one-step retrieval and integrative analysis of public single-cell RNA sequencing data.
PMID 41458999 · PMC12738925 · Frontiers in bioinformatics · 2025 · 8 claims · 6 setups
Celline is a Python package that automates the full scRNA-seq workflow (retrieval, metadata extraction, preprocessing, cell-type annotation, batch correction, trajectory inference) via single-line commands.
-
Has reproduction · 50
DeeReCT-APA: Prediction of Alternative Polyadenylation Site Usage Through Deep Learning.
PMID 33662629 · PMC9801043 · Genomics, proteomics & bioinformatics · 2022 · 8 claims · 8 setups
DeeReCT-APA quantitatively predicts the usage of all competing PASs of a gene simultaneously, rather than casting the problem as pairwise comparison like prior methods.
-
Has reproduction · 83
Multimodal data integration for biologically-relevant artificial intelligence to guide adjuvant chemotherapy in stage II colorectal cancer.
PMID 40472802 · PMC12171563 · EBioMedicine · 2025 · 6 claims · 7 setups
AI-derived radiological clustering identifies stage II CRC patients with significantly different survival benefit from adjuvant chemotherapy
-
Full-text index only
Improved mutation tagging with gene identifiers applied to membrane protein stability prediction.
PMID 19758467 · PMC2745585 · BMC bioinformatics · 2009 · 8 claims · 4 setups
MutationTagger achieves 87% F-measure for the mutation retrieval task on a benchmark dataset
-
Full-text index only
Extraction of human kinase mutations from literature, databases and genotyping studies.
PMID 19758464 · PMC2745582 · BMC bioinformatics · 2009 · 7 claims · 6 setups
A literature mining pipeline combining MutationFinder, false-positive filtering, and SVM-based classification can extract and disambiguate single-point mutation mentions from abstracts and full text
-
Full-text index only
Local combinational variables: an approach used in DNA-binding helix-turn-helix motif prediction with sequence information.
PMID 19651875 · PMC2761287 · Nucleic acids research · 2009 · 8 claims · 7 setups
The LCV approach predicts HTH motifs with 93.29% accuracy, 93.93% sensitivity and 92.66% specificity using only primary sequence information
-
Full-text index only
WBT-DC pipeline: a cross-cohort and cross-platform disease classification pipeline based on whole-blood transcriptomics.
PMID 42116144 · PMC13173921 · Journal of translational medicine · 2026 · 8 claims · 8 setups
WBT-DC integrates rank-based GSVA feature extraction with an ensemble random forest framework using cross-validation and hyperparameter optimization to mitigate batch effects.
-
Has reproduction
Artificial Intelligence Meets Whole Slide Images: Deep Learning Model Shapes an Immune-Hot Tumor and Guides Precision Therapy in Bladder Cancer.
PMID 36245985 · PMC9553530 · Journal of oncology · 2022 · 8 claims · 8 setups
A three-class WSI cluster (C0/C1/C2) derived via mini batch K-means clustering on Inception V3-extracted image features is significantly associated with overall survival and is an independent prognostic predictor in BLCA.
-
Full-text index only
Optimized library preparation, sequencing, and data analysis protocols for the generation of orbivirus consensus sequences.
PMID 41527034 · PMC12809950 · BMC genomics · 2026 · 8 claims · 8 setups
Optimized sample and library preparation protocols achieved comparable results to established methods while requiring simpler sample preparation.
-
Full-text index only
Classification of real and pseudo microRNA precursors using local structure-sequence features and support vector machine.
PMID 16381612 · PMC1360673 · BMC bioinformatics · 2005 · 7 claims · 7 setups
A 32-dimensional triplet structure-sequence feature vector combined with SVM (triplet-SVM) can distinguish real human pre-miRNAs from pseudo pre-miRNA hairpins with ~90% accuracy.
-
Full-text index only
ORFer--retrieval of protein sequences and open reading frames from GenBank and storage into relational databases or text files.
PMID 12493080 · PMC139979 · BMC bioinformatics · 2002 · 6 claims · 6 setups
ORFer retrieves protein and nucleic acid sequences and annotations from NCBI GenBank using the XML sequence format
-
Full-text index only
Genome assembly comparison identifies structural variants in the human genome.
PMID 17115057 · PMC2674632 · Nature genetics · 2006 · 7 claims · 7 setups
Genome assembly comparison is a robust approach for identifying all classes of genetic variation, with no lower size limit.
-
Full-text index only
Detection of atovaquone-proguanil resistance conferring mutations in Plasmodium falciparum cytochrome b gene in Luanda, Angola.
PMID 16597338 · PMC1513587 · Malaria journal · 2006 · 6 claims · 4 setups
No pfcytb mutations associated with atovaquone-proguanil treatment failure (codon 268 wild type, T802A, A803C) were found in the Luanda study population.
-
Full-text index only
Discovery and identification of potential biomarkers of papillary thyroid carcinoma.
PMID 19785722 · PMC2761863 · Molecular cancer · 2009 · 8 claims · 7 setups
A 3-peak (m/z 9190, 6631, 8697 Da) SVM classification model discriminates PTC from non-cancer controls with high sensitivity and specificity
-
Full-text index only
CaHoT-GRN: context-aware high-order topology learning for robust single-cell gene regulatory network inference.
PMID 42059479 · PMC13130071 · Briefings in bioinformatics · 2026 · 7 claims · 5 setups
CaHoT-GRN integrates pretrained biological language model embeddings (DNABERT for DNA, ESM for protein) with scRNA-seq expression data to improve GRN inference
-
Full-text index only
Identification of novel prognostic markers in cervical intraepithelial neoplasia using LDMAS (LOH Data Management and Analysis Software).
PMID 15673474 · PMC548130 · BMC bioinformatics · 2005 · 8 claims · 3 setups
LDMAS software integrates LOH molecular data with clinico-pathological data for prognostic marker discovery
-
Has reproduction · 100
Smart spatial omics (S2-omics) optimizes region of interest selection to capture molecular heterogeneity in diverse tissues.
PMID 41298871 · PMC12662399 · Nature cell biology · 2025 · 7 claims · 6 setups
S2-omics is an end-to-end workflow that automatically selects ROIs from H&E histology images to maximize molecular information content for spatial omics profiling.
-
Full-text index only
A high throughput method for genome-wide analysis of retroviral integration.
PMID 17028098 · PMC1636494 · Nucleic acids research · 2006 · 8 claims · 8 setups
VITA uses MmeI to cleave DNA at a fixed distance from its recognition site, generating 21-22 bp genomic tags that serve as signatures of lentiviral integration sites.
-
Full-text index only
Radiogenomics predicts immune microenvironment heterogeneity and response to combination immunotherapy in hepatocellular carcinoma.
PMID 41555384 · PMC12895822 · Journal of translational medicine · 2026 · 8 claims · 8 setups
A parsimonious 2-gene immune-related signature (IRS = 0.425×KPNA2 + 0.302×SMG5) is significantly associated with immune heterogeneity and response to ICI plus anti-angiogenic combination therapy in HCC