Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 76
Correcting scale distortion in RNA sequencing data.
PMID 39875825 · PMC11776150 · BMC bioinformatics · 2025 · 8 claims · 8 setups
Local averaging reveals expression-level-dependent biases that differ from sample to sample across all RNA-seq datasets studied, and are not corrected by conventional normalization (TPM/FPKM)
-
Full-text index only
Classification of real and pseudo microRNA precursors using local structure-sequence features and support vector machine.
PMID 16381612 · PMC1360673 · BMC bioinformatics · 2005 · 7 claims · 7 setups
A 32-dimensional triplet structure-sequence feature vector combined with SVM (triplet-SVM) can distinguish real human pre-miRNAs from pseudo pre-miRNA hairpins with ~90% accuracy.
-
Full-text index only
Analysis of protein sequence and interaction data for candidate disease gene prediction.
PMID 17020920 · PMC1636487 · Nucleic acids research · 2006 · 8 claims · 7 setups
Combining CPS and CMP using known disease genes as input achieves sensitivity 0.52 and specificity 0.97, reducing candidate lists 13-fold
-
Full-text index only
Identification of diagnostic markers for tuberculosis by proteomic fingerprinting of serum.
PMID 16980117 · PMC7159276 · Lancet (London, England) · 2006 · 8 claims · 5 setups
An SVM classifier trained on serum proteomic profiles discriminated patients with active tuberculosis from controls with clinically overlapping conditions
-
Full-text index only
Detection of venous thromboembolism by proteomic serum biomarkers.
PMID 17579716 · PMC1891085 · PloS one · 2007 · 5 claims · 8 setups
A neural network-based classifier built from direct MALDI-TOF MS serum protein expression profiles can diagnose VTE with sensitivity/specificity that exceeds D-dimer assays
-
Full-text index only
Usefulness of cancer-testis antigens as biomarkers for the diagnosis and treatment of hepatocellular carcinoma.
PMID 17244360 · PMC1797003 · Journal of translational medicine · 2007 · 8 claims · 8 setups
HCC is a highly heterogeneous, non-linear disease driven by complex, multi-stage genetic and environmental alterations
-
Full-text index only
Pol II promoter prediction using characteristic 4-mer motifs: a machine learning approach.
PMID 18834544 · PMC2575220 · BMC bioinformatics · 2008 · 8 claims · 8 setups
128 discriminating 4-mer motifs combined with an SVM (RBF kernel, LIBSVM) can distinguish promoter from non-promoter DNA sequences
-
Full-text index only
The cancer secretome: a reservoir of biomarkers.
PMID 18796163 · PMC2562990 · Journal of translational medicine · 2008 · 8 claims · 8 setups
Cancer secretome analysis is a promising reservoir for identifying novel, non-invasive cancer biomarkers, addressing limitations of whole blood/serum proteomics
-
Full-text index only
Design and analysis issues in genome-wide somatic mutation studies of cancer.
PMID 18692126 · PMC2820387 · Genomics · 2009 · 6 claims · 4 setups
Two-stage (discovery + validation) sequencing designs efficiently allocate resources and can produce highly informative candidate driver gene lists even with relatively small sample sizes.
-
Full-text index only
Development of proteomic patterns for detecting lung cancer.
PMID 14757945 · PMC3851077 · Disease markers · 2003 · 8 claims · 3 setups
A decision tree classification algorithm built on three serum protein mass peaks (8122Da, 1452Da, 1610Da) can discriminate lung cancer patients from healthy controls
-
Full-text index only
Limitations in SELDI-TOF MS whole serum proteomic profiling with IMAC surface to specifically detect colorectal cancer.
PMID 19689818 · PMC2743709 · BMC cancer · 2009 · 7 claims · 3 setups
The previously reported classifier (m/z 8,132 and 4,002) failed to discriminate CRC patients from healthy volunteers in this independent validation cohort
-
Has reproduction
Unlocking the microbial studies through computational approaches: how far have we reached?
PMID 36920617 · PMC10016191 · Environmental science and pollution research international · 2023 · 8 claims · 8 setups
Metagenomics enables culture-independent study of microbial communities directly from their natural environments, bypassing the need for clonal isolation.
-
Has reproduction · 78
GenTB: A user-friendly genome-based predictor for tuberculosis resistance powered by machine learning.
PMID 34461978 · PMC8407037 · Genome medicine · 2021 · 8 claims · 6 setups
GenTB is a free, open, web-based application offering two ML predictors (Random Forest and WDNN) that predict resistance to 13 and 10 anti-TB drugs, respectively.
-
Full-text index only
Diagnostic proteomics: serum proteomic patterns for the detection of early stage cancers.
PMID 15258335 · PMC3851082 · Disease markers · 2003 · 8 claims · 8 setups
Proteomic pattern analysis of serum mass spectra, without identifying the underlying proteins, can distinguish cancer patients from healthy controls with high sensitivity and specificity.
-
Full-text index only
JIGSAW, GeneZilla, and GlimmerHMM: puzzling out the features of human genes in the ENCODE regions.
PMID 16925843 · PMC1810558 · Genome biology · 2006 · 8 claims · 4 setups
Adding model states for specific biological features (signal peptides, CpG islands, etc.) to non-comparative GHMM gene finders did little or nothing to enhance predictive accuracy, sometimes reducing it.
-
Full-text index only
Exogean: a framework for annotating protein-coding genes in eukaryotic genomic DNA.
PMID 16925841 · PMC1810556 · Genome biology · 2006 · 8 claims · 5 setups
Exogean is a framework using directed acyclic coloured multigraphs (DACMs) to represent biological objects (mRNA, ESTs, protein alignments, exons) and iteratively combine them into complex protein-coding transcript models.
-
Full-text index only
AceView: a comprehensive cDNA-supported gene and transcripts annotation.
PMID 16925834 · PMC1810549 · Genome biology · 2006 · 8 claims · 4 setups
At the mRNA level, AceView transcripts are the closest match to Gencode transcripts among all evaluated methods, including alternative splice variants
-
Full-text index only
Using ESTs to improve the accuracy of de novo gene prediction.
PMID 16817966 · PMC1534067 · BMC bioinformatics · 2006 · 8 claims · 8 setups
TWINSCAN_EST combines EST alignments with TWINSCAN via a trainable 'ESTseq' representation and improves exact gene structure prediction accuracy on the whole C. elegans genome
-
Full-text index only
Identification of serum biomarkers for colon cancer by proteomic analysis.
PMID 16755300 · PMC2361335 · British journal of cancer · 2006 · 8 claims · 8 setups
Complement C3a des-arg, α1-antitrypsin and transferrin were identified as serum proteins with diagnostic potential for CRC.
-
Full-text index only
High resolution melting for mutation scanning of TP53 exons 5-8.
PMID 17764544 · PMC2025602 · BMC cancer · 2007 · 7 claims · 5 setups
HRM is an effective technique for rapid scanning of TP53 mutations that can markedly reduce required sequencing