Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 87
A robust data scaling algorithm to improve classification accuracies in biomedical data.
PMID 27612635 · PMC5016890 · BMC bioinformatics · 2016 · 8 claims · 2 setups
Models trained on data scaled by the GL algorithm outperform models trained on data scaled by the Min-max or Z-score algorithms across 16 binary classification tasks, measured by AUROC and percentage of correct classification
-
Full-text index only
Uncovering information on expression of natural antisense transcripts in Affymetrix MOE430 datasets.
PMID 17598913 · PMC1929078 · BMC genomics · 2007 · 8 claims · 4 setups
Standard Affymetrix expression GeneChips (MOE430, HG-U133) contain probe sets that detect natural antisense transcripts (NATs)
-
Full-text index only
Optimality driven nearest centroid classification from genomic data.
PMID 17912341 · PMC1991588 · PloS one · 2007 · 7 claims · 5 setups
A theoretical result determines the subset of features of a given size that minimizes the misclassification rate for a nearest-centroid (LDA) classifier, based on equation (4).
-
Full-text index only
Development of proteomic patterns for detecting lung cancer.
PMID 14757945 · PMC3851077 · Disease markers · 2003 · 8 claims · 3 setups
A decision tree classification algorithm built on three serum protein mass peaks (8122Da, 1452Da, 1610Da) can discriminate lung cancer patients from healthy controls
-
Has reproduction · 68
Cell-type annotation with accurate unseen cell-type identification using multiple references.
PMID 37379341 · PMC10335708 · PLoS computational biology · 2023 · 8 claims · 4 setups
mtANN integrates multiple reference datasets and eight gene selection methods via ensemble learning (multiple deep classification models + majority voting) to improve cell-type annotation accuracy
-
Has reproduction · 86
The selection of software and database for metagenomics sequence analysis impacts the outcome of microbial profiling and pathogen detection.
PMID 37027361 · PMC10081788 · PloS one · 2023 · 7 claims · 7 setups
Obtaining an accurate species-level microbial profile using current direct-read metagenomics profiling software is still a challenging task.
-
Has reproduction · 83
ConNIS and labeling instability: New statistical methods for improving the detection of essential genes in TraDIS libraries.
PMID 41790830 · PMC12991369 · PLoS computational biology · 2026 · 8 claims · 3 setups
ConNIS provides an analytic solution for the probability of observing the longest insertion-free sequence within a gene given its length and number of insertion sites under non-essentiality.
-
Has reproduction · 100
Intratumoral heterogeneity in microsatellite instability status at single-cell resolution.
PMID 41767255 · PMC12936829 · iScience · 2026 · 8 claims · 8 setups
MSI status can be heterogeneous at the single-cell level within a tumor, challenging its use as a binary biomarker
-
Full-text index only
Diagnostic proteomics: serum proteomic patterns for the detection of early stage cancers.
PMID 15258335 · PMC3851082 · Disease markers · 2003 · 8 claims · 8 setups
Proteomic pattern analysis of serum mass spectra, without identifying the underlying proteins, can distinguish cancer patients from healthy controls with high sensitivity and specificity.
-
Full-text index only
Genomics, molecular imaging, bioinformatics, and bio-nano-info integration are synergistic components of translational medicine and personalized healthcare research.
PMID 18831773 · PMC3226104 · BMC genomics · 2008 · 8 claims · 8 setups
Genomics, molecular imaging, bioinformatics, and bio-nano-info integration are synergistic components of translational medicine and personalized healthcare
-
Full-text index only
CLEAN: CLustering Enrichment ANalysis.
PMID 19640299 · PMC2734555 · BMC bioinformatics · 2009 · 8 claims · 4 setups
The gene-specific CLEAN score improves reproducibility of cluster analysis conclusions across independent datasets compared to the traditional cluster-wide score (cwCLEAN).
-
Has reproduction · 83
Accurate prediction of metagenome-assembled genome completeness by MAGISTA, a random forest model built on alignment-free intra-bin statistics.
PMID 35248155 · PMC8898458 · Environmental microbiome · 2022 · 7 claims · 7 setups
MAGISTA, a random forest model built on alignment-free intra-bin distance-distribution statistics, can estimate MAG completeness and purity without relying on reference marker genes.
-
Full-text index only
Mitochondrial diversity within modern human populations.
PMID 17439969 · PMC1888801 · Nucleic acids research · 2007 · 8 claims · 5 setups
Modern humans show extremely low divergence from the mitochondrial consensus sequence, differing on average by only 21.6 nucleotide sites
-
Full-text index only
Have microarrays failed to deliver for developmental biology?
PMID 12225576 · PMC139405 · Genome biology · 2002 · 8 claims · 8 setups
Despite predictions that microarrays would transform biology, very few published developmental biology microarray studies have generated novel insights.
-
Has reproduction · 91
Genome-wide identification of conserved and novel microRNAs in one bud and two tender leaves of tea plant (Camellia sinensis) by small RNA sequencing, microarray-based hybridization and genome survey scaffold sequences.
PMID 29157210 · PMC5697157 · BMC plant biology · 2017 · 7 claims · 8 setups
175 conserved and 83 novel miRNAs were identified mainly in one bud and two tender leaves of tea plant via small RNA sequencing combined with genome survey data