Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
The specificity and polymorphism of the MHC class I prevents the global adaptation of HIV-1 to the monomorphic proteasome and TAP.
PMID 18949050 · PMC2569417 · PloS one · 2008 · 6 claims · 5 setups
Within individual hosts, proteasome and TAP escape mutations in HIV-1 occur frequently
-
Full-text index only
Aggregation propensity of the human proteome.
PMID 18927604 · PMC2557143 · PLoS computational biology · 2008 · 8 claims · 7 setups
Long proteins have, on average, less intense/pronounced aggregation peaks than short proteins
-
Full-text index only
PlasmoDraft: a database of Plasmodium falciparum gene function predictions based on postgenomic data.
PMID 18925948 · PMC2605471 · BMC bioinformatics · 2008 · 8 claims · 4 setups
Gonna, a supervised k-nearest-neighbor Guilt-By-Association predictor, proposes GO annotations for a gene based on similarity of its transcriptome, proteome, or interactome profile to genes already annotated by GeneDB
-
Full-text index only
Improving the specificity of exon prediction using comparative genomics.
PMID 18831778 · PMC2559877 · BMC genomics · 2008 · 8 claims · 6 setups
A log-odds ratio scoring method based on codon conservation across human-mouse/human-dog alignments and adjacent-codon dependency can classify putative exons as coding vs non-coding.
-
Full-text index only
Zebrafish whole-adult-organism chemogenomics for large-scale predictive and discovery chemical biology.
PMID 18618001 · PMC2442223 · PLoS genetics · 2008 · 8 claims · 6 setups
Zebrafish whole-adult-organism chemogenomics generates robust prediction models that discriminate P(H)AHs from ECs across independent experiments
-
Full-text index only
An efficient method for the prediction of deleterious multiple-point mutations in the secondary structure of RNAs using suboptimal folding solutions.
PMID 18445289 · PMC2386494 · BMC bioinformatics · 2008 · 8 claims · 6 setups
Using RNAsubopt suboptimal solutions computed once for the wild-type sequence, specific multiple-point mutations likely to cause conformational rearrangement can be selected without brute-force enumeration.
-
Full-text index only
Integrated reiterative pipeline for rapid epitope-based pan-alphavirus vaccines.
PMID 41811958 · PMC12978219 · Science advances · 2026 · 6 claims · 7 setups
An integrated pipeline combining ML-based epitope prediction (netMHCpan, EpiDope, BepiPred), TCRpMHC structural modeling/docking, and JessEV vaccine design can prioritize viral peptides by immunogenicity, allele coverage, solubility, and stability.
-
Full-text index only
Separating selection from mutation in antibody language models.
PMID 41944291 · PMC13056363 · eLife · 2026 · 8 claims · 6 setups
Masked antibody language models such as AbLang2 are biased by nucleotide-level mutation processes (germline memorization, codon table, SHM rate variation)
-
Full-text index only
Critical evaluation of drug response prediction models with DrEval.
PMID 42120410 · PMC13168506 · Nature communications · 2026 · 8 claims · 6 setups
DrEval is a living open-source benchmarking pipeline for unbiased, biologically meaningful evaluation of cancer drug response prediction models, integrating standardized preprocessing, hyperparameter tuning, statistically rigorous evaluation, cross-study benchmarks, and ablation studies.
-
Full-text index only
Uncertainty-aware genomic deep learning with knowledge distillation.
PMID 41523993 · PMC12779563 · NPJ artificial intelligence · 2026 · 7 claims · 6 setups
DEGU distills an ensemble of teacher DNNs into a single student model by jointly predicting the ensemble mean and the variability (epistemic uncertainty) across ensemble predictions.
-
Full-text index only
Learning collective multicellular dynamics with an interacting mean field neural SDE model.
PMID 41564117 · PMC12854464 · PLoS computational biology · 2026 · 7 claims · 5 setups
scIMF models multicellular dynamics as interacting diffusion processes using a McKean-Vlasov SDE solved via Neural SDE, with a Transformer-based cell-wise attention mechanism approximating the distribution-dependent drift term
-
Full-text index only
Development of lead hammerhead ribozyme candidates against human rod opsin mRNA for retinal degeneration therapy.
PMID 19094986 · PMC3388947 · Experimental eye research · 2009 · 8 claims · 3 setups
Three lead hhRz candidates (CUC↓266, CUC↓1411, AUA↓1414) significantly knock down human RHO protein expression relative to control (p<0.05)
-
Full-text index only
SNAP predicts effect of mutations on protein function.
PMID 18757876 · PMC2562009 · Bioinformatics (Oxford, England) · 2008 · 8 claims · 3 setups
SNAP is a publicly available web-server implementation predicting functional effects (neutral/non-neutral) of single amino acid substitutions.
-
Full-text index only
ASPIC: a web resource for alternative splicing prediction and transcript isoforms characterization.
PMID 16845044 · PMC1538898 · Nucleic acids research · 2006 · 8 claims · 2 setups
The ASPIC algorithm, using an optimization procedure that minimizes splice site predictions and transcript isoforms from multiple EST-genome alignments, outperforms other similar AS-prediction tools in sensitivity and selectivity
-
Has reproduction · 81
SEMdag: Fast learning of Directed Acyclic Graphs via node or layer ordering.
PMID 39775401 · PMC11709272 · PloS one · 2025 · 8 claims · 5 setups
SEMdag() is a two-step order-based algorithm for fast learning of high-dimensional linear SEMs, using knowledge-based (KB) or data-driven bottom-up (BU) node/layer ordering followed by penalized (L1) DAG estimation
-
Has reproduction · 95
Mouse-Geneformer: A deep learning model for mouse single-cell transcriptome and its cross-species utility.
PMID 40106407 · PMC11964219 · PLoS genetics · 2025 · 7 claims · 6 setups
Mouse-Geneformer, a Transformer Encoder model pre-trained via masked-token self-supervised learning on mouse-Genecorpus-20M, was successfully constructed following the original human Geneformer architecture.
-
Has reproduction · 90
The COMBAT-TB Workbench: Making Powerful Mycobacterium tuberculosis Bioinformatics Accessible.
PMID 35138128 · PMC8827006 · mSphere · 2022 · 8 claims · 5 setups
The COMBAT-TB Workbench combines the IRIDA web platform and the Galaxy workflow platform into a single easy-to-install, Docker-based application
-
Has reproduction · 90
Streaming Long-Read Sequence Alignments for HLA Predictions Using HLAminer.
PMID 40145684 · PMC11948951 · Current protocols · 2025 · 6 claims · 5 setups
Streaming minimap2 alignment output directly into HLAminer via Unix pipe (-a stream mode) enables HLA class I and II allele prediction without storing bulky SAM alignment files on disk
-
Full-text index only
Analysis of protein sequence and interaction data for candidate disease gene prediction.
PMID 17020920 · PMC1636487 · Nucleic acids research · 2006 · 8 claims · 7 setups
Combining CPS and CMP using known disease genes as input achieves sensitivity 0.52 and specificity 0.97, reducing candidate lists 13-fold
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs