Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 78
Machine learning and free energy clustering reveal PAH protein binding linked to AD risk.
PMID 41953002 · PMC13053772 · iScience · 2026 · 7 claims · 8 setups
An integrated framework of bioinformatics, machine learning, and ΔG clustering can prioritize PAHs for AD-associated neurotoxicity.
-
Full-text index only
Novel selection and genetic characterisation of an etoposide-resistant human leukaemic CCRF-CEM cell line.
PMID 8382508 · PMC1968246 · British journal of cancer · 1993 · 8 claims · 5 setups
CEM/VP-1 is 15-fold more resistant to etoposide than parental CCRF-CEM cells
-
Full-text index only
Improved tagging strategy for protein identification in mammalian cells.
PMID 16138932 · PMC1250225 · BMC genomics · 2005 · 7 claims · 7 setups
EGFP tagging via the artificial exon did not affect the subcellular localization of the tagged endogenous proteins
-
Has reproduction · 58
Genome-wide identification and characterization of germin-like protein family in Brassica juncea reveals their role against biotic stress.
PMID 41327045 · PMC12763953 · BMC plant biology · 2025 · 8 claims · 8 setups
102 GLPs were identified in B. juncea, 51 in B. nigra, and 48 in B. rapa via genome-wide in-silico analysis
-
Full-text index only
A survey of integral alpha-helical membrane proteins.
PMID 19760129 · PMC2780624 · Journal of structural and functional genomics · 2009 · 8 claims · 8 setups
An automated annotation pipeline defines the integral membrane genome and family associations for 21,379 proteins from 34 genomes, most belonging to 598 Pfam-derived membrane protein families.
-
Full-text index only
Towards the identification of essential genes using targeted genome sequencing and comparative analysis.
PMID 17052348 · PMC1624830 · BMC genomics · 2006 · 8 claims · 8 setups
Phyletic retention (ortholog presence across organisms) is the single most predictive feature of gene essentiality in both E. coli and S. cerevisiae.
-
Full-text index only
Proteomics of the human malaria parasite Plasmodium falciparum.
PMID 16445353 · PMC2721975 · Expert review of proteomics · 2006 · 8 claims · 8 setups
Completion of the P. falciparum genome sequence together with advances in mass spectrometry has enabled large-scale proteomic analysis of the parasite that was previously limited by inability to identify proteins from 2D gels
-
Full-text index only
Identification and characterization of insect-specific proteins by genome data analysis.
PMID 17407609 · PMC1852559 · BMC genomics · 2007 · 8 claims · 7 setups
Comparative genome analysis across five holometabolous insects and three non-insect eukaryotes (opisthokonts) identifies 154 insect-specific orthologous groups (refined to 51 proteins) and 466 eukaryote/opisthokont-core orthologous groups
-
Full-text index only
Systems integration of biodefense omics data for analysis of pathogen-host interactions and identification of potential targets.
PMID 19779614 · PMC2745575 · PloS one · 2009 · 8 claims · 8 setups
A protein-centric data integration approach (Master Protein Directory) enables integration and mining of heterogeneous pathogen-host omics data across multiple research centers
-
Full-text index only
Applications for protein sequence-function evolution data: mRNA/protein expression analysis and coding SNP scoring tools.
PMID 16912992 · PMC1538848 · Nucleic acids research · 2006 · 7 claims · 8 setups
PANTHER HMMs built from family/subfamily multiple sequence alignments can classify novel protein sequences into functional groups based on statistically significant HMM match scores
-
Full-text index only
Protein ranking by semi-supervised network propagation.
PMID 16723003 · PMC1810311 · BMC bioinformatics · 2006 · 8 claims · 5 setups
RankProp, a diffusion-based network propagation algorithm on a PSI-BLAST-derived protein similarity network, significantly outperforms local search methods (BLAST/PSI-BLAST) at detecting remote homologs.
-
Full-text index only
Structure SNP (StSNP): a web server for mapping and modeling nsSNPs on protein structures with linkage to metabolic pathways.
PMID 17537826 · PMC1933130 · Nucleic acids research · 2007 · 7 claims · 5 setups
StSNP integrates dbSNP, PDB, KEGG, and NCBI Entrez data into a single web server for nsSNP analysis
-
Full-text index only
Boosting accuracy of automated classification of fluorescence microscope images for location proteomics.
PMID 15207009 · PMC449699 · BMC bioinformatics · 2004 · 8 claims · 8 setups
New classifiers (SVMs, ensembles) and new wavelet-derived (Gabor, Daubechies) features improve recognition of protein subcellular location patterns over the previous neural network approach
-
Full-text index only
Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.
PMID 16790052 · PMC1534064 · BMC bioinformatics · 2006 · 8 claims · 7 setups
The Sequential Minimal Optimization (SMO) SVM algorithm was the best-performing classifier among 26 WEKA classifiers for predicting catalytic residues
-
Full-text index only
Generation of a restriction minus enteropathogenic Escherichia coli E2348/69 strain that is efficiently transformed with large, low copy plasmids.
PMID 18681975 · PMC2518929 · BMC microbiology · 2008 · 8 claims · 7 setups
E2348/69 possesses a type I restriction-modification system encoded by an hsdMSR-like operon identified by homology to known Hsd proteins.
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Full-text index only
Analysis of protein sequence and interaction data for candidate disease gene prediction.
PMID 17020920 · PMC1636487 · Nucleic acids research · 2006 · 8 claims · 7 setups
Combining CPS and CMP using known disease genes as input achieves sensitivity 0.52 and specificity 0.97, reducing candidate lists 13-fold
-
Full-text index only
Retroposition and evolution of the DNA-binding motifs of YY1, YY2 and REX1.
PMID 17478514 · PMC1904287 · Nucleic acids research · 2007 · 8 claims · 5 setups
62 YY1-related sequences were identified across genomes ranging from flying insects to humans, with high zinc finger domain conservation
-
Full-text index only
Genomics: applications in mechanism elucidation.
PMID 19166886 · PMC2698023 · Advanced drug delivery reviews · 2009 · 8 claims · 8 setups
Genomic tools require no a priori knowledge of a compound's mode of action and can reveal biological pathways (metabolism, distribution, off-target effects) in addition to the precise mechanism of action.
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes