Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Genome comparison without alignment using shortest unique substrings.
PMID 15910684 · PMC1166540 · BMC bioinformatics · 2005 · 8 claims · 8 setups
A number of sequence comparison tasks, including detection of unique genomic regions, can be accomplished efficiently without an alignment step using shortest unique substrings.
-
Full-text index only
Optimal step length EM algorithm (OSLEM) for the estimation of haplotype frequency and its application in lipoprotein lipase genotyping.
PMID 12529185 · PMC149347 · BMC bioinformatics · 2003 · 5 claims · 4 setups
OSLEM (Optimal Step Length EM), which approximates an optimal step length via a fixed-point search (D_N = D_{N-1} + λ(D_preN - D_{N-1})), runs about twice as fast as standard EM while producing the same haplotype frequency estimates.
-
Has reproduction · 71
RNAmountAlign: Efficient software for local, global, semiglobal pairwise and multiple RNA sequence/structure alignment.
PMID 31978147 · PMC6980424 · PloS one · 2020 · 7 claims · 6 setups
RNAmountAlign performs pairwise local, global, and semiglobal (query search) alignment and progressive multiple alignment (global and local) using incremental ensemble mountain height, running in O(n^3) time and O(n^2) space for two sequences of length n
-
Full-text index only
An investigation of polymorphisms in the 17q11.2-12 CC chemokine gene cluster for association with multiple sclerosis in Australians.
PMID 16872505 · PMC1550395 · BMC medical genetics · 2006 · 7 claims · 7 setups
Marginally significant (uncorrected) transmission distortion was found for four SNPs after stratification by HLA-DRB1*1501 status, disease course, or gender.
-
Has reproduction · 73
Vespucci: a system for building annotated databases of nascent transcripts.
PMID 24304890 · PMC3936758 · Nucleic acids research · 2014 · 8 claims · 7 setups
Existing ChIP-seq and RNA-seq analysis platforms (e.g. Cufflinks, peak callers) are unsuited to GRO-seq because they assume spliced/exonic reads, uniform density and paired-end data, and cannot identify transcriptional units de novo across the whole genome.
-
Full-text index only
Genotyping of genetically monomorphic bacteria: DNA sequencing in Mycobacterium tuberculosis highlights the limitations of current methodologies.
PMID 19915672 · PMC2772813 · PloS one · 2009 · 8 claims · 8 setups
MLSA of 89 genes across 108 global MTBC strains yields a single, highly robust phylogeny with virtually no homoplasy, congruent across parsimony, NJ, ML, and Bayesian methods.
-
Full-text index only
Implementation of a data repository-driven approach for targeted proteomics experiments by multiple reaction monitoring.
PMID 19121650 · PMC2706936 · Journal of proteomics · 2009 · 7 claims · 5 setups
A new MRM worksheet was implemented in The Global Proteome Machine database (GPMDB) that provides all information needed to design MRM transitions based solely on archived observations from previous experiments by other researchers.
-
Full-text index only
SW-ARRAY: a dynamic programming solution for the identification of copy-number changes in genomic DNA using array comparative genome hybridization data.
PMID 15961730 · PMC1151590 · Nucleic acids research · 2005 · 7 claims · 5 setups
SW-ARRAY, an adaptation of the Smith-Waterman dynamic programming algorithm, provides a sensitive and robust method for identifying copy-number changes in array CGH data
-
Full-text index only
MALDI profiling of human lung cancer subtypes.
PMID 19890392 · PMC2767501 · PloS one · 2009 · 8 claims · 8 setups
PIMAC/MALDI-TOF peptide profiles combined with classification models can distinguish normal lung from tumor and differentiate NSCLC histological subtypes
-
Full-text index only
Genomic divergences among cattle, dog and human estimated from large-scale alignments of genomic sequences.
PMID 16759380 · PMC1525190 · BMC genomics · 2006 · 8 claims · 6 setups
Overall pairwise genomic divergences among cattle, dog and human are relatively constant (0.32–0.37 change/site)
-
Has reproduction · 78
Widespread mono- and oligoadenylation direct small noncoding RNA maturation versus degradation fates.
PMID 41350938 · PMC12811392 · The EMBO journal · 2026 · 8 claims · 2 setups
Post-transcriptional adenylation of newly transcribed sncRNAs is widespread and occurs in two distinct varieties: transient oligoadenylation and stable monoadenylation
-
Has reproduction · 90
Systematic clustering algorithm for chromatin accessibility data and its application to hematopoietic cells.
PMID 33253153 · PMC7728210 · PLoS computational biology · 2020 · 7 claims · 5 setups
A systematic clustering algorithm for ATAC-seq data can be built by binarizing the genome into open/closed chromatin (1/0) strings and computing Hamming distances between samples for hierarchical clustering.
-
Has reproduction · 89
Comparative genomics of dairy-associated Staphylococcus aureus from selected sub-Saharan African regions reveals milk as reservoir for human-and animal-derived strains and identifies a putative animal-related clade with presumptive novel siderophore.
PMID 36046020 · PMC9421002 · Frontiers in microbiology · 2022 · 7 claims · 8 setups
Milk serves as a reservoir for both human- and animal-derived S. aureus strains in sub-Saharan Africa
-
Has reproduction · 78
A network-based model of Aspergillus fumigatus elucidates regulators of development and defensive natural products of an opportunistic pathogen.
PMID 41505094 · PMC12781895 · Nucleic acids research · 2026 · 7 claims · 6 setups
MERLIN-P-TFA network inference on 18 curated public RNA-seq datasets produced a genome-wide GRN resource for A. fumigatus called GRAsp.
-
Has reproduction · 30
Minimal metabolic pathway structure is consistent with associated biomolecular interactions.
PMID 24987116 · PMC4299494 · Molecular systems biology · 2014 · 8 claims · 8 setups
MinSpan, a mixed-integer linear optimization algorithm, computes the shortest, linearly independent pathways (sparsest basis of the null space of the stoichiometric matrix S) for genome-scale metabolic networks, which convex approaches (extreme pathways, elementary flux modes) cannot do at genome scale.
-
Has reproduction · 87
A target enrichment method for gathering phylogenetic information from hundreds of loci: An example from the Compositae.
PMID 25202605 · PMC4103609 · Applications in plant sciences · 2014 · 8 claims · 8 setups
A custom sequence capture probe set (9678 baits targeting 1061 orthologous genes) was designed to enrich COS loci across the Compositae.
-
Full-text index only
Malaria genomics meets drug-resistance phenotyping in the field.
PMID 19664295 · PMC2745762 · Genome biology · 2009 · 8 claims · 8 setups
Significantly longer parasite clearance times in Pailin (western Cambodia) versus Wang Pha (Thailand-Myanmar border) establish the presence of artemisinin resistance in western Cambodia.
-
Full-text index only
Insights into the coupling of duplication events and macroevolution from an age profile of animal transmembrane gene families.
PMID 16895434 · PMC1534073 · PLoS computational biology · 2006 · 8 claims · 7 setups
The density of transmembrane gene duplicates positively correlates with the estimated maximum number of cell types of common ancestors
-
Full-text index only
Statistical challenges in preprocessing in microarray experiments in cancer.
PMID 18829474 · PMC3529914 · Clinical cancer research : an official journal of the American Association for Cancer Research · 2008 · 8 claims · 7 setups
Choice of pre-processing method materially changes which features are found significantly associated with survival in the Beer et al. lung cancer microarray dataset
-
Full-text index only
Development of an integrated genome informatics, data management and workflow infrastructure: a toolbox for the study of complex disease genetics.
PMID 15601538 · PMC3525068 · Human genomics · 2004 · 8 claims · 8 setups
An integrated system combining Ensembl, ACeDB, Gbrowse and custom relational databases provides a scalable genome informatics and workflow infrastructure for complex disease gene discovery.