Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 57
Analysis and comprehensive comparison of PacBio and nanopore-based RNA sequencing of the Arabidopsis transcriptome.
PMID 32536962 · PMC7291481 · Plant methods · 2020 · 8 claims · 8 setups
ONT Pc produces higher raw data quality (higher alignment rate, lower error rate) than ONT Dc, while PacBio generates the longest reads
-
Full-text index only
A re-annotation pipeline for Illumina BeadArrays: improving the interpretation of gene expression data.
PMID 19923232 · PMC2817484 · Nucleic acids research · 2010 · 8 claims · 7 setups
A Perl-based pipeline that BLASTs/BLATs Illumina probe sequences against genomes and transcript databases (RefSeq, UCSC Known Genes, UniGene/GenBank, Ensembl) can classify probes by quality grade (Perfect/Good/Bad/No match) and is applicable across 8 BeadArray platforms and other array types
-
Has reproduction · 88
pwrEWAS: a user-friendly tool for comprehensive power estimation for epigenome wide association studies (EWAS).
PMID 31035919 · PMC6489300 · BMC bioinformatics · 2019 · 7 claims · 8 setups
pwrEWAS is a user-friendly tool (R package and Shiny web interface) for comprehensive power estimation in two-group EWAS using Illumina HumanMethylation BeadChip technology.
-
Full-text index only
Gene expression study on peripheral blood identifies progranulin mutations.
PMID 18551524 · PMC2773201 · Annals of neurology · 2008 · 7 claims · 3 setups
PGRN is highly expressed in peripheral blood (97th percentile of all array genes)
-
Has reproduction · 90
Assessment of genotyping array performance for genome-wide association studies and imputation in African cattle.
PMID 36057548 · PMC9441065 · Genetics, selection, evolution : GSE · 2022 · 7 claims · 6 setups
Commercially available bovine arrays are ineffective at capturing variants segregating among African indicine animals, with only 6% of high-LD (r2>0.8) variants captured by the best arrays versus 17% in African taurine and 25% in European taurine.
-
Full-text index only
Efficacy assessment of SNP sets for genome-wide disease association studies.
PMID 17726055 · PMC2034459 · Nucleic acids research · 2007 · 6 claims · 4 setups
τ, derived from Shannon entropy and swept radius ɛ, approximates the relative sample size efficiency of a marker set for mapping a causal variant at a given map position compared to a maximally polymorphic SNP
-
Has reproduction · 83
Accurate prediction of metagenome-assembled genome completeness by MAGISTA, a random forest model built on alignment-free intra-bin statistics.
PMID 35248155 · PMC8898458 · Environmental microbiome · 2022 · 7 claims · 7 setups
MAGISTA, a random forest model built on alignment-free intra-bin distance-distribution statistics, can estimate MAG completeness and purity without relying on reference marker genes.
-
Has reproduction · 87
A target enrichment method for gathering phylogenetic information from hundreds of loci: An example from the Compositae.
PMID 25202605 · PMC4103609 · Applications in plant sciences · 2014 · 8 claims · 8 setups
A custom sequence capture probe set (9678 baits targeting 1061 orthologous genes) was designed to enrich COS loci across the Compositae.
-
Has reproduction · 78
GenTB: A user-friendly genome-based predictor for tuberculosis resistance powered by machine learning.
PMID 34461978 · PMC8407037 · Genome medicine · 2021 · 8 claims · 6 setups
GenTB is a free, open, web-based application offering two ML predictors (Random Forest and WDNN) that predict resistance to 13 and 10 anti-TB drugs, respectively.
-
Has reproduction · 67
Optimal scaling of digital transcriptomes.
PMID 24223126 · PMC3819321 · PloS one · 2013 · 8 claims · 8 setups
Fifteen existing and novel transcript-count normalization algorithms can be compared with two novel, mutually independent metrics: the number of "uniform" genes (sufficiently low coefficient of variation after normalization) and low average Spearman correlation between normalized expression profiles of gene pairs.
-
Has reproduction · 100
Integrative transcriptome sequencing identifies trans-splicing events with important roles in human embryonic stem cell pluripotency.
PMID 24131564 · PMC3875859 · Genome research · 2014 · 8 claims · 8 setups
TSscan, a computational pipeline integrating long- and short-read transcriptome sequencing from multiple hESC lines, can detect trans-splicing while minimizing false positives from experimental artifacts and genetic rearrangements.
-
Has reproduction · 84
Expression Atlas update--a database of gene and transcript expression from microarray- and sequencing-based functional genomics experiments.
PMID 24304889 · PMC3964963 · Nucleic acids research · 2014 · 8 claims · 6 setups
Expression Atlas is a value-added database providing gene, protein and splice variant expression across cell types, organism parts, developmental stages, diseases and other biological/experimental conditions, built from manually curated high-quality microarray and RNA-sequencing experiments from ArrayExpress.
-
Has reproduction · 83
De novo identification of CD4(+) T cell epitopes.
PMID 38658646 · PMC11093748 · Nature methods · 2024 · 7 claims · 8 setups
SABR-IIs (chimeric receptors linking a covalently attached peptide-MHC-II to CD28-CD3ζ signaling domains) present epitopes to CD4+ T cells and induce a readable NFAT-GFP/CD69 signal upon cognate TCR recognition
-
Full-text index only
Personalized copy number and segmental duplication maps using next-generation sequencing.
PMID 19718026 · PMC2875196 · Nature genetics · 2009 · 5 claims · 5 setups
mrFAST maps short reads to all possible locations in the reference genome, enabling read-depth-based prediction of absolute copy number in both unique and duplicated sequence, including discrimination between highly identical gene paralogs.
-
Full-text index only
Variation resources at UC Santa Cruz.
PMID 17151077 · PMC1781230 · Nucleic acids research · 2007 · 8 claims · 8 setups
The UCSC Genome Browser variation resources integrate polymorphism data from public collections (dbSNP, HapMap, Affymetrix, Perlegen, SeattleSNPs) into a common format with additional annotations and genomic context.