Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Comparative gene finding in chicken indicates that we are closing in on the set of multi-exonic widely expressed human genes.
PMID 15809229 · PMC1074396 · Nucleic acids research · 2005 · 8 claims · 6 setups
Comparative gene finding (SGP2) between human and chicken, followed by RT-PCR verification, adds at most ~0.2% new genes to the multi-exonic human gene catalog
-
Full-text index only
Large-scale analysis of human alternative protein isoforms: pattern classification and correlation with subcellular localization signals.
PMID 15860772 · PMC1087780 · Nucleic acids research · 2005 · 8 claims · 8 setups
Constructed a large-scale dataset of 6876 human alternative protein isoforms from 2624 genes by combining H-Invitational full-length cDNA data and SwissProt VARSPLIC entries
-
Has reproduction
miRge3.0: a comprehensive microRNA and tRF sequencing analysis pipeline.
PMID 34308351 · PMC8294687 · NAR genomics and bioinformatics · 2021 · 6 claims · 5 setups
miRge3.0 with 12 CPUs consistently has the best execution speed compared to miRge2.0, Chimira and sRNAbench
-
Full-text index only
Functional importance of different patterns of correlation between adjacent cassette exons in human and mouse.
PMID 18439302 · PMC2432081 · BMC genomics · 2008 · 8 claims · 7 setups
Adjacent cassette exon pairs can be categorized by EST-derived correlation coefficient into three groups: mutually exclusive (ME, r<=-0.7), independent (IND, -0.2<=r<=0.2), and linked (LNK, r>=0.7)
-
Full-text index only
CorGen--measuring and generating long-range correlations for DNA sequence analysis.
PMID 16845099 · PMC1538783 · Nucleic acids research · 2006 · 8 claims · 3 setups
CorGen is a web server that measures long-range correlations in DNA sequences and generates random sequences with the same (or user-specified) correlation and composition parameters
-
Full-text index only
FeatureScan: revealing property-dependent similarity of nucleotide sequences.
PMID 16845077 · PMC1538849 · Nucleic acids research · 2006 · 6 claims · 5 setups
FeatureScan transforms nucleotide sequences into numerical signals of physico-chemical/conformational properties and compares them via a convolution/correlation (Fourier transform) method rather than comparing letters
-
Has reproduction · 65
SPEAQeasy: a scalable pipeline for expression analysis and quantification for R/bioconductor-powered RNA-seq analyses.
PMID 33932985 · PMC8088074 · BMC bioinformatics · 2021 · 8 claims · 5 setups
SPEAQeasy is a portable, easy-to-install, Nextflow-powered RNA-seq processing pipeline that lowers the computational entry barrier for biologists/clinicians
-
Full-text index only
Genome-wide identification of human functional DNA using a neutral indel model.
PMID 16410828 · PMC1326222 · PLoS computational biology · 2006 · 8 claims · 8 setups
A neutral indel model predicting a geometric distribution of intergap segment (IGS) lengths fits human-mouse ancestral repeat (AR) alignment data excellently
-
Full-text index only
Genome-wide analysis of human disease alleles reveals that their locations are correlated in paralogous proteins.
PMID 18989397 · PMC2565504 · PLoS computational biology · 2008 · 7 claims · 5 setups
The locations of sequence variants are correlated between paralogous human proteins more than expected by chance.
-
Full-text index only
IDEAL-Q, an automated tool for label-free quantitation analysis using an efficient peptide alignment approach and spectral data validation.
PMID 19752006 · PMC2808259 · Molecular & cellular proteomics : MCP · 2010 · 6 claims · 5 setups
IDEAL-Q predicts the elution time of peptides unidentified in a given LC-MS/MS run (but identified in others) using a computation-efficient linear regression plus fragmental refining function, avoiding costly whole-dataset pattern recognition
-
Full-text index only
Sequence similarity network reveals common ancestry of multidomain proteins.
PMID 18475320 · PMC2377100 · PLoS computational biology · 2008 · 8 claims · 6 setups
Traditional homology definitions do not capture multidomain evolution; the authors extend the definition to include domain insertion via a common ancestral locus model.
-
Full-text index only
Proteomic analysis of stage I primary lung adenocarcinoma aimed at individualisation of postoperative therapy.
PMID 18212748 · PMC2243141 · British journal of cancer · 2008 · 5 claims · 6 setups
LC-MS/MS proteomic analysis of stage I lung adenocarcinoma specimens identified myosin IIA and vimentin as candidate biomarker proteins with signal intensities that differed significantly among patient outcome groups
-
Full-text index only
The ENCODE Project at UC Santa Cruz.
PMID 17166863 · PMC1781110 · Nucleic acids research · 2007 · 8 claims · 4 setups
The UCSC ENCODE portal serves as the primary repository and access point for sequence-based ENCODE pilot phase data
-
Has reproduction · 98
Massively parallel genomic perturbations with multi-target CRISPR interrogates Cas9 activity and DNA repair at endogenous sites.
PMID 36064968 · PMC9481459 · Nature cell biology · 2022 · 8 claims · 6 setups
Multi-target gRNAs (mgRNAs) can direct Cas9 to over a hundred well-mapped endogenous genomic sites simultaneously, enabling massively parallel, high-throughput interrogation of Cas9 activity via short-read sequencing
-
Full-text index only
Recurring genomic breaks in independent lineages support genomic fragility.
PMID 17090315 · PMC1636669 · BMC evolutionary biology · 2006 · 6 claims · 6 setups
The propensity of a chromosomal region to break is significantly correlated among independent lineages, even after accounting for covariates like region length and functional class.
-
Full-text index only
Analysis of the prostate cancer cell line LNCaP transcriptome using a sequencing-by-synthesis approach.
PMID 17010196 · PMC1592491 · BMC genomics · 2006 · 8 claims · 7 setups
High-throughput 454 sequencing-by-synthesis of LNCaP cDNA can profile transcript abundance across the transcriptome
-
Full-text index only
Sequence determinants of human microsatellite variability.
PMID 20015383 · PMC2806349 · BMC genomics · 2009 · 6 claims · 4 setups
Mean and maximum number of repeats across individuals are positively correlated with heterozygosity
-
Full-text index only
A rigorous method for multigenic families' functional annotation: the peptidyl arginine deiminase (PADs) proteins family example.
PMID 16271148 · PMC1310624 · BMC genomics · 2005 · 8 claims · 5 setups
Integrating EST-based expression data with phylogenetic analysis is a valid new method for functionally annotating multigenic protein families
-
Full-text index only
COMUS: Clinician-Oriented locus-specific MUtation detection and deposition System.
PMID 19958500 · PMC2788389 · BMC genomics · 2009 · 8 claims · 6 setups
COMUS is a bioinformatics system for detecting and depositing new mutations from patient DNA with a clinician-friendly interface
-
Has reproduction · 75
Identification of Novel Therapeutic Candidates Against SARS-CoV-2 Infections: An Application of RNA Sequencing Toward mRNA Based Nanotherapeutics.
PMID 35983322 · PMC9378778 · Frontiers in microbiology · 2022 · 6 claims · 7 setups
RPL29 (60S ribosomal protein L29) is highly/consistently expressed across all COVID-19 infected groups regardless of severity, suggesting it as a novel host therapeutic target for mRNA-based nanomedicines.