Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Using ESTs to improve the accuracy of de novo gene prediction.
PMID 16817966 · PMC1534067 · BMC bioinformatics · 2006 · 8 claims · 8 setups
TWINSCAN_EST combines EST alignments with TWINSCAN via a trainable 'ESTseq' representation and improves exact gene structure prediction accuracy on the whole C. elegans genome
-
Full-text index only
CONTRAST: a discriminative, phylogeny-free approach to multiple informant de novo gene prediction.
PMID 18096039 · PMC2246271 · Genome biology · 2007 · 8 claims · 5 setups
CONTRAST predicts exact coding region structures for 65% more human genes than the previous state-of-the-art de novo predictor (N-SCAN)
-
Full-text index only
Evolution of genomic sequence inhomogeneity at mid-range scales.
PMID 19891785 · PMC2779198 · BMC genomics · 2009 · 7 claims · 3 setups
MRI regions have comparable levels of de novo mutations to control genomic sequences with average base composition.
-
Has reproduction · 86
LMAS: evaluating metagenomic short de novo assembly methods through defined communities.
PMID 36576131 · PMC9795473 · GigaScience · 2022 · 8 claims · 5 setups
LMAS (Last Metagenomic Assembler Standing) is a flexible, Nextflow-based, Docker-containerized automated workflow for benchmarking de novo metagenomic assemblers against defined mock communities, producing an interactive HTML report.
-
Has reproduction · 75
Graph-Based Approaches Significantly Improve the Recovery of Antibiotic Resistance Genes From Complex Metagenomic Datasets.
PMID 34690959 · PMC8528159 · Frontiers in microbiology · 2021 · 8 claims · 6 setups
GraphAMR, a Nextflow pipeline that aligns AMR profile HMMs (or AA sequences) to metagenomic assembly graphs via PathRacer, then dereplicates and annotates hits, recovers more and more complete AMR genes than contig-based or read-based methods.
-
Full-text index only
Using several pair-wise informant sequences for de novo prediction of alternatively spliced transcripts.
PMID 16925842 · PMC1810557 · Genome biology · 2006 · 8 claims · 4 setups
MARS, an extension of the Twinscan algorithm, uses multiple pairwise informant genomes to predict human alternatively spliced transcripts de novo without expressed sequence information.
-
Full-text index only
Combining comparative genomics with de novo motif discovery to identify human transcription factor DNA-binding motifs.
PMID 17217514 · PMC1780116 · BMC bioinformatics · 2006 · 6 claims · 4 setups
A novel method combining 8-species comparative genomics with de novo motif discovery identifies human TF DNA-binding motifs overrepresented and conserved in upstream regions of co-regulated genes
-
Full-text index only
Pairagon+N-SCAN_EST: a model-based gene annotation pipeline.
PMID 16925839 · PMC1810554 · Genome biology · 2006 · 7 claims · 5 setups
Pairagon+N-SCAN_EST, using only native alignments, was as accurate as ENSEMBL and ExoGean in the EGASP mRNA/EST evidence assessment
-
Full-text index only
CompMoby: comparative MobyDick for detection of cis-regulatory motifs.
PMID 18950538 · PMC2605473 · BMC bioinformatics · 2008 · 7 claims · 4 setups
CompMoby identifies cis-regulatory binding sites at both transcriptional and post-transcriptional levels in metazoans without prior knowledge of the trans-acting factor
-
Has reproduction · 83
Accurate prediction of metagenome-assembled genome completeness by MAGISTA, a random forest model built on alignment-free intra-bin statistics.
PMID 35248155 · PMC8898458 · Environmental microbiome · 2022 · 7 claims · 7 setups
MAGISTA, a random forest model built on alignment-free intra-bin distance-distribution statistics, can estimate MAG completeness and purity without relying on reference marker genes.
-
Has reproduction · 95
Identification and Characterization of Small Noncoding RNAs in Genome Sequences of the Edible Fungus Pleurotus ostreatus.
PMID 27703969 · PMC5040776 · BioMed research international · 2016 · 8 claims · 7 setups
254 small noncoding RNAs (snRNAs, snoRNAs, tRNAs, miRNAs) were detected in the P. ostreatus CCEF00389 genome assembly, the first genome-scale identification of sncRNAs for a basidiomycete.
-
Has reproduction · 91
De Novo Assembly and Annotation of the Larval Transcriptome of Two Spadefoot Toads Widely Divergent in Developmental Rate.
PMID 31217263 · PMC6686947 · G3 (Bethesda, Md.) · 2019 · 8 claims · 8 setups
De novo transcriptome assemblies were generated for larval P. cultripes and S. couchii, providing new genomic resources for spadefoot toads
-
Has reproduction · 91
Whole genome and transcriptome maps of the entirely black native Korean chicken breed Yeonsan Ogye.
PMID 30010758 · PMC6065499 · GigaScience · 2018 · 6 claims · 7 setups
A hybrid de novo assembly combining high-depth Illumina short reads (376.6X) and low-depth PacBio long reads (9.7X) produced the YO draft genome Ogye_1.1 with contig and scaffold NG50 of 362.3 Kbp and 16.8 Mbp.
-
Has reproduction · 94
A Deluge of Complex Repeats: The Solanum Genome.
PMID 26241045 · PMC4524691 · PloS one · 2015 · 8 claims · 7 setups
~50–60% of the S. tuberosum and S. lycopersicum genomes are composed of repetitive elements
-
Has reproduction · 83
MetaGT: A pipeline for de novo assembly of metatranscriptomes with the aid of metagenomic data.
PMID 36386613 · PMC9651917 · Frontiers in microbiology · 2022 · 7 claims · 4 setups
MetaGT is a pipeline that combines metatranscriptomic and metagenomic data from the same sample to assemble complete transcript sequences
-
Has reproduction · 85
An extensive evaluation of read trimming effects on Illumina NGS data analysis.
PMID 24376861 · PMC3871669 · PloS one · 2013 · 8 claims · 8 setups
Read trimming increases the quality and reliability of downstream NGS analyses (RNA-Seq mapping, SNP identification, genome assembly) while reducing execution time and computational resources.
-
Full-text index only
The sequence and de novo assembly of the giant panda genome.
PMID 20010809 · PMC3951497 · Nature · 2010 · 8 claims · 8 setups
A draft giant panda genome was successfully generated and assembled de novo using only Illumina Genome Analyser short-read sequencing
-
Has reproduction · 79
pyrpipe: a Python package for RNA-Seq workflows.
PMID 34085037 · PMC8168212 · NAR genomics and bioinformatics · 2021 · 8 claims · 3 setups
pyrpipe enables development of flexible, reproducible, and easy-to-debug RNA-Seq computational pipelines purely in Python, in an object-oriented manner
-
Has reproduction · 45
De novo transcriptomic analysis of leaf and fruit tissue of Cornus officinalis using Illumina platform.
PMID 29451882 · PMC5815590 · PloS one · 2018 · 7 claims · 7 setups
This is the first de novo transcriptomic analysis of Cornus officinalis, providing fundamental gene and biosynthetic pathway information.
-
Has reproduction · 69
Manual curation for improved genome annotation of the functionally extinct northern white rhinoceros (Ceratotherium simum cottoni).
PMID 41490125 · PMC12768360 · PloS one · 2026 · 7 claims · 7 setups
Manual curation of RNA-seq-derived de novo transcripts increased the number of functional genes in the NWR annotation by 81% (from 8,701 to 15,738).