Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 87
R2DT is a framework for predicting and visualising RNA secondary structure using templates.
PMID 34108470 · PMC8190129 · Nature communications · 2021 · 8 claims · 6 setups
R2DT is a template-based computational framework/pipeline that predicts and visualises RNA 2D structure in standardised, community-accepted layouts
-
Full-text index only
Computing Ka and Ks with a consideration of unequal transitional substitutions.
PMID 16740169 · PMC1552089 · BMC evolutionary biology · 2006 · 7 claims · 7 setups
MYN, a modified version of the Yang-Nielsen (YN) algorithm based on the Tamura-Nei Model, allows unequal transitional substitution rates between purines (κR) and pyrimidines (κY) plus codon frequency bias
-
Full-text index only
The diploid genome sequence of an Asian individual.
PMID 18987735 · PMC2716080 · Nature · 2008 · 8 claims · 8 setups
First diploid genome sequence of an Asian (Han Chinese) individual generated using massively parallel Illumina sequencing
-
Full-text index only
Empirical codon substitution matrix.
PMID 15927081 · PMC1173088 · BMC bioinformatics · 2005 · 8 claims · 5 setups
The authors present the first empirical codon substitution matrix built entirely from alignments of vertebrate coding DNA sequences.
-
Has reproduction · 24
MiGPC: a comprehensive catalog of enzybiotics from environmental metagenomes.
PMID 41888223 · PMC13172421 · Scientific reports · 2026 · 8 claims · 8 setups
MiGPC is the first genome-resolved metagenomic gene and protein catalog specifically targeted to enzybiotics
-
Full-text index only
GeneTide--Terra Incognita Discovery Endeavor: a new transcriptome focused member of the GeneCards/GeneNote suite of databases.
PMID 15608261 · PMC540076 · Nucleic acids research · 2005 · 8 claims · 7 setups
GeneTide integrates UniGene, DoTS, AceView, BLAT/GeneLoc genomic alignment, and GeneAnnot probe-set data into a unified Consensus/Uniqueness/Score scheme to associate ESTs with GeneCards genes
-
Full-text index only
Genome wide identification of recessive cancer genes by combinatorial mutation analysis.
PMID 18846217 · PMC2557123 · PloS one · 2008 · 7 claims · 4 setups
A combinatorial mutation analysis identified 154 candidate recessive cancer genes (pRecessiveCancer<1.5x10-7, FDR=0.39)
-
Full-text index only
Integrative functional genomics.
PMID 15239826 · PMC463286 · Genome biology · 2004 · 8 claims · 8 setups
Ultra-conserved noncoding elements exist across human, mouse and rat genomes at very high sequence identity, often far from genes
-
Full-text index only
PigGIS: Pig Genomic Informatics System.
PMID 17090590 · PMC1669765 · Nucleic acids research · 2007 · 7 claims · 7 setups
PigGIS identified 15,700 pig consensus sequences covering 18.5 Mb of homologous human exons
-
Full-text index only
The global landscape of sequence diversity.
PMID 17996061 · PMC2258180 · Genome biology · 2007 · 7 claims · 5 setups
Eukaryotic sequence datasets show substantially greater genetic diversity (higher sequence/gene family discovery rates) than bacterial datasets, likely related to differences in modes of genetic inheritance.
-
Has reproduction · 67
A consensus approach to vertebrate de novo transcriptome assembly from RNA-seq data: assembly of the duck (Anas platyrhynchos) transcriptome.
PMID 25009556 · PMC4070175 · Frontiers in genetics · 2014 · 8 claims · 8 setups
Multiple k-mer (MK) assemblies are more complete than single k-mer (SK) assemblies, showing higher reads-mapped-back-to-transcripts (RMBT) and higher CEGMA complete-gene percentages for all three tools.
-
Has reproduction · 83
Macrel: antimicrobial peptide screening in genomes and metagenomes.
PMID 33384902 · PMC7751412 · PeerJ · 2020 · 8 claims · 8 setups
Macrel is an end-to-end pipeline that predicts high-quality AMP candidates from peptides, contigs, or reads of (meta)genomes
-
Full-text index only
Vertebrate gene finding from multiple-species alignments using a two-level strategy.
PMID 16925840 · PMC1810555 · Genome biology · 2006 · 8 claims · 5 setups
DOGFISH cleanly separates a multi-species alignment classifier (RVM cascade) from an HMM-based structure predictor, avoiding tight coupling of alignment complexity with HMM formalism
-
Full-text index only
The Neandertal genome and ancient DNA authenticity.
PMID 19661919 · PMC2725275 · The EMBO journal · 2009 · 8 claims · 6 setups
Only direct assays of DNA sequence positions where Neandertals differ from all contemporary humans can reliably estimate human contamination.
-
Full-text index only
ChimerDB 2.0--a knowledgebase for fusion genes updated.
PMID 19906715 · PMC2808913 · Nucleic acids research · 2010 · 8 claims · 4 setups
ChimerDB 2.0 is an updated knowledgebase integrating fusion transcripts from GenBank transcriptome analysis with Sanger CGP, OMIM, PubMed, and Mitelman's database data.
-
Has reproduction · 67
Satellitome Analysis and Transposable Elements Comparison in Geographically Distant Populations of Spodoptera frugiperda.
PMID 35455012 · PMC9026859 · Life (Basel, Switzerland) · 2022 · 8 claims · 5 setups
Most transposable elements are commonly shared across all eight geographically distant S. frugiperda samples, except Maverick and PIF/Harbinger elements which show divergent repeat copies
-
Has reproduction · 87
Insights into the Evolution of the New World Diploid Cottons (Gossypium, Subgenus Houzingenia) Based on Genome Sequencing.
PMID 30476109 · PMC6320677 · Genome biology and evolution · 2019 · 8 claims · 8 setups
Subgenus Houzingenia originated via transoceanic dispersal from Africa ~6.6 Ma, with most biodiversity arising from rapid mid-Pleistocene (0.5–2.0 Ma) diversification plus multiple long-distance dispersals.
-
Has reproduction · 65
FusionQ: a novel approach for gene fusion detection and quantification from paired-end RNA-Seq.
PMID 23768108 · PMC3691734 · BMC bioinformatics · 2013 · 8 claims · 8 setups
FusionQ is a novel tool that detects gene fusions, constructs chimerical transcript structures, and estimates their abundances from paired-end RNA-Seq data.
-
Has reproduction · 50
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization.
PMID 41266599 · PMC12635123 · Communications biology · 2025 · 8 claims · 8 setups
BiRNA-BERT uses adaptive dual-tokenization that dynamically selects nucleotide-level (NUC) or byte-pair encoding (BPE) tokens based on input sequence length
-
Has reproduction · 69
A comparison across non-model animals suggests an optimal sequencing depth for de novo transcriptome assembly.
PMID 23496952 · PMC3655071 · BMC genomics · 2013 · 8 claims · 8 setups
Representative de novo transcriptome assemblies are generated with as few as ~20 million reads for single-tissue samples and ~30 million reads for whole animals at the mRNA-coverage level.