Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 79
Detection of Virus-Related Sequences Associated With Potential Etiologies of Hepatitis in Liver Tissue Samples From Rats, Mice, Shrews, and Bats.
PMID 34177835 · PMC8221242 · Frontiers in microbiology · 2021 · 8 claims · 6 setups
Viral metagenomics of liver tissue from rats, mice, shrews, and bats revealed a diverse set of sequences related to herpesviruses, orthomyxoviruses, anelloviruses, hepeviruses, hepadnaviruses, flaviviruses, parvoviruses, and picornaviruses
-
Has reproduction · 69
Discovery and characterization of Alu repeat sequences via precise local read assembly.
PMID 26503250 · PMC4666360 · Nucleic acids research · 2015 · 7 claims · 8 setups
Combining Alu-supporting read detection (RetroSeq) with local de novo assembly (CAP3) reconstructs the full sequence of non-reference Alu insertions from Illumina paired-end WGS reads
-
Full-text index only
A re-annotation pipeline for Illumina BeadArrays: improving the interpretation of gene expression data.
PMID 19923232 · PMC2817484 · Nucleic acids research · 2010 · 8 claims · 7 setups
A Perl-based pipeline that BLASTs/BLATs Illumina probe sequences against genomes and transcript databases (RefSeq, UCSC Known Genes, UniGene/GenBank, Ensembl) can classify probes by quality grade (Perfect/Good/Bad/No match) and is applicable across 8 BeadArray platforms and other array types
-
Has reproduction · 61
Comprehensive transcriptome study to develop molecular resources of the copepod Calanus sinicus for their potential ecological applications.
PMID 24982883 · PMC4055022 · BioMed research international · 2014 · 8 claims · 8 setups
Illumina RNA-Seq with Trinity de novo assembly produced a C. sinicus transcriptome of 69,751 contigs (average 928.8 bp, N50 1,127 bp) from 58.9 million reads.
-
Has reproduction · 87
A target enrichment method for gathering phylogenetic information from hundreds of loci: An example from the Compositae.
PMID 25202605 · PMC4103609 · Applications in plant sciences · 2014 · 8 claims · 8 setups
A custom sequence capture probe set (9678 baits targeting 1061 orthologous genes) was designed to enrich COS loci across the Compositae.
-
Has reproduction · 67
A consensus approach to vertebrate de novo transcriptome assembly from RNA-seq data: assembly of the duck (Anas platyrhynchos) transcriptome.
PMID 25009556 · PMC4070175 · Frontiers in genetics · 2014 · 8 claims · 8 setups
Multiple k-mer (MK) assemblies are more complete than single k-mer (SK) assemblies, showing higher reads-mapped-back-to-transcripts (RMBT) and higher CEGMA complete-gene percentages for all three tools.
-
Has reproduction · 23
Analysis of whole-genome re-sequencing data of ducks reveals a diverse demographic history and extensive gene flow between Southeast/South Asian and Chinese populations.
PMID 33849442 · PMC8042899 · Genetics, selection, evolution : GSE · 2021 · 8 claims · 8 setups
Whole-genome resequencing reveals three geographically distinct genetic groups: local Chinese, wild, and local Southeast/South Asian duck populations
-
Full-text index only
SeqBuster, a bioinformatic tool for the processing and analysis of small RNAs datasets, reveals ubiquitous miRNA modifications in human embryonic cells.
PMID 20008100 · PMC2836562 · Nucleic acids research · 2010 · 8 claims · 6 setups
SeqBuster is a versatile web-based and stand-alone bioinformatic toolkit for processing and analyzing large-scale small RNA deep sequencing datasets.
-
Has reproduction · 57
Analysis and comprehensive comparison of PacBio and nanopore-based RNA sequencing of the Arabidopsis transcriptome.
PMID 32536962 · PMC7291481 · Plant methods · 2020 · 8 claims · 8 setups
ONT Pc produces higher raw data quality (higher alignment rate, lower error rate) than ONT Dc, while PacBio generates the longest reads
-
Has reproduction · 50
GAL08, an Uncultivated Group of Acidobacteria, Is a Dominant Bacterial Clade in a Neutral Hot Spring.
PMID 35087491 · PMC8787282 · Frontiers in microbiology · 2021 · 8 claims · 8 setups
GAL08 is a dominant bacterial clade in a neutral hot spring, comprising up to 29.2% of the community by relative read abundance and up to 4.7 × 10^5 16S rRNA gene copies per gram sediment.
-
Has reproduction · 50
Genome-wide identification of Hfq-regulated small RNAs in the fire blight pathogen Erwinia amylovora discovered small RNAs with virulence regulatory function.
PMID 24885615 · PMC4070566 · BMC genomics · 2014 · 8 claims · 8 setups
A total of 40 candidate Hfq-dependent sRNAs were identified genome-wide in E. amylovora by combining RNA-seq with a Rho-independent terminator search.
-
Full-text index only
Enrichment of sequencing targets from the human genome by solution hybridization.
PMID 19835619 · PMC2784331 · Genome biology · 2009 · 8 claims · 5 setups
Solution hybridization with 120-mer capture probes efficiently enriches targeted genomic sequences for next-generation sequencing
-
Has reproduction · 76
What the Phage: a scalable workflow for the identification and analysis of phage sequences.
PMID 36399058 · PMC9673492 · GigaScience · 2022 · 8 claims · 7 setups
WtP combines 11 tools (14 approaches) for phage prediction in a parallel, containerized Nextflow workflow
-
Has reproduction · 80
SLDMS: A Tool for Calculating the Overlapping Regions of Sequences.
PMID 35046988 · PMC8761809 · Frontiers in plant science · 2021 · 8 claims · 5 setups
SLDMS is a novel method for computing overlapping regions of sequencing reads using suffix array (SA), longest common prefix (LCP) array, document array (DA), and a monotonic stack.
-
Has reproduction · 89
A near complete genome for goat genetic and genomic research.
PMID 34507524 · PMC8434745 · Genetics, selection, evolution : GSE · 2021 · 8 claims · 8 setups
Saanen_v1 is a near-complete de novo goat genome assembly generated from 117x PacBio and 118x Hi-C data, including the first goat Y chromosome scaffold
-
Has reproduction · 54
Population structure analysis of Salmonella serovar Muenchen to redefine geno-serotyping using genome indexing approaches.
PMID 41743541 · PMC12929376 · Frontiers in microbiology · 2025 · 6 claims · 6 setups
Integrating genome-indexing (bettercallsal, DNA sketching + genome proximity) with SeqSero2 yields complementary serovar calls that improve discrimination of genomically distinct but antigenically similar serovars while retaining historical nomenclature
-
Has reproduction · 87
De Novo Transcriptome Meta-Assembly of the Mixotrophic Freshwater Microalga Euglena gracilis.
PMID 34072576 · PMC8227486 · Genes · 2021 · 6 claims · 8 setups
A consensus transcriptome assembled by combining reads from five independent studies is the most complete E. gracilis transcriptome released to date, outperforming the two previously available transcriptomes (GEFR01 and GDJR01).
-
Has reproduction · 50
MEDUSA: A Pipeline for Sensitive Taxonomic Classification and Flexible Functional Annotation of Metagenomic Shotgun Sequences.
PMID 35330728 · PMC8940201 · Frontiers in genetics · 2022 · 6 claims · 6 setups
MEDUSA is an automated, Conda-installable and Snakemake-managed pipeline performing preprocessing, assembly, alignment, taxonomic classification, and functional annotation on shotgun data.
-
Has reproduction · 20
Expression and Secretion of Circular RNAs in the Parasitic Nematode, Ascaris suum.
PMID 35711944 · PMC9194832 · Frontiers in genetics · 2022 · 8 claims · 7 setups
A. suum expresses 1,997 distinct circRNAs identified via next-generation sequencing in adult female body wall and ovary-enriched tissue
-
Has reproduction · 67
Optimal scaling of digital transcriptomes.
PMID 24223126 · PMC3819321 · PloS one · 2013 · 8 claims · 8 setups
Fifteen existing and novel transcript-count normalization algorithms can be compared with two novel, mutually independent metrics: the number of "uniform" genes (sufficiently low coefficient of variation after normalization) and low average Spearman correlation between normalized expression profiles of gene pairs.