Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 86
The selection of software and database for metagenomics sequence analysis impacts the outcome of microbial profiling and pathogen detection.
PMID 37027361 · PMC10081788 · PloS one · 2023 · 7 claims · 7 setups
Obtaining an accurate species-level microbial profile using current direct-read metagenomics profiling software is still a challenging task.
-
Full-text index only
Gene loss rate: a probabilistic measure for the conservation of eukaryotic genes.
PMID 17158152 · PMC1802574 · Nucleic acids research · 2007 · 8 claims · 8 setups
GLR is a novel maximum-likelihood measure of gene loss rate that probabilistically weighs all possible ancestral phyletic patterns rather than relying on a single parsimonious reconstruction.
-
Full-text index only
Meta-analysis of inter-species liver co-expression networks elucidates traits associated with common human diseases.
PMID 20019805 · PMC2787626 · PLoS computational biology · 2009 · 8 claims · 8 setups
A novel semi-parametric meta-analysis method (based on a gene-centric Glass's d effect size) outperforms existing parametric and non-parametric meta-analysis methods at identifying functionally coherent gene pairs across species.
-
Has reproduction · 99
Evaluation of taxonomic classification and profiling methods for long-read shotgun metagenomic sequencing datasets.
PMID 36513983 · PMC9749362 · BMC bioinformatics · 2022 · 8 claims · 7 setups
Long-read classifiers generally performed best among the 11 methods tested
-
Has reproduction · 85
PowerBacGWAS: a computational pipeline to perform power calculations for bacterial genome-wide association studies.
PMID 35338232 · PMC8956664 · Communications biology · 2022 · 8 claims · 8 setups
Two computational approaches (sub-sampling and phenotype-simulation) can be implemented to perform power calculations for bacterial GWAS using existing genome collections, packaged as the PowerBacGWAS pipeline
-
Full-text index only
Assessing the gene space in draft genomes.
PMID 19042974 · PMC2615622 · Nucleic acids research · 2009 · 6 claims · 7 setups
The proportion of mapped CEGs in a draft genome assembly is a useful metric for describing gene space completeness, complementing N50 and x-fold coverage.
-
Has reproduction · 59
Comparing time series transcriptome data between plants using a network module finding algorithm.
PMID 31164912 · PMC6544932 · Plant methods · 2019 · 8 claims · 6 setups
Converting time-series expression data into co-expression networks and applying network module finding (OrthoClust) enables cross-species comparison without requiring one-to-one developmental stage mapping.
-
Has reproduction · 95
Mouse-Geneformer: A deep learning model for mouse single-cell transcriptome and its cross-species utility.
PMID 40106407 · PMC11964219 · PLoS genetics · 2025 · 7 claims · 6 setups
Mouse-Geneformer, a Transformer Encoder model pre-trained via masked-token self-supervised learning on mouse-Genecorpus-20M, was successfully constructed following the original human Geneformer architecture.
-
Has reproduction · 66
RiboTaxa: combined approaches for rRNA genes taxonomic resolution down to the species level from metagenomics data revealing novelties.
PMID 36159175 · PMC9492272 · NAR genomics and bioinformatics · 2022 · 8 claims · 6 setups
RiboTaxa, combining BBTools, FastQC, SortMeRNA, MetaRib, EMIRGE, VSEARCH, BBMap and QIIME 2's Sklearn classifier, was built as a pipeline for SSU rRNA-based taxonomic profiling of metagenomics data.
-
Full-text index only
Structural organization and interactions of transmembrane domains in tetraspanin proteins.
PMID 15985154 · PMC1190194 · BMC structural biology · 2005 · 8 claims · 5 setups
TM1, TM2 and TM3 of human tetraspanins display a distinct heptad repeat motif (abcdefg)n, while TM4 lacks this motif.
-
Has reproduction · 89
Improved eukaryotic detection compatible with large-scale automated analysis of metagenomes.
PMID 37032329 · PMC10084625 · Microbiome · 2023 · 8 claims · 7 setups
MAPQ ≥30 filtering improves precision but substantially reduces recall, especially for unrepresented/divergent eukaryotic taxa
-
Full-text index only
Effect of the assignment of ancestral CpG state on the estimation of nucleotide substitution rates in mammals.
PMID 18826599 · PMC2576242 · BMC evolutionary biology · 2008 · 7 claims · 4 setups
CpG/non-CpG assignment based on presence/absence of a CpG dinucleotide seriously biases substitution rate estimates, overestimating CpG changes and underestimating non-CpG changes.
-
Full-text index only
A space-efficient and accurate method for mapping and aligning cDNA sequences onto genomic sequence.
PMID 18344523 · PMC2377433 · Nucleic acids research · 2008 · 7 claims · 6 setups
Spaln maps and aligns large cDNA sequence sets onto whole mammalian genomes using substantially less memory than comparable existing tools
-
Full-text index only
Patrocles: a database of polymorphic miRNA-mediated gene regulation in vertebrates.
PMID 19906729 · PMC2808989 · Nucleic acids research · 2010 · 8 claims · 6 setups
Patrocles is a database compiling DSPs predicted to perturb miRNA-mediated gene regulation across seven vertebrate species, covering targets, miRNA precursors and silencing machinery.
-
Has reproduction · 42
CanCellCap: robust cancer cell capture across tissue types on single-cell RNA-seq data by multi-domain learning.
PMID 40739511 · PMC12312500 · BMC biology · 2025 · 8 claims · 8 setups
CanCellCap, a multi-domain learning framework integrating domain adversarial learning and Mixture of Experts, identifies cancer cells across all tissues, cancers, and sequencing platforms by extracting tissue-common and tissue-specific gene expression patterns.
-
Has reproduction · 88
Wochenende - modular and flexible alignment-based shotgun metagenome analysis.
PMID 36368923 · PMC9650795 · BMC genomics · 2022 · 8 claims · 6 setups
Wochenende is a modular, transparent alignment-based pipeline for shotgun metagenome analysis supporting short and long reads across all kingdoms of life
-
Full-text index only
A third approach to gene prediction suggests thousands of additional human transcribed regions.
PMID 16543943 · PMC1391917 · PLoS computational biology · 2006 · 8 claims · 7 setups
A third basic concept for gene prediction exists, based on detecting strand-specific 'transcription footprints' (mutational and selectional biases) rather than gene structure or sequence similarity.
-
Full-text index only
Genome-wide identification of human functional DNA using a neutral indel model.
PMID 16410828 · PMC1326222 · PLoS computational biology · 2006 · 8 claims · 8 setups
A neutral indel model predicting a geometric distribution of intergap segment (IGS) lengths fits human-mouse ancestral repeat (AR) alignment data excellently
-
Has reproduction · 53
Estimates of recent and historical effective population size in turbot, seabream, seabass and carp selective breeding programmes.
PMID 34742227 · PMC8572424 · Genetics, selection, evolution : GSE · 2021 · 7 claims · 7 setups
Current effective population size for all four farmed fish populations is small (≤50 fish), potentially threatening breeding-programme sustainability
-
Has reproduction · 86
LMAS: evaluating metagenomic short de novo assembly methods through defined communities.
PMID 36576131 · PMC9795473 · GigaScience · 2022 · 8 claims · 5 setups
LMAS (Last Metagenomic Assembler Standing) is a flexible, Nextflow-based, Docker-containerized automated workflow for benchmarking de novo metagenomic assemblers against defined mock communities, producing an interactive HTML report.