Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 89
TrEMOLO: accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches.
PMID 37013657 · PMC10069131 · Genome biology · 2023 · 6 claims · 6 setups
TrEMOLO combines an assembly-based INSIDER module and a mapping-based OUTSIDER module to detect TE insertions/deletions from long-read sequencing data and estimate their allele frequency
-
Has reproduction · 77
spotter: a single-nucleotide resolution stochastic simulation model of supercoiling-mediated transcription and translation in prokaryotes.
PMID 37602419 · PMC10516669 · Nucleic acids research · 2023 · 8 claims · 4 setups
spotter is the first simulation model to integrate transcription, DNA supercoiling, and translation simultaneously in a single stochastic framework for prokaryotes.
-
Full-text index only
Design and analysis issues in genome-wide somatic mutation studies of cancer.
PMID 18692126 · PMC2820387 · Genomics · 2009 · 6 claims · 4 setups
Two-stage (discovery + validation) sequencing designs efficiently allocate resources and can produce highly informative candidate driver gene lists even with relatively small sample sizes.
-
Has reproduction · 60
UNMF: a unified nonnegative matrix factorization for multi-dimensional omics data.
PMID 37478378 · PMC10516365 · Briefings in bioinformatics · 2023 · 5 claims · 3 setups
UNMF is designed for tidy data format and structure, allowing it to handle a wide range of data structures and formats in a unified manner without requiring format-specific preprocessing.
-
Has reproduction · 50
RNA modifications detection by comparative Nanopore direct RNA sequencing.
PMID 34893601 · PMC8664944 · Nature communications · 2021 · 7 claims · 5 setups
Nanocompore is a model-free comparative method that uses a 2-component Gaussian mixture model (GMM) and univariate statistical tests on signal intensity/dwell time to detect RNA modifications in Nanopore direct RNA sequencing data without needing a training set
-
Full-text index only
Simultaneous analysis of all SNPs in genome-wide and re-sequencing association studies.
PMID 18654633 · PMC2464715 · PLoS genetics · 2008 · 8 claims · 5 setups
A Bayesian-inspired penalised maximum likelihood stochastic search method can simultaneously analyse all SNPs (up to 500K) from a GWA study in a few hours on a desktop workstation
-
Full-text index only
Population history and natural selection shape patterns of genetic variation in 132 genes.
PMID 15361935 · PMC515367 · PLoS biology · 2004 · 7 claims · 5 setups
Developed a rigorous computational approach that corrects for multiple hypothesis testing and models population demographic history to test for natural selection
-
Full-text index only
Inter-individual variation of DNA methylation and its implications for large-scale epigenome mapping.
PMID 18413340 · PMC2425484 · Nucleic acids research · 2008 · 8 claims · 8 setups
CpG-rich regions (CpG islands) show low and similar methylation levels across individuals, but the sequential order of the few methylated CpGs among the many unmethylated ones varies randomly between individuals.
-
Full-text index only
The stem cell population of the human colon crypt: analysis via methylation patterns.
PMID 17335343 · PMC1808490 · PLoS computational biology · 2007 · 8 claims · 3 setups
A coalescent-based, full probabilistic model with MCMC Bayesian inference provides a more powerful alternative to prior forward-simulation approaches for analyzing methylation pattern data from crypts.
-
Full-text index only
Decoding of superimposed traces produced by direct sequencing of heterozygous indels.
PMID 18654614 · PMC2429969 · PLoS computational biology · 2008 · 7 claims · 3 setups
A dynamic programming method (implemented as web app Indelligent) can decode superimposed allelic sequences from a single mixed trace, using only the observed string of ambiguous peak calls, without a reference sequence or reverse trace.
-
Has reproduction · 90
LoRA-TV: read depth profile-based clustering of tumor cells in single-cell sequencing.
PMID 38877886 · PMC11179121 · Briefings in bioinformatics · 2024 · 6 claims · 2 setups
LoRA-TV jointly processes read-depth profiles of all cells by stacking them into a matrix and applying low-rank approximation plus total-variation smoothing to capture shared genomic signatures for clustering.
-
Has reproduction · 53
Estimates of recent and historical effective population size in turbot, seabream, seabass and carp selective breeding programmes.
PMID 34742227 · PMC8572424 · Genetics, selection, evolution : GSE · 2021 · 7 claims · 7 setups
Current effective population size for all four farmed fish populations is small (≤50 fish), potentially threatening breeding-programme sustainability
-
Has reproduction · 67
Generative and integrative modeling for transcriptomics with formalin fixed paraffin embedded material.
PMID 41029822 · PMC12486589 · Journal of translational medicine · 2025 · 8 claims · 5 setups
fRNA-seq transcript counts are best fit by the negative binomial distribution, with little evidence supporting zero-inflated extensions
-
Has reproduction · 66
RiboTaxa: combined approaches for rRNA genes taxonomic resolution down to the species level from metagenomics data revealing novelties.
PMID 36159175 · PMC9492272 · NAR genomics and bioinformatics · 2022 · 8 claims · 6 setups
RiboTaxa, combining BBTools, FastQC, SortMeRNA, MetaRib, EMIRGE, VSEARCH, BBMap and QIIME 2's Sklearn classifier, was built as a pipeline for SSU rRNA-based taxonomic profiling of metagenomics data.
-
Has reproduction · 44
Detecting DNA modifications from SMRT sequencing data by modeling sequence context dependence of polymerase kinetic.
PMID 23516341 · PMC3597545 · PLoS computational biology · 2013 · 8 claims · 7 setups
Local sequence context strongly determines position-specific polymerase kinetic rate: roughly 80% of IPD variation is explained by a 10 bp context (7 bases upstream, 2 bases downstream of the incorporation site), saturating at 7 bases upstream.
-
Full-text index only
Glucokinase gene mutations: structural and genotype-phenotype analyses in MODY children from South Italy.
PMID 18382660 · PMC2270336 · PloS one · 2008 · 8 claims · 6 setups
16 of 30 patients with suspected MODY (53%) carry GCK mutations, confirming GCK MODY diagnosis
-
Full-text index only
No evidence of a Neanderthal contribution to modern human diversity.
PMID 18304371 · PMC2374707 · Genome biology · 2008 · 8 claims · 7 setups
There is no evidence of any Neanderthal contribution to modern human genetic diversity
-
Has reproduction · 96
A bioinformatic pipeline for simulating viral integration data.
PMID 35496474 · PMC9046613 · Data in brief · 2022 · 7 claims · 3 setups
A snakemake-based pipeline was developed to simulate integration of a viral or vector genome into a host genome, including sub-genomic fragment integration, structural variation, and host-site deletions.
-
Has reproduction
Genomic prediction based on selective linkage disequilibrium pruning of low-coverage whole-genome sequence variants in a pure Duroc population.
PMID 37853325 · PMC10583454 · Genetics, selection, evolution : GSE · 2023 · 8 claims · 6 setups
Selective linkage disequilibrium pruning (SLDP) refines whole-genome SNP sets using GWAS prior information to improve genomic prediction accuracy.
-
Has reproduction · 64
Nimbus: a design-driven analyses suite for amplicon-based NGS data.
PMID 29538618 · PMC6084620 · Bioinformatics (Oxford, England) · 2018 · 7 claims · 4 setups
Nimbus is an end-to-end software suite for amplicon-based NGS data that tracks source amplicons through alignment and variant calling, with tools for trimming, alignment, SNP/InDel calling, QC and visualization.