Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 87
Forseti: a mechanistic and predictive model of the splicing status of scRNA-seq reads.
PMID 38940130 · PMC11256924 · Bioinformatics (Oxford, England) · 2024 · 7 claims · 5 setups
Forseti is the first probabilistic model for resolving the splicing status of exonic scRNA-seq reads by scoring putative fragments linking read alignments to proximate priming sites
-
Has reproduction · 75
stDyer-image improves clustering analysis of spatially resolved transcriptomics and proteomics with morphological images.
PMID 41692960 · PMC12960910 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 5 setups
stDyer-image directly associates the image modality with predicted cluster labels rather than using images to enhance/impute gene expression data
-
Has reproduction · 82
Landscape of allele-specific transcription factor binding in the human genome.
PMID 33980847 · PMC8115691 · Nature communications · 2021 · 8 claims · 6 setups
A novel statistical framework (ADASTRA) calls allele-specific TF binding from existing ChIP-Seq alignments by jointly correcting for background allelic dosage (BAD, from aneuploidy/CNVs) and reference mapping bias.
-
Has reproduction · 77
Accurate chromatin marks peak calling with Omnipeak.
PMID 41521664 · PMC12784980 · Nucleic acids research · 2026 · 8 claims · 6 setups
Omnipeak is a universal unsupervised peak-calling algorithm based on a constrained three-state hidden Markov model (zero, noise, signal states)
-
Has reproduction · 97
Determination of complete chromosomal haplotypes by bulk DNA sequencing.
PMID 33957932 · PMC8101039 · Genome biology · 2021 · 8 claims · 8 setups
A hierarchical computational strategy that first builds high-confidence local haplotype blocks from long-range/linked-read linkage and then concatenates them into whole-chromosome haplotypes using Hi-C contacts
-
Has reproduction · 50
MoDLE: high-performance stochastic modeling of DNA loop extrusion interactions.
PMID 36451166 · PMC9710047 · Genome biology · 2022 · 8 claims · 6 setups
MoDLE is a high-performance stochastic model/software for simulating DNA-DNA contacts generated by loop extrusion genome-wide
-
Has reproduction
Fast, accurate, and racially unbiased pan-cancer tumor-only variant calling with tabular machine learning.
PMID 36611079 · PMC9825621 · NPJ precision oncology · 2023 · 8 claims · 8 setups
Tree-based (XGBoost, LightGBM) and deep-learning (TabNet) tabular ML classifiers achieve state-of-the-art somatic vs germline classification in tumor-only WES samples, outperforming PureCN.
-
Has reproduction · 73
Genetic polyploid phasing from low-depth progeny samples.
PMID 35692633 · PMC9184567 · iScience · 2022 · 8 claims · 7 setups
WH-PPG phases polyploid parental samples by scoring informative variant pairs with a Bayesian log-likelihood model of progeny allele depths, clustering alleles by co-occurrence likelihood, and assigning clusters to haplotypes via interval scheduling
-
Full-text index only
Exome sequencing of a multigenerational human pedigree.
PMID 20011588 · PMC2788131 · PloS one · 2009 · 8 claims · 6 setups
Microarray-based exome capture combined with 454 GS FLX NGS is an efficient and reliable method to enrich for chromosomal regions of interest, validated on eight individuals from a three-generation pedigree
-
Full-text index only
Profiling critical cancer gene mutations in clinical tumor samples.
PMID 19924296 · PMC2774511 · PloS one · 2009 · 7 claims · 4 setups
OncoMap, a panel of ~400 mass-spectrometric genotyping assays targeting 33 cancer genes, enables robust mutation profiling of clinical fresh-frozen and FFPE tumor DNA.
-
Full-text index only
Glioblastoma stem cells show transcriptionally correlated spatial organization.
PMID 41577992 · PMC12894897 · Communications biology · 2026 · 8 claims · 5 setups
GSCs from different patient samples exhibit diverse, characteristic multicellular spatial patterns in culture (e.g., anisotropic vs isotropic packing, overlapping vs non-overlapping space utilization)
-
Full-text index only
scLong: a billion-parameter foundation model for capturing long-range gene context in single-cell transcriptomics.
PMID 41639087 · PMC12982784 · Nature communications · 2026 · 7 claims · 4 setups
scLong performs self-attention across all ~27,874 human genes, including lowly expressed ones, to capture long-range gene dependencies missed by models restricted to highly expressed gene subsets
-
Full-text index only
Effective quantitative real-time polymerase chain reaction analysis of the parkin gene (PARK2) exon 1-12 dosage.
PMID 17324265 · PMC1810516 · BMC medical genetics · 2007 · 8 claims · 3 setups
Developed a real-time TaqMan PCR method that quantifies PARK2 exon 1-12 copy number by comparing amplification signal to the β-globin internal control gene
-
Has reproduction · 67
HArmonized single-cell RNA-seq Cell type Assisted Deconvolution (HASCAD).
PMID 37907883 · PMC10619225 · BMC medical genomics · 2023 · 6 claims · 4 setups
Removal of batch effects in reference scRNA-seq datasets (via Harmony-Symphony) benefits the task of cell composition deconvolution
-
Full-text index only
Addressing pandemic-wide systematic errors in the SARS-CoV-2 phylogeny.
PMID 41663577 · PMC12982125 · Nature methods · 2026 · 6 claims · 3 setups
Most SARS-CoV-2 genomes were sequenced using tiled amplicon methods, which introduce systematic errors unless assembly software is aware of the amplicon scheme and the error modes of amplicon sequencing
-
Has reproduction · 50
RNA-Seq alignment to individualized genomes improves transcript abundance estimates in multiparent populations.
PMID 25236449 · PMC4174954 · Genetics · 2014 · 8 claims · 7 setups
Genetic variants distinguishing an individual genome from the reference cause read misalignment and biased transcript abundance estimates, and fine-tuning of alignment algorithms does not correct this problem.
-
Full-text index only
Cancer genome standards for long-read sequencing using cancer cell line mixtures.
PMID 41934171 · PMC13137868 · GigaScience · 2026 · 8 claims · 6 setups
Long-read variant calling tools achieve recall rates comparable to short-read gold standards
-
Full-text index only
FracFixR: a compositional statistical framework for absolute proportion estimation between fractions in RNA sequencing data.
PMID 41264734 · PMC12866640 · Bioinformatics (Oxford, England) · 2026 · 7 claims · 5 setups
FracFixR reconstructs original fraction proportions by modeling the compositional relationship between whole and fractionated RNA samples using non-negative least squares (NNLS) regression on selected transcripts
-
Has reproduction · 50
DeeReCT-APA: Prediction of Alternative Polyadenylation Site Usage Through Deep Learning.
PMID 33662629 · PMC9801043 · Genomics, proteomics & bioinformatics · 2022 · 8 claims · 8 setups
DeeReCT-APA quantitatively predicts the usage of all competing PASs of a gene simultaneously, rather than casting the problem as pairwise comparison like prior methods.