Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 50
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization.
PMID 41266599 · PMC12635123 · Communications biology · 2025 · 8 claims · 8 setups
BiRNA-BERT uses adaptive dual-tokenization that dynamically selects nucleotide-level (NUC) or byte-pair encoding (BPE) tokens based on input sequence length
-
Full-text index only
Advancing codon language modeling with synonymous codon constrained masking.
PMID 41736545 · PMC12956333 · Nucleic acids research · 2026 · 8 claims · 7 setups
SynCodonLM introduces synonymous codon-constrained masking, restricting masked-codon prediction to only synonymous codon options via logit masking before softmax
-
Has reproduction · 57
Diapause vs. reproductive programs: transcriptional phenotypes in a keystone copepod.
PMID 33782539 · PMC8007741 · Communications biology · 2021 · 8 claims · 7 setups
t-SNE clustering of all-gene expression data groups field-collected (diapause program) samples into one cluster while early and late culture (reproductive program) samples separate into two distinct phenotypes
-
Has reproduction · 76
Tracing human genetic histories and natural selection with precise local ancestry inference.
PMID 40379651 · PMC12084304 · Nature communications · 2025 · 7 claims · 7 setups
Orchestra, a two-stage LAI method combining a recombination-distance base layer with a deep learning (convolutional + attention) smoothing module, outperforms RFmix, FLARE and Gnomix in precision and recall across simulated admixture generations.
-
Has reproduction · 85
ScLRTC: imputation for single-cell RNA-seq data via low-rank tensor completion.
PMID 34844559 · PMC8628418 · BMC genomics · 2021 · 8 claims · 8 setups
scLRTC imputes dropout entries closest to the original expression values on simulated datasets, outperforming other state-of-the-art methods by SSE and PCC.
-
Has reproduction · 94
Genome-wide insights into population structure and host specificity of Campylobacter jejuni.
PMID 33990625 · PMC8121833 · Scientific reports · 2021 · 8 claims · 6 setups
Both core and accessory genome characteristics show strong association with distinct host animal species, indicating multiple independent adaptive trajectories rather than a single common evolutionary path
-
Full-text index only
Retentive Network promotes efficient RNA language modeling of long sequences.
PMID 41814064 · PMC13111708 · Communications biology · 2026 · 8 claims · 6 setups
RNAret, a RetNet-based RNA language model with O(n) complexity, achieves training parallelism and low computational overhead while processing long RNA sequences
-
Full-text index only
scArchon: a scalable benchmarking framework for assessing single-cell perturbation models.
PMID 42121287 · PMC13162514 · Genome biology · 2026 · 8 claims · 8 setups
scArchon is a reproducible, modular, Snakemake-based benchmarking platform that evaluates perturbation response prediction tools in a standardized, containerized, extensible manner.
-
Has reproduction · 73
GREIN: An Interactive Web Platform for Re-analyzing GEO RNA-seq Data.
PMID 31110304 · PMC6527554 · Scientific reports · 2019 · 8 claims · 7 setups
GREIN is a web application providing user-friendly interfaces to manipulate, visualize, and analyze GEO RNA-seq data.
-
Has reproduction · 50
Grad-seq identifies KhpB as a global RNA-binding protein in Clostridioides difficile that regulates toxin production.
PMID 37223250 · PMC10117727 · microLife · 2021 · 8 claims · 9 setups
Grad-seq resolves in-gradient sedimentation profiles for ~87-88% of annotated C. difficile transcripts and ~50% of annotated proteins, providing a comprehensive RNA-protein complexome resource