Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
DDBJ dealing with mass data produced by the second generation sequencer.
PMID 18927114 · PMC2686496 · Nucleic acids research · 2009 · 8 claims · 7 setups
DDBJ collected and released 2,368,110 entries (1,415,106,598 bases) of original DNA sequence data from July 2007 to June 2008.
-
Full-text index only
DiagHunter and GenoPix2D: programs for genomic comparisons, large-scale homology discovery and visualization.
PMID 14519203 · PMC328457 · Genome biology · 2003 · 7 claims · 5 setups
DiagHunter identifies large-scale synteny blocks within or between genomes efficiently despite background noise and genomic discontinuities, without performing sequence alignment
-
Full-text index only
Integrating alternative splicing detection into gene prediction.
PMID 15705189 · PMC550657 · BMC bioinformatics · 2005 · 8 claims · 4 setups
An integrative intrinsic/extrinsic method was implemented in the gene finder EuGÈNE (as EuGÈNE-M) to detect AS evidence from aligned transcripts and generate alternative optimal gene predictions consistent with each detected AS event.
-
Has reproduction · 45
Identifying and classifying trait linked polymorphisms in non-reference species by walking coloured de bruijn graphs.
PMID 23536903 · PMC3607606 · PloS one · 2013 · 8 claims · 9 setups
Bubbleparse detects sequence variants directly from NGS reads without a reference genome, using the coloured de Bruijn graph implementation of Cortex plus a new depth-first bubble-finding module.
-
Has reproduction · 59
Comparing time series transcriptome data between plants using a network module finding algorithm.
PMID 31164912 · PMC6544932 · Plant methods · 2019 · 8 claims · 6 setups
Converting time-series expression data into co-expression networks and applying network module finding (OrthoClust) enables cross-species comparison without requiring one-to-one developmental stage mapping.
-
Full-text index only
Genomics, molecular imaging, bioinformatics, and bio-nano-info integration are synergistic components of translational medicine and personalized healthcare research.
PMID 18831773 · PMC3226104 · BMC genomics · 2008 · 8 claims · 8 setups
Genomics, molecular imaging, bioinformatics, and bio-nano-info integration are synergistic components of translational medicine and personalized healthcare
-
Full-text index only
The most frequent short sequences in non-coding DNA.
PMID 19966278 · PMC2831315 · Nucleic acids research · 2010 · 8 claims · 2 setups
Short frequent sequences (9-14 bases) in non-coding DNA may play a role in maintaining chromosome structure and function
-
Full-text index only
Genomic variability within an organism exposes its cell lineage tree.
PMID 16261192 · PMC1274291 · PLoS computational biology · 2005 · 8 claims · 5 setups
Somatic mutations accumulated during normal development implicitly encode an organism's entire cell lineage tree with very high precision.
-
Full-text index only
Duplication count distributions in DNA sequences.
PMID 19256873 · PMC3121164 · Physical review. E, Statistical, nonlinear, and soft matter physics · 2008 · 8 claims · 8 setups
Duplication count distributions N(c) for complex 40-mers show power-law-like decay for c roughly 3 to 50 (or higher) across human, C. elegans, A. thaliana, and D. melanogaster genomes.
-
Full-text index only
Comparing whole genomes using DNA microarrays.
PMID 18347592 · PMC7097741 · Nature reviews. Genetics · 2008 · 8 claims · 6 setups
DNA microarrays offer a relatively inexpensive and efficient alternative to genome sequencing for comparing all known classes of genomic diversity between closely related genomes.
-
Full-text index only
From microarrays to genome duplications.
PMID 12914655 · PMC193639 · Genome biology · 2003 · 8 claims · 8 setups
Gene3D shows that most genes across sequenced genomes can be assigned to known structural domain families, many of which are shared across kingdoms of life
-
Has reproduction · 85
An extensive evaluation of read trimming effects on Illumina NGS data analysis.
PMID 24376861 · PMC3871669 · PloS one · 2013 · 8 claims · 8 setups
Read trimming increases the quality and reliability of downstream NGS analyses (RNA-Seq mapping, SNP identification, genome assembly) while reducing execution time and computational resources.
-
Full-text index only
PeroxisomeDB: a database for the peroxisomal proteome, functional genomics and disease.
PMID 17135190 · PMC1747181 · Nucleic acids research · 2007 · 8 claims · 6 setups
PeroxisomeDB integrates the complete peroxisomal proteome of Homo sapiens and Saccharomyces cerevisiae into interrelated 'Genes', 'Functions', 'Metabolic pathways' and 'Diseases' sections with links to NCBI, ENSEMBL and UCSC
-
Full-text index only
Towards a comprehensive structural coverage of completed genomes: a structural genomics viewpoint.
PMID 17349043 · PMC1829165 · BMC bioinformatics · 2007 · 8 claims · 6 setups
A combined target-selection approach — pursuing both structurally uncharacterised domain families and additional targets from large structurally characterised superfamilies — is essential for comprehensive structural coverage of the genomes.
-
Full-text index only
Assessing the gene space in draft genomes.
PMID 19042974 · PMC2615622 · Nucleic acids research · 2009 · 6 claims · 7 setups
The proportion of mapped CEGs in a draft genome assembly is a useful metric for describing gene space completeness, complementing N50 and x-fold coverage.
-
Full-text index only
The protein-phosphatome of the human malaria parasite Plasmodium falciparum.
PMID 18793411 · PMC2559854 · BMC genomics · 2008 · 8 claims · 8 setups
P. falciparum possesses 27 putative protein phosphatase sequences across the four major PP families (PPP, PPM, PTP, NIF), plus 7 additional sequences predicted to dephosphorylate non-protein substrates, totaling 34.
-
Has reproduction
Systematic analysis of CNGCs in cotton and the positive role of GhCNGC32 and GhCNGC35 in salt tolerance.
PMID 35931984 · PMC9356423 · BMC genomics · 2022 · 7 claims · 8 setups
114 CNGC genes were identified across 4 cotton species (G. arboreum 20, G. raimondii 20, G. hirsutum 38, G. barbadense 36), clustering into 5 groups (I, II, III, IVa, IVb).
-
Has reproduction · 58
Genome-wide identification and characterization of germin-like protein family in Brassica juncea reveals their role against biotic stress.
PMID 41327045 · PMC12763953 · BMC plant biology · 2025 · 8 claims · 8 setups
102 GLPs were identified in B. juncea, 51 in B. nigra, and 48 in B. rapa via genome-wide in-silico analysis
-
Full-text index only
Molecular phylogeny of the kelch-repeat superfamily reveals an expansion of BTB/kelch proteins in animals.
PMID 13678422 · PMC222960 · BMC bioinformatics · 2003 · 8 claims · 8 setups
The human genome encodes at least 71 kelch-repeat proteins
-
Full-text index only
Genome informatics: taming the avalanche of genomic data.
PMID 15642109 · PMC549058 · Genome biology · 2005 · 8 claims · 7 setups
Ultraconserved regions (>100 bp, 100% conserved among mammals) exist in the genome and their function remains unknown