Reconsidering the significance of genomic word frequency

Gregory Kucherov; Laurent No\'e; Mikl\'os Cs\H{u}r\"os

arxiv: q-bio/0609022 · v1 · submitted 2006-09-14 · 🧬 q-bio.GN

Reconsidering the significance of genomic word frequency

Mikl\'os Cs\H{u}r\"os , Laurent No\'e , Gregory Kucherov This is my paper

classification 🧬 q-bio.GN

keywords distributiongenomicsequencesignificancewordacrossallowsassessment

0 comments

read the original abstract

We propose that the distribution of DNA words in genomic sequences can be primarily characterized by a double Pareto-lognormal distribution, which explains lognormal and power-law features found across all known genomes. Such a distribution may be the result of completely random sequence evolution by duplication processes. The parametrization of genomic word frequencies allows for an assessment of significance for frequent or rare sequence motifs.

This paper has not been read by Pith yet.

Reconsidering the significance of genomic word frequency

discussion (0)