节点文献
一种高效的基于CHMM的哼唱式旋律检索方法
A CHMM Based Efficient Approach for Melody Retrieval by Humming
【Author】 CHEN Zhi-kun1,2, XU Ming2 and HUANG Yun-sen2 1(College of Information Engineering, Shenzhen University, 518060, China) 2(Information Center, Shenzhen University, 518060, China)
【机构】 深圳大学信息工程学院; 深圳大学信息中心;
【摘要】 本文以CHMM为基础进行音乐哼唱检索算法的研究,实现了模型的建立、模型的训练和旋律识别过程。与已有建模方法不同,本文利用从左到右、没有跳转的CHMM结构建立声学模型,使旋律模型得到简化,明显提高了识别效率。用经过音调转换的音高序列表示旋律特征,利用CHMM的二重随机特性隐含表示音长信息,从而避免了音符切分,使哼唱方式更自然。实验表明本文利用CHMM进行音乐哼唱检索的算法是高效的,并且取得了较高的识别率。
【Abstract】 Traditional retrieval method for music information is via the search of composer, performer, keywords of titles, singers or lyrics, etc., it is obviously, however, not a user friendly method for music retrieval, for these information is actually not the music content itself. Humans can identify musical just using melody alone, even in cases where the melody is cut short, corrupted, or rendered inaccurately. To avoid the traditional way of song retrieval, this paper presents a query by humming approach based on CHMM as its comparison engine. The approach facilitates the content-based song database retrieval via users’ acoustic inputs, which means that it allows the users to retrieve songs based on a few notes hummed naturally to the microphone. For efficiency purpose, the CHMM used is from left to right and no skipping. Furthermore, for this simple model, the training and the recognition method can be efficient. After pitch tracking, to guarantee better performance, we use a method named key transposition to shift the entire pitch vector to a suitable position that can generate maximum accumulative probability. To avoid note segmenting which constrains user to hum like “di di di” or “da da da”, we use frame level processing to train model and recognize melody, and “hide” the rhythm into the CHMM. Using VC 6.0 on a AMD 2500+ and 512M PC, we develop a training tool and a retrieval program. Whit the training tool, we construct 20 models (each one stands for a segment of melody), and use 800 recorded segments to train them. Use the retrieval program and other 400 recorded segments to test efficiency and correctness of the proposed approach. In the experimental results, retrieving one melody segment consumes 0.74s on average, which from pitch tracking (0.60s) and model recognize (0.14s). The recognition rate (TOP3) is 94.92%. The experimental results show that the proposed approach is efficiency and useful for query by humming application.
- 【会议录名称】 第三届和谐人机环境联合学术会议(HHME2007)论文集
- 【会议名称】第三届和谐人机环境联合学术会议(HHME2007)
- 【会议时间】2007-10
- 【会议地点】中国山东济南
- 【分类号】TP391.3
- 【主办单位】山东大学计算机科学与技术学院