节点文献

基于小波和隐马尔可夫模型的音频分类

Audio Classification Based on Wavelet and Hidden Markov Mdoels

【作者】 王超

【导师】 吴亚锋;

【作者基本信息】 西北工业大学 , 环境工程, 2007, 硕士

【摘要】 音频处理在多媒体信息处理中占有重要地位。原始音频数据是一种非语义符号表示和非结构化的二进制流,如何提取音频中的结构化信息和内容语义是音频信息深度处理、基于内容的音频检索以及辅助视频分析等应用的关键。基于内容的音频分类作为解决音频结构化问题的核心技术,是当前音频内容自动分析领域的一个研究热点。 本文围绕音频分类的两大技术难点一特征分析与抽取以及分类器设计展开研究,主要内容如下: 概要地介绍了HMM的基本理论和主要算法。深入研究了语音、音乐的区别性特征及其计算方法,采用了音频clip和音频帧相结合的方法进行音频特征抽取。提出了一种基于各态历经混合高斯密度隐马尔可夫模型(EMGD HMM)的音频分类器,用于语音、音乐以及它们混合声音的分类。该分类器采用了全连接Markov链,从而能够有效地描述音频中的状态反复情况。对比实验结果表明,该分类器具有很高的分类精度。尝试了结合小波分析和傅立叶分析进行音频特征抽取,其中对子带能量和基音周期采用小波分析抽取,对频谱中心、带宽等特征则采用傅立叶分析抽取,并在本文提出的EMGD HMM音频分类器上进行了实验考察,结果表明该方法也是一种有效的音频特征抽取方法。

【Abstract】 Audio information processing plays an important role in multimedia applications. Raw audio data is non-semantic and non-structured binary stream, how to extract the structure information and semantic content from raw audio is crucial to deeper processing of audio information, content-based audio retrieval and video parsing with audio assistance. As the core technology of audio structuring, content-based audio classification is a current studied hotspot of audio content automatic analysis.This paper focuses on the two key points of audio classification: feature analysis and extraction, classification algorithm. Main aspects of the paper as follows:We introduce the basic theory and algorithms of Hidden Markov models briefly. Discriminating features between speech and music are deeply analyzed and calculated, which are extracted at frame-level and clip-level. A classifier based on EMGD_HMM (Ergodic Mixed Gaussian Density HMM) is proposed to classify speech, music, and their mixed audio. The classifier uses ergodic Markov chains, which can better describe the variety characteristic of audio states. Experimental comparisons show that the classifier has a high accuracy for audio classification. A method that collaborates wavelet analysis and Fourier analysis is used to extract audio features. Sub-band energy and pitch are extracted by wavelet, the others by Fourier. EMGD_HMM is used to evaluated the performance of feature set, experimental results show that the method is efficient for audio feature extraction.

  • 【分类号】TN912.3
  • 【被引频次】14
  • 【下载频次】612
节点文献中: 

本文链接的文献网络图示:

本文的引文网络