节点文献
基于锚空间的音频场景识别方法研究
Research on Audio Scene Recognition Based on Anchor Space
【作者】 杨静;
【导师】 韩纪庆;
【作者基本信息】 哈尔滨工业大学 , 计算机科学与技术, 2011, 硕士
【摘要】 随着现代信息技术,特别是数字信号处理技术、网络多媒体技术的迅猛发展,越来越多的声音信号被数字化处理,并以各种音频格式存在。基于此,人们迫切地需要能够在音频数据流中对音频内容进行识别和理解的有效技术手段,从而高效地利用这些音频资源,并为各种智能系统提供基于声音的决策依据信息。音频场景是指语义上相关,时间上相邻的若干声学事件所组成的一个音频片段,此片段总是蕴含着高层抽象概念和特定的语义表达。音频场景识别是对音频语义内容高层次的识别和理解,该技术可广泛应用于信息内容安全、智能监控、无人驾驶车辆、智能会议室等领域。传统的音频场景识别方法,如高斯混合模型方法等,一般在短时上进行建模和识别,在长时上根据短时得分进行综合判决。这种方法忽略了声学内容在长时上的分布特性,且不适用于目标声学内容与非目标内容混杂的情况。本文提出了三种在长时上进行建模的锚空间音频场景识别方法,并设计了一个识别任务对这三种方法的性能进行了验证,在一段娱乐节目中根据音频寻找“令人激动”的场景片段,该场景一般对应较激烈的欢笑声和鼓掌声等。锚可以看作一个类别的原型表示,是根据信号产生的矢量到类别的一种映射关系。本文提出了三种面向音频场景的锚空间构造方法,并设计了相应的场景识别方法:1)基于状态变化统计量的锚空间音频场景识别方法。此方法将音频特征在时序上的变化量转化为若干变化状态,基于这些变化状态的统计信息张成锚空间,每个目标音频文件在此锚空间中映射成一个锚矢量,将此锚矢量当作目标场景的一个模板,从而构成目标场景库;2)基于高斯混合模型的锚空间音频场景识别方法。训练数据的目标音频文件训练得到目标高斯混合模型,集外音频文件训练得到集外高斯混合模型,基于各高斯分量的均值矢量张成锚空间,通过计算余弦距离将音频帧映射到锚空间中的一个点,求全部目标场景文件各帧在锚空间中的样本均值作为锚模板,目标场景由此锚模板表示;3)基于稀疏分解的锚空间音频场景识别方法。训练数据的目标音频文件训练得到目标字典,集外音频文件训练得到集外字典,基于其字典原子张成锚空间,稀疏分解得到的稀疏系数为此锚空间的坐标。实验数据为从网络上下载的娱乐节目,实验结果表明,三种基于锚空间的方法对节目中令人激动的场景都有很好的识别效果。特别是基于状态变化统计量的锚空间音频场景识别方法,其召回率达到85.67%时,其对应的错误接收率仅为9.57%。最后通过系统总结,提出了尚需完善和改进的方面。
【Abstract】 With the rapid development of modern information technology, especially network multimedia technology, digital signal processing technology, more and more voice signal is digitized, and stored in a variety of audio formats. Based on this, people urgently need an effective method to recognize the content of audio from the audio data streams in order to efficiently use the audio resources and supply suggestions for intelligent systems.Audio scene is an audio fragment composed of several acoustic events which are relevant in semantic and adjacent in time-domain. This audio fragment always contains high-level abstraction conception and specific semantic expression. Audio scene recognition is to recognize and understand the audio semantic content in high level, this technology is widely used in domains of information content security, smart surveillance, unmanned vehicle, smart meeting rooms, and so on. The traditional audio scene recognition methods, such as, Gaussian Mixture Model, model and recognize in short time, give a final response in long time according to the goodness in short time. This method not only neglects the distribution feature in long time of the acoustic content, but also fails to recognize the situation of chaos both in target acoustic event and non-target acoustic event. This paper proposes three audio scene recognition methods to model in long time based on anchor space, and designs a recognition task which to find the excited audio scene fragment in entertainment programs to test the performance of the three methods. The excited audio scene fragment means the intense applause acoustic events and laugh acoustic events.Anchor can be seen as a prototype of a category, a connection of vector formed by input signal mapping into category. This paper proposes three methods to construct the anchor space, and also designs corresponding audio scene recognition methods. Firstly, anchor space is based on the statistics of state changes. This method transforms the changed magnitude from audio feature in time sequences into changed state. The anchor space gets from the statistics of changed state. The projection of each target audio file forms one anchor vector which can be seen as one model of target scene, further more, all anchor vectors forms the target scene library; Secondly, anchor space is based on the Gaussian Mixture Model. The target audio data from the train data trains a target GMM while the non-target audio data trains a non-target GMM. The anchor space gets from the parameter of mean vectors of the two GMM, one audio frame can project to a point in this anchor space by cosine distance, the means of all points generated by all target audio frames can be seen as the target anchor model; Finally, the anchor space is based on the sparse decomposition. The target audio data from the train data learns a target dictionary while the non-target audio data learns a non-target dictionary. The anchor space gets from the atoms of the two dictionaries, the coefficients by sparse decomposition are the coordinates of this anchor space.The experimental data are entertainment programs downloaded from the Internet. The experimental results show that all three methods can recognize the excited scene from programs very well. Especially the method based on the statistics of the state changes, when the recall is 85.67%, the false alarm rate is only 9.57%. After systematic summarization, there is still a lot of room to improve.
【Key words】 Audio Scene Recognition; Anchor Space; Gaussian Mixture Model; Sparse Decomposition;
- 【网络出版投稿人】 哈尔滨工业大学 【网络出版年期】2012年 05期
- 【分类号】TN912.34
- 【下载频次】121