节点文献

基于响度特性加权的噪声下语音识别方法

A WEIGHTED METHOD FOR NOISY SPEECH RECOGNITION BASED ON LOUDNESS PROPERTY

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 蒋文建林耀荣韦岗

【Author】 Jiang Wenjian, Lin Yaorong, Wei Gang (Department of Electronic Engineering, South China University of Technology, Guangzhou 510640)

【机构】 华南理工大学电子与通信工程系

【摘要】 提出了一种将人耳听觉响度特性应用于噪声下语音识别的前端特征提取方法。本文使用分层遗传算法设计响度加权滤波器,对频谱进行听觉特性加权。将本方法应用于TIMIT数据包的E-SET在NoiseX92的各种噪声条件下的识别实验。实验结果表明,在各种噪声的不同信噪比下,对LPCC和MFCC特征,采用响度加权平均识别率分别有7%和11%的提高,证明本方法是有效的。

【Abstract】 Human Auditory System (HAS) is robust in speech perception. HAS can extract valid linguistic information from noisy speech which is stained by non-linguistic factors of information such as speaker variability, channel variability and the environment noise. The characteristics of HAS are widely employed in speech signal processing such as speech coding and speech recognition. A new weighted method for noisy speech recognition based on loudness property of human auditory system is presented in this paper. This method is used in the front-end parameterization of speech recognition system.Equal loudness contours are descriptions of the frequency dependence of the loudness of pure tones and the perceptive characertistic of human auditory system. We can deduce the frequency response of human auditory system from equal loudness contours. There is 5dB gain at the range of 3kHz-5kHz, so speech spectrum should be emphasized at this range of frequency. In this paper, we use Hierarchical Genetic Algorithm to design an IIR filter, which describe the loudness characteristic of HAS. The filter is used to weight speech spectrum. According to speech signal analysis, most of the signal energy is contained in its low frequency, but those low frequency components make little contribution to the speech articulation. Thus those components of speech which make little contribution to speech articulation can be canceled and those components of speech which make most contribution to speech articulation can be emphasized when loudness-weigh ted function is used. When noises stain some components of speech that human auditory system is not perceptible, these components will stain the whole speech feature vector if all components are dealt equally, and consequently the performance of speech recognition system will be greatly degraded. By using loudness-weighted function, we can restrict those components of speech which include little linguistic information and are stained by noises, and thus improve the performance of speech recognition system.In this paper, the new method is evaluated by a task on an English E-Set database, which includes nine easily confusable sounds, /b/, /c/, /d/, /e/, /g/, /p/, /t/, /v/ and /z/. White noise and F16 noise from Noi-seX92 database are added to the original speech to simulate noisy speech at different SNRs. MFCC and LPCC are extracted as speech features, respectively. Each sound is modeled by a three-state left-right Hidden Markov Model. Experimental results show that average increases of 7% and 11% in recognition accuracy rate are obtained by weighted LPCC and weighted MFCC respectively.

【关键词】 响度噪声语音识别
【Key words】 LoudnessNoiseSpeech Recognition
【基金】 国家自然科学基金;国家教育部博士点基金
  • 【文献出处】 模式识别与人工智能 ,Pattern Recognition and Artificial Intelligence , 编辑部邮箱 ,2001年02期
  • 【分类号】TN912.3
  • 【被引频次】9
  • 【下载频次】77
节点文献中: