节点文献

改进的ZCPA语音识别特征提取算法研究

Research on Improved Zcpa Speech Recognition Feature Extraction Algorithm

【作者】 焦志平

【导师】 张雪英;

【作者基本信息】 太原理工大学 , 信号与信息处理, 2005, 硕士

【摘要】 目前大多数语音识别系统在静音环境下具有较高的识别率,但在噪声环境下,系统的性能会严重下降,为了使语音识别系统实用化,抗噪语音识别研究具有重要意义。 人耳具有很强的识别能力,即使在噪声环境下也如此。因此研究人耳的听觉特性,进行语音特征参数的提取,有利于提高系统的鲁棒性。 本文围绕抗噪语音识别这个中心,完成了以下研究工作。 首先实现了具有过零峰值幅度(ZCPA:Zero-crossing with Peak Amplitude)特征的语音识别系统,它是基于人耳的听觉模型建立起来的。该模型通过分析和计算语音信号相邻上升过零点间的间隔,并将之分配到对应的频率箱,以此反映信号的频率信息;再通过检测相邻上升过零点间的峰值幅度并进行非线性压缩,对频率箱幅度进行加权。论文分析了该系统的抗噪性能,通过实验证明了这种系统的抗噪性能优于常用的由LPCC,MFCC作为识别特征的系统性能。 接着,论文以上述系统为基础,提出了改进ZCPA特征,

【Abstract】 Recently, most of speech recognition systems have much more recognition rates in the clean environment, however, the performances of these systems are severely degraded when there exists noisy environment. For the practicability of the speech recognition technology, it has the important significance to study the robustness of the speech recognition.The recognition ability of the human ear is very well, even in the noisy environment. So some researchers have devoted to the study of the auditory model to extract speech feature parameters that will improve the robustness of the system.This paper focuses on the robust noise speech recognition, completing the following research works.Firstly, this paper accomplished the speech recognition withthe Zero-crossings with Peak Amplitudes features, which is based on the auditory model of the human ear. This model reflects the speech signals’ frequency information by analyzing and computing the adjacent upward-going zero-crossing intervals, and allocates them to the corresponding frequency bins. Then it is detected that the peak amplitudes of the successive upward-going intervals. Finally, it goes along the compressive nonlinearity to weight the frequency bins’ peaks. This paper has also analyzed the robust noise performance. The results of many experiments showed that the robust noise performance of this system outperforms the other recognition system that uses the LPCC, MFCC as the recognition features.Secondly, this paper presented the improved ZCPA features based on the above system, namely, combining difference ZCPA features. These features use the characteristic of the difference signals, adding the difference information to the ZCPA features. The new features can extract the high frequency information mixed in the low frequency information, so the deficiency of the ZCPA features are made up for, and the improved recognition results are obtained.And this paper studied the front-end filters of this recognition system, introducing to use the Bark wavelet filters instead of the FIR filters. However, most of wavelet transforms, whether they aredyadic wavelets, wavelet packets or M-band wavelet transform, their frequency allocations all are based on octave relation, however, these frequency allocations are much more different with critical frequency band allocations of the human. So, if there is a wavelet that can allocate frequency according to the critical band, this wavelet will much more accord with the perception of the human to the speech and will improve the performance of the system. The basal thought of construction Bark wavelet is that the selected wavelet mother function satisfies the minimum of the time and bandwidth product, namely, the gauss function of the Bark fields, the mother wavelet has the equal bandwidth in the Bark fields. This paper analyzed the decomposition and reconstruction of this wavelet, presented the characteristic of time and frequency fields about this wavelet, and introduced the theory of this wavelet used in the front-end preprocessing.Finally, this paper simulated the speech recognition based on the Bark wavelet filters and ZCPA features, obtained the improved results and increased the recognition rates of the system.

  • 【分类号】TN912.3
  • 【被引频次】18
  • 【下载频次】357
节点文献中: