节点文献

HMM和神经网络用于语音识别的算法研究

Study of Speech Recognition Algorithm Based on HMM and Neural Network

【作者】 赵姝彦

【导师】 张雪英;

【作者基本信息】 太原理工大学 , 计算机应用技术, 2005, 硕士

【摘要】 语音识别是语音信号处理领域的研究热点,它长期以来一直是一项难题,尤其是对于噪声环境下以及非特定人语音识别。为此,本文讨论了几种常用的语音识别方法包括经典的隐马尔可夫模型以及目前比较流行的人工神经网络,并引入一种新的用于抗噪的特征参数:过零率与峰值幅度特征(简称ZCPA特征)组成一个鲁棒的语音识别系统。 文中首先介绍了几种常用的特征提取方法如线性预测倒谱系数(LPCC)和Mel频率倒谱系数(MFCC),这两种特征在静音环境下有很好的识别效果,但在噪音环境下,性能就会严重下降。为此论文重点介绍了一种抗噪特征:ZCPA特征,并分析了其抗噪原理。接下来论文讨论了隐马尔可夫模型的原理及用于语音识别的系统实现过程。经典的Baum-Welch训练算法在软件实现中存在下溢问题,文献中没有给出正确的针对下溢问题的重估公式。因此,论文使用定标算法,重新推导了Baum-Welch训练算法的重估公式。实验结果表明修改后的公式收敛速度很快,并且得到了较好的识别效果,充分证明了重新推导后公式的正确性,而使用原公式在训练时无法收敛。然后

【Abstract】 Speech recognition is the research hotspot in the field of speech signal processing. It has been a difficult problem for a long time, especially for the recognition of person-independent and in noisy environment. This paper discussed several common speech recognition methods including classical Hidden Markov Model and artificial neural network which is very popular currently. It also introduced a new anti-noise feature parameter, Zero-crossings with Peak-amplitudes feature (ZCPA feature), which can be used to construct a robust speech recognition system.This paper presented several familiar feature extracting methods such as Linear Prediction Cepstrum Coefficient (LPCC) and Mel Frequency Cepstrum Coefficient (MFCC). They have got excellent recognition results under clean environment, but their performance will deteriorate severely in noisy condition. So most part is devoted to introduce ZCPA feature and analyze its anti-noise principle. Then this paper discussed HMM theory which is used in speech recognition and its implementation process. There areunderflow problems in software implementation procedure for the classical Baum-Welch training algorithm, and a lot of literatures did not presented an explicit method. With respect to this problem, this paper inducted the scaling algorithm and derived the reestimate formulae of Baum-Welch algorithm again. The experiments showed that it can converge rapidly and the recognition results are good, which proved the correctness of the new reestimate formulae, while the old formulae can not converge in experiments. Then the paper studied several feed-forward neural networks used in classification including BP network, RBF network and wavelet network. It discussed their theories, learning process and the modeling method for speech recognition respectively. Centroid selecting of RBF hidden nodes has great influence for the network performance. The common K-means clustering is a kind of unsupervised learning method, the paper proposed to cluster the input data by the classification information of training samples and calculate their centroids to be the centers of each hidden function. Experiment results showed that the recognition rate by selecting the centroids of hidden functions supervised is better than K-means clustering method. Finally, the paper introduced wavelet transform theory, and the Gaussian basis function of RBF network was taken place by a wavelet basis function, so a wavelet neural network can be formed. Experiments showed that the wavelet network also can get excellent recognition performance.

  • 【分类号】TN912.3
  • 【被引频次】19
  • 【下载频次】778
节点文献中: 

本文链接的文献网络图示:

本文的引文网络