节点文献

高效简约的语音识别声学模型

Towards High Performance and Parsimonious Acoustic Modeling in Speech Recognition

【作者】 李小兵

【导师】 王仁华; 宋謌平;

【作者基本信息】 中国科学技术大学 , 信号与信息处理, 2006, 博士

【摘要】 当前连续密度HMM模型的语音识别系统性能良好,但其存储和计算需求过大。针对这一问题,本论文专注于语音识别系统的核心——声学模型。本文分别从训练方法、特征降维、模型参数压缩三个方面研究如何获得高效小巧的声学模型,在保证模型精度的前提下使用尽小可能的参数量,降低系统资源需求。基于已有的方法,我们提出及推广了一系列新方法,以实验证明了它们的有效性。这些方法主要集中在以下几个方面。 首先,本文研究了最小分类错误方法,实现了基于N-best解码的训练方法。实验证实,在保证模型精度的前提下,经MCE训练的模型可显著减小。我们并将其推广到子空间分布聚类HMM模型上,在很大程度上弥补了在将CDHMM转换成SDCHMM的过程中由于特征空间分裂和子空间分布聚类带来的性能降低。与直接由CDHMM转换而成的SDCHMM相比,性能提升15-80%。 其次,为了解决特征降维方法通常也降低识别性能的问题,我们提出了在区分性特征提取框架下按照最小分类错误准则调整模型参数和特征降维变换的方法,效果极为明显。更进一步,我们提出了以LDA变换执行的集去相关与降维于一体的新的特征提取方法,并将该方法同样纳入区分性特征提取框架之中。利用该方法,14维特征获得了与39维MFCC同样的性能,显著降低了计算和存储的需求。 再次,针对声学模型中各个状态对系统性能的贡献不同,提出了以贪心算法实现的基于似然度、Kullback-Leibler散度和状态间分散度的HMM模型各状态高斯分布数的确定方法。在总高斯分布数目给定前提下,分别最大化训练数据的似然度,最小化当前模型与“真正”模型之间的距离和最大化模型各状态间之分散度。其中基于状态间分散度的方法融入了状态间的竞争信息,具有区分性的特性。实验结果表明这几种方法相较基于贝叶斯信息准则的方法性能更佳。在相同模型精度的前提下,都可不同程度地减少参数。 最后,本文对声学模型特征级参数聚类进行了研究。在进行特征级参数聚类时我们提出采用具有信息熵意义的KLD作LBG聚类,聚类性能良好。而基于不同维的特征区分性信息多寡的不同,我们分别提出了各标量维高斯核的基于KLD和似然度的非均一分配法。在总高斯核数不变原则下,利用贪心算法在不同维之间进行高斯核的优化分配来最小化压缩模型与原始模型间的KLD和最大化训练数据的似然度。这两种非均一分配方法比均一分配性能更佳。而基于似然度的方法又优于基于KLD的方法。这些方法在保证模型性能基本不降的同时将模型参数压缩到原来的15%左右。此时加减需求为原来的50%左右,而乘除的需求则可大幅减少为1%以内。对于孤立词任务,相应的乘除运算更降到未压缩模型的0.05%左右。

【Abstract】 Current state-of-the-art, continuous density HMM-based large vocabulary speech recognition system delivers a fairly decent recognition performance in a benign environment but usually at a price of large memory and high computation complexities. In this thesis we explore the possibilities to obtain parsimonious acoustic model while maintaining the same performance as the complex model. They are explored in: 1) training algorithm; 2) dimensionality reduction; 3) model compression. Novel and efficient algorithms are proposed.In model training, the N-best based minimum classification error training is developed. Experimental results show that a high performance, parsimonious model can be obtained. This MCE is then extended to optimize subspace distribution clustering HMM. Experimental results show that performance degradation resulted from converting CDHMM to SDCHMM can be recovered and 15-80% word error rate reduction is obtained.In dimensionality reduction, we jointly optimize feature reduction transformation and the model parameters with MCE criterion. A new feature extraction, which uses LDA to perform feature decorrelation and dimensionality reduction, is proposed and developed into a discriminative feature extraction framework. A 14-dimension features gives almost the same performance as the 39-dimension MFCC features.In model compression, we found that different states contribute non-uniformly to recognition. Likelihood, Kullback-Leibler divergence, and state divergence are used to allocate Gaussian components to HMM states. The state divergence-based approach considers the discrimination of states. A greedy search is proposed to optimize Gaussian component allocation. Compared with Bayesian information criterion-based determination, the proposed approaches show improved performance.Also, we study feature-level model compression. Optimal clustering and non-uniform allocation of Gaussian kernels in the scalar feature dimension are proposed. Symmetric KLD is adopted to cluster Gaussian kernel, and KLD-based and likelihood-based non-uniform allocation are developed by using a Greedy search. Our non-uniform allocation gives better performance than uniform allocation, especially at larger compression ratios; likelihood-based allocation also outperforms KLD-based one. With almost negligible recognition performance degradation, the original HMMs can be compressed to 15% of its original size, which needs about 1% of the original multiplication/division operations. For the isolated-word recognition task tested, the multiplication/division operations can be further reduced to 0.05%.

  • 【分类号】TN912.34
  • 【被引频次】15
  • 【下载频次】1084
  • 攻读期成果
节点文献中: 

本文链接的文献网络图示:

本文的引文网络