节点文献
噪声环境下说话人识别研究
Research on Speaker Identification in Noisy Environment
【作者】 芮贤义;
【导师】 俞一彪;
【作者基本信息】 苏州大学 , 通信与信息系统, 2005, 硕士
【摘要】 说话人识别就是从说话人的一段语音中提取出说话人的个性特征,通过对这些个性特征的分析和识别,从而达到对说话人进行辨认或者确认的目的。矢量量化(VQ)方法是文本无关说话人识别中广泛应用的建模方法之一。在矢量量化过程中,经典的LBG 算法收敛速度快,但极易收敛于局部最优点,无法保证根据有限样本数据得到最优码本,并最终影响系统识别性能。考虑到遗传算法是一种具有全局化寻优搜索能力的算法,本文提出了遗传算法和K 均值算法相结合的综合分析方法GA-K 进行码本设计,改善码本的质量。讨论了具体的算法实现,分析了在不同的特征参数LPCC 及MFCC、不同测试语音长度下的说话人识别性能。实验结果表明,GA-K 方法优于传统的LBG 算法,可以很好地协调收敛性和识别率之间的关系。针对在干净语音环境下识别率很高的说话人识别系统,在噪声环境下识别率显著降低的缺点,本文结合具有多分辨率分析特点的小波变换技术,提出一种基于小波变换的鲁棒型特征提取算法,以提高说话人识别系统在噪声环境下的识别性能。实验结果表明,本文提出的鲁棒型特征提取算法可以有效地提高说话人识别系统在噪声环境下的识别性能。
【Abstract】 Speaker recognition is task of identifying or verifying who is speaking by analyzing and recognizing specific information abstracted from speech waves of speaker. Vector Quantization is one of popular methods for text-independent speaker identification at current. In the process of Vector Quantization, traditional LBG algorithm owns the advantage of fast convergence, but it is easy to get the local optimal result, so the codebook designed by LBG is not optimal and recognition performance will be influenced. According to the understanding that Genetic Algorithm has the capability of getting the global optimal result, a hybrid clustering method GA-K based on Genetic Algorithm and K-means algorithm is proposed to improve the codebook quality. Some feature parameters such as LPCC and MFCC are analyzed with GA-K in identification experiments via test voice utterance length. Experiments show the proposed GA-K method is effective. A new robust feature-extraction algorithm based on wavelet transform is proposed in this paper, because a speaker recognition system with high performance in clean environment will become deficient with unacceptable recognition performance in noisy environment. Benefit from its multi-resolution analysis abilities, the cepstrum features detected from several different time-frequency channels are integrated with a statistical entropy values. Experiments show that the proposed algorithm is quite efficient for speaker identification in noisy environment.
【Key words】 Speaker recognition; Vector quantization; Genetic algorithm; Wavelet transform; Robust features;
- 【网络出版投稿人】 苏州大学 【网络出版年期】2006年 04期
- 【分类号】TN912.34
- 【被引频次】8
- 【下载频次】212