节点文献

说话人识别系统设计研究

【作者】 刘刚

【导师】 张琴珠;

【作者基本信息】 华东师范大学 , 计算机应用技术, 2004, 硕士

【摘要】 说话人识别与语音识别一样,都是通过对所收到的语音信号进行处理,然后据此作出判断,不同之处在于说话人识别希望从语音中提取说话人的特征,加以利用,而语音识别刚好相反。近几十年来,特别是9.11事件以来,由于在安全访问控制、身份自动鉴别等相关领域的现实意义,说话人识别系统得到了大量的关注和研究。 说话人识别系统中,最重要的是说话人特征提取。本文详细地讨论了基音频率、线性预测系数及美倒谱系数提取方法,这些特征分别反映了说话人声带振动、声道响应的特性。之后,还必须对特征样本建模。本文涉及了一些说话人的语音模型,主要就这些模型的实现方法作了详尽分析,他们中包括矢量量化、混合高斯模型、人工神经网及隐马尔科夫模型。并且,从机器学习角度给这些建模方法以理论上的支持。 本文首次将数据挖掘概念引入说话人识别的建模研究。提出分析特征向量的分量相关度法,用线性回归、中心度量趋势及离散度量趋势描述特征向量中两两分量间的关系,得到相应说话人的模型。 而且,还从软件设计角度分析了说话人识别系统架构。引用了主体理论,定义了什么是主体,主体有什么特征,主体系统的功能,系统中主体间的通讯方式。然后,本文利用主体思想设计了说话人识别的实现架构,并详细描述了架构内主体间协同的规则。 本文从语音原理层次、具体算法设计层次、软件架构层次及所涉及机器学习理论和软件设计理论等多方面分析设计说话人识别系统,力图全面认识这个智能系统,从而能够利于它的开发实践。

【Abstract】 Like speech recognition, speaker recognition is that the computer dealing with the pronunciation signal received, then make the judgment in view of the above. Their difference consist that the aim of speaker recognition is drawing the speaker’s characteristic of the pronunciation and utilizing it while the one of speech recognition is just opposite. In near these decades, especially since 9.11 incidents, the speaker recognition system has received a large number of concerns and studying because of the realistic need of security access and identity automatically etc.In speaker recognition system, the most important thing is that speaker’s characteristic is drawn. This text discuss tonic frequency , Linear Prediction Coefficient and Mel-Frequency Cepstral Coefficient in detail, which reflect speaker’s characteristic of vocal cord vibration that sound channel respond separately. And then, the characteristic sample will be modeled. The text analyses and describes some modeling methods and their applications, which include VQ ,GMM , ANN and HMM. And they are thrown light on in theory from the angle of machine learning .And the text have also analysed speaker’s recognition system framework in terms of software design. It is defined what is a subject, which feature the subject has and the ways of communication among the subjects of the system are recounted. We design the system framework of speaker recognition utilizing the idea of subject while recount the rules of coordination among the subjects in the system framework.This text assays speaker recognition system from many aspects such as the principle level of the pronunciation, the level of concrete algorithm design , the level of software framework ,the conceptions of machine learning and the theory of software design to try to well know the intelligence system, thereby can be advantaged to its development in practice .

  • 【分类号】TP391.4
  • 【被引频次】1
  • 【下载频次】342
节点文献中: 

本文链接的文献网络图示:

本文的引文网络