节点文献

说话人识别系统研究

【作者】 李轶

【导师】 童勤业;

【作者基本信息】 浙江大学 , 生物医学工程, 2003, 硕士

【摘要】 以前往往采用线性的方法如频谱分析来分析语音信号,而这些线性方法只适用于平稳的、一致的、平衡的线性的时间序列,对于非平稳、不一致、非平衡的非线性的语音时间序列,这些传统的线性方法就往往丢掉了许多蕴涵本质的重要信息。语音信号证明是非线性的混沌的,本文在此基础上提出了一种能够有效反映语音信号非线性特征的处理方法,并由此提出一种能够反映说话人个人特性的KC复杂性特征,作为说话人识别系统的一种新的特征,经实验证明这种特征是简洁合理并且是有效的。我们认为,复杂性分析方法能够用于说话人识别特征分析,具有很宽广的应用前景,这是很有前途的研究课题。 本文基于传统的LPC倒谱特征和KC复杂性特征建立了一个说话人确认系统,采用了YOHO speaker verification数据库,Enroll阶段:采用138说话人4个session每个session有10个语音样本数据,Verify阶段:采用138说话人10个session每个session有4个语音样本数据,训练模板和测试该说话人确认系统,取得了较好的说话人确认效果。 实验数据分析证明:因为传统的基于线性理论基础上的传统的说话人识别特征提取方法(例如:LPC倒谱特征),与基于非线性理论基础上的KC复杂性特征基本无相关性,因此这两类特征如能相互结合,有着良好的互补特性,能够大幅度的提高系统性能。这说明新提出的说话人的KC复杂性特征即使不能单独应用于说话人识别系统,也能够成为一个非常有用的传统线性特征的有效辅助特征。

【Abstract】 The Automatic speaker verification in this paper introduces one novel feature selection and extraction method, computation complexity features. As traditional linear features are mainly based on frequency analysis, and the assumptions used to extract traditional linear features do not describe the nonlinear dynamic evolution of the system, merely applicable to the steady, coherent and balanced linear time series, they generally ignored the most important information, which contained the essence of the unsteady, incoherent and unbalanced nonlinear time series of speech. Computation complexity feature can extract that nonlinear character of speech signal, which overcomes the disadvantage of the traditional linear feature extraction method. In this paper, the computation complexity theory is applied to feature extraction of speech signal; this is a creative thought and trial.This automatic speaker verification system mainly included three modules: speech signal preprocess, features extraction (traditional feature and complexity feature), pattern recognition (distance matching). The primary database for this work is known as the YOHO Speaker-Verification Corpus, which was collected by ITT under a U.S government contract. The YOHO data base was the first large-scale, scientifically controlled and collected, high-quality speech data base for speaker-verification testing at high confidence levels. Through the experiments, one conclusion comes that the combination of the nonlinear features and traditional linear features can reduce the verification errors markedly and gain more accurate results.This work suggests new ideas to construct speaker recognition systems more robust and reliable. Extract new information that specifically distinguish different speakers is very important to continue the development of this area. In the other hand, the introduction of new techniques and new features to characterize a speaker will bring an intrinsic computational processing overhead. Many applications where the speaker recognition technology can potentially be introduced are still searching for more accurate systems. The nonlinear dynamic analysis can analyze the speech production differently, as the result of a nonlinear dynamic process, bringing up new information to characterize it in a more complete way.

  • 【网络出版投稿人】 浙江大学
  • 【网络出版年期】2003年 03期
  • 【分类号】TN912.3
  • 【被引频次】6
  • 【下载频次】380
节点文献中: 

本文链接的文献网络图示:

本文的引文网络