节点文献

多通道语音增强与语音盲分离算法的研究

Study of Speech Enhancement and Blind Separation Algorithm with Multi-input

【作者】 欧世峰

【导师】 赵晓晖;

【作者基本信息】 吉林大学 , 信号与信息处理, 2005, 硕士

【摘要】 本文提出了一种频域语音活动性检测(VAD)算法。该方法通过短时傅立叶变换把信号在时域中的卷积混合形式转化为频域中的瞬时混合形式,然后利用峭度的特性对语音信号进行VAD 判决。该方法可在对信号源方向及位置等空间特性没有任何先验知识的情况下,应用于复杂背景噪声下的语音活动性检测,且方法简单有效。通过同时对角化麦克风阵列接收信号中语音信号和噪声信号的全局协方差矩阵来推导纯净语音信号子空间,本文改进了一种基于信号子空间分解方法的多通道语音增强算法。算法弥补了原始算法只限于白噪声背景下语音增强的不足,实现了色噪声背景下语音信号的最优估计。本文对独立分量分析技术和语音盲分离算法进行了全面细致的研究,详细地分析和讨论了语音信号频域盲分离算法中所涉及的几个关键问题,并对它们进行了合理的解释,为以后的工作打下了良好的基础。

【Abstract】 In our lives, Speech is often corrupted acoustically by ambient noise whichproduces aesthetically undesirable effects on the performance of digital voiceprocessor and even diminishes communication system ability to conveyinformation across the interface. Therefore a speech enhancement system isstrongly needed whose responsibility is improving the speech quality and ensuringreliability of digital voice communication systems.Depending on the number of microphone, speech enhancement algorithmscan be divided into two categories, one is single-input and other is multiple-inputalgorithm. There is only one channel of speech signal needed in conventionalsingle-input system which is banded with the attributes of speech signals and basedon the attenuation of the varieties of noise, such as spectral magnitude subtraction,adaptive noise canceling and auditory masking and so on. These algorithms whichare easy to realize with hardware and simple for construction have been applied tomodern communication systems. While capable of improving speech quality inrestrictive environments (additive noise, no multiple channel, high to moderatesignal-to-noise ratio (SNR), single source), these approaches do not perform well inthe cases of reverberant distortions, competing sources, and severe noise conditions.In recent years, the use of microphone arrays has received considerable attenuationas a means for dramatically improving the performance of traditional singlemicrophone systems and amounts of approaches have been proposed, such as delayand sum beamformer, adaptive beamformer, and speech signal model etc..In addition, there is a very important accessory measure, namely voice activitydetection (VAD) in digital speech processing. By use of VAD, we can obtain oranew the statistical attributes of noise demanded in speech signal enhancementtechnique which is based on distinguishing the inactive interval of speech signaland after that, we can better trace the change of speech signal and improve theenhancement of speech signal at last. The essence of VAD is subtracting theattributes of the speech signal which are different from the ones of noise. In general,the parameters of VAD includes short time energy、autocorrelation of speechsignals and short time zero crossing.After general presentation of the algorithms for speech signal enhancementand VAD, based on array signal processing technique and high order statistics, avoice activity detection algorithm is proposed, which transforms convolutive mixedsignals in time domain to the instantaneous mixed in frequency domain using ashort-time discrete Fourier transform. Then voice activity detection in accordancewith certain characteristics of HOS is achieved. The proposed algorithm is not assame as other VAD algorithms with multi-input. It can be applied in a complexnoise environment without a priori knowledge on the direction and location of thesource signal. Furthermore, it is simple and effective. Simulation resultsdemonstrate that the proposed algorithm possesses good performance withstationary noise at –10dB SNR and keeps its performance in non-stationary noise at0dB SNR. An improved speech enhancement algorithm based on signal subspace withmulti-input is presented in chapter four. Through simultaneous diagonalization ofthe overall covariance matrices of clean speech and noise signal observed bymicrophone array, clean speech signal subspace is estimated without anypreassumption on the stochastic property of noise signal. The proposed methoddoes not rely on any signal model and attains the optimal estimation of speechcorrupted by colored noise, which overcomes the disadvantage of original methodonly suitable for white noise case. Simulation results demonstrate that thealgorithm possesses good performance both in objective and subjective tests. Blind signal separation (BSS) is a hot research topic in the signal processingand has been applied to lot of different fields. It can recover the source signals bymeans of the sensor signals, though there is any transcendental information forneither source signal nor transmitting channel. This proceeding can be named asindependent component analysis (ICA), too. At present, the algorithms of ICA include: second order statistics blindidentification based temporal structure of sources; learning algorithms based highorder statistics and blind source separation based on information theory. The formertwo among the three algorithms are the classical ones which use the eigenvaluedecomposition(EVD)generalized eigenvalue decomposition (GEVD) and CentralLimit Theorem as its radical principles and accomplish the blind separation ofseveral source signal under the satisfied assumption in the processing. Thealgorithms based on information theory are so popular recently that most people

  • 【网络出版投稿人】 吉林大学
  • 【网络出版年期】2005年 06期
  • 【分类号】TN912.3
  • 【被引频次】6
  • 【下载频次】764
节点文献中: