节点文献
基于奇异值分解的低速率波形内插语音编码算法的研究
Research on Low Bit Rates Waveform Interpolation Speech Coding Based on Singular Value Decomposition
【作者】 王贵平;
【导师】 鲍长春;
【作者基本信息】 北京工业大学 , 信号与信息处理, 2005, 硕士
【摘要】 一直以来无线通信、卫星通信和语音存储的需求都在不断增长,随着通信和因特网技术的发展,又涌现了许多新的话音业务,所以语音压缩编码仍然是通信领域中的关键技术。本文的主要目标是在现有的波形内插(WI—Waveform Interpolation)语音编码算法的基础上,开发一个速率为2.4kbps 的WI 语音编码器,并用C 语言在计算机上模拟实现。传统WI 编码器将残差信号表示为渐变的特征波形(CW—Characteristic Waveform),然后分解为慢渐变波形(SEW—Slowly Evolving Waveform)和快渐变波形(REW—Rapidly Evolving Waveform),分别表示语音的准周期成分和类噪声成分。其中,CW 的分解是通过线性相位非因果FIR 低通滤波器沿着时间轴完成的,不仅增加了额外的一帧延时,同时也难以控制分解的精度。本论文正是针对以上的问题提出了一种基于奇异值分解(SVD-Singular Value Decomposition)的特征波形分解方法,减少了算法的延时,提高了分解精度。首先,为了降低计算复杂度,将CW 的幅度谱分块处理,分成基本矩阵、过渡矩阵和补充矩阵。其次,对基本矩阵进行SVD,按照编码比特数的要求由近似矩阵表示;对过渡矩阵采用离散余弦变换(DCT)近似表示;对补充矩阵通过计算各列均值粗糙表示。最后,对近似矩阵、DCT 系数和均值矢量量化。主观A/B 测试表明,基于上述分解与量化的2.4kbps SVD-WI 编码器的质量略好于2.4kbps MELP 编码器。
【Abstract】 Recently, there is more and more demand for wireless communication and speech storage. With the development of communication and Internet technology, lots of new voice-based services have appeared. Therefore, speech coding is still the key technique in the field of communication. The goal of this thesis is to develop a kind of 2.4kbps WI speech coder based on Waveform Interpolation (WI) coding, and have it simulated on computer by C programming language. In traditional WI coder, the residual signal is represented by an evolving characteristic waveform. The CW is decomposed into a slowly evolving waveform (SEW) and a rapidly evolving waveform (REW), representing the quasi-periodic and non-periodic components of speech signal, respectively. The decomposition of CW is accompolished along the time axis by non-causality linear FIR low-pass filter. Therefore, an additional delay of one frame is added and it’s hard to control the decomposed precision. According to above problems, the Singular Value Decomposition (SVD) is proposed in this thesis to decompose characteristic Waveform (CW), which reduced the delay of speech coder and improved the precision of decomposition. Firstly, in order to reduced the computational complexity, the magnitude spectrum of CW is divided into basic matrix、transitional matrix and supplemental matrix. Then, the basic matrix is decomposed by SVD, which is represented by underlying matrix according to the requirement of bit rates; and the transitional matrix and the supplemental matrix are approximated by caculating the mean of each column and Discrete Cosine Transform (DCT), respectively. Finally, the underlying matrix、the DCT coefficients and mean-value are vector quantized. Subjective A/B listening tests indicated that the reconstructed speech quality of the 2.4kbps SVD-WI codec is a little better than that of 2.4kbps MELP coder.
【Key words】 Speech Coding; Linear Prediction; Vector Quantization; Waveform Interpolation; Singular Value Decomposition;
- 【网络出版投稿人】 北京工业大学 【网络出版年期】2005年 06期
- 【分类号】TN912.3
- 【被引频次】9
- 【下载频次】156