节点文献

波形内插语音编码算法中相位问题的研究

Research on Phase Information in Waveform Interpolation Speech Coding

【作者】 陈悦

【导师】 鲍长春;

【作者基本信息】 北京工业大学 , 信号与信息处理, 2006, 硕士

【摘要】 近年来低速率语音编码得到了巨大的发展,目前,在4kb/s以下速率实现具有通信质量的编码器已成为当今语音编码领域的一大研究热点。波形内插(WI——Waveform Interpolation)作为一种极具潜力的语音编码方法受到了人们的关注。在传统的低比特率语音编码中,考虑到人耳对相位信息不敏感而经常忽略相位信息,这将导致语音粗糙、刺耳甚至音调发生改变。为了获得高质量的声码器,语音的相位信息是不能不考虑的。本文基于感觉加权相位谱分析合成(AbS- Analysis-by-Synthesis)矢量量化方法,给出了一种WI编码器中慢渐变波形(SEW- Slowly Evolving Waveform)的相位信息量化及合成端相位的三次多项式插值重建方法。主观A/B测试结果显示,当用4~6比特量化相位信息时,该方法合成的语音质量明显好于固定相位法和倒谱法。此外,本文在此基础上提出了一种相位预测式矢量量化方案,使得女声的语音合成质量有所改进。另外,本文给出了一种改进的WI编码器合成方案。在该方案中,当帧间的基音周期连续变化时,语音残差信号由幅度谱和相位轨迹直接合成,而当基音周期发生跳变时,则利用相位过渡过程合成语音残差信号。该方法大大降低了WI解码器的复杂度,同时保证了合成语音质量没有变化。

【Abstract】 Recently, low bit-rate speech coding has developed greatly. The research on low bit-rate coder at 4kb/s and below with communication quality has become one of the most attention at present. Waveform Interpolation as a great potential speech coder has got much attention.In traditional low bit-rate speech coding, considering that ears are not sensitive to phase information, the phase information is often neglected, and this will result in coarse and harsh speech quality, and it even may lead to inflection in pitch. In order to obtain a high-quality speech codec, the phase information of speech should be included in codec. In this thesis, a method for quantizing the phase of SEW (Slowly Evolving Waveform) and reconstructing SEW’s phase with cubic polynomial interpolation is given based on the perceptual weighting analysis-by-synthesis (A-b-S) vector quantizer for the phase spectrum in WI coder. The subjective A/B listening tests indicate that the reconstructed speech’s quality of this scheme is better than that of fixed phase and Complex Cepstrum with 4-6 bits. Moreover, a predictive phase vector quantizer is proposed in this thesis based on the above method. The synthesis speech is slightly improved for female speakers.In addition, an improved synthesis scheme of WI coder is presented in this thesis. In this scheme, the speech residual signal is synthesized directly using magnitude spectrum and phase track where there are continuous changes for pitch between frames, while the speech residual signal is synthesized using phase intergradations with a burst of the pitch. The computing complexity of WI decoder is reduced greatly by this method, and the reconstructed speech quality keeps invariable meanwhile.

  • 【分类号】TN912.3
  • 【下载频次】132
节点文献中: