节点文献
基于语音参数和深度学习的极低速率语音编码算法研究
Research on Very Low Rate Speech Coding Algorithm Based on Speech Parameters and Deep Learning
【作者】 肖立;
【导师】 涂卫平;
【作者基本信息】 武汉大学 , 通信与信息系统, 2022, 硕士
【摘要】 随着北斗三号定位导航系统的全面建成与开放,其特有的短报文通信为普通移动网络无法覆盖的地方提供了信息交流的巨大便利,如:户外救援、偏远地区作业以及军事通信等等。但是北斗三号系统的通信信道极其狭窄,因此极低速率、高质量的语音编码研究具有迫切需求。但现有极低速率语音编码技术存在以下问题:(1)在极低速率下,传统基于模型的语音编码器受限于模型参数的准确度,解码语音质量一般较差。神经声码器作为极低速率下的解码器,有效提升了解码语音质量,但是基于现有神经声码器的语音编解码方案将编码器提取的全部参数作为神经声码器的输入信息,对于神经声码器来说,这些参数存在冗余;(2)当前基于端到端的极低速率编码器通过神经网络层提取隐向量特征,解码端利用隐向量合成语音,解码语音质量较高,但是用全新的编码系统替换已有的参数编码器,则存在着成本高昂的问题。针对上述问题,本文开展基于语音参数和深度学习的极低语音速率语音编码算法的研究,主要创新如下:(1)基于线性预测系数的极低速率语音编解码模型针对现有基于神经声码器的语音编解码方案中编码参数存在冗余的问题,本文研究基于Fast Speech2的线性预测系数-梅尔谱的转换模型,仅利用编码器提取的线性预测系数,生成神经声码器需要的梅尔谱,进而获得高质量的重建语音。由于编码端需要的量化编码的参数减少,故有效降低编码比特率。本文以Codec2 1.2 kb/s编码算法为例,编码器仅提取和量化线性预测系数,将编码比特率降低至0.675 kb/s。实验结果表明,相较于Codec2 1.2 kb/s编码器和LPCNet 1.6 kb/s编码器,本方法的主观MUSHRA平均分分别提高30左右和5左右。(2)基于梅尔谱参数的极低速率语音解码后处理模型针对全面更换语音编解码系统成本高的问题,本文研究基于HIFI-GAN的解码后处理模型,在不影响原有编解码器系统兼容性的前提下,建立解码语音后处理模型提升语音质量。后处理模型以解码语音的梅尔谱作为输入信息,以重构尽可能接近原始语音信号的增强语音为目的。为尽可能降低重构语音的失真度,本文提出多子带多分辨率STFT损失,用于学习语音信号频率成分的精细信息。本文在Codec 1.2 kb/s解码器后端添加上述后处理器,处理后的语音质量的主观平均MUSHRA分较Codec2 1.2kb/s和Lyra 3 kb/s分别提高约48和13。
【Abstract】 With the full completion and opening of the Beidou-3 positioning and navigation system,its unique short message communication provides great convenience for information exchange in places not covered by ordinary mobile networks,possible application scenarios include outdoor rescue,remote area operations,and military communication,etc.However,the communication channel of Bei Dou-3 system is extremely narrow,thus there is an urgent need for very low bit rate,high-quality speech coding research.However,the existing very low bit rate speech coding technology has the following problems:(1)At very low bit rates,the traditional model-based speech encoder is limited by the accuracy of model parameters,and the decoded speech quality is generally poor.Neural vocoder has effectively improved the quality of decoded speech as a decoder at very low bit rates.However,existing neural-vocoder-based speech coding and decoding schemes use all the parameters extracted by the encoder as input information to the neural vocoder,which are redundant to the neural vocoder.(2)The current end-to-end based very low bit rate encoder extracts the hidden vector features through the neural network layer,and the decoder uses the hidden vector to synthesize the speech with high decoded speech quality.But replacing the existing parametric encoder with a brand new encoding system has the high cost problem.To address the above problems,this thesis carries out the research of very low bit rate for speech coding algorithm based on speech parameters and deep learning.The main innovations are as follows.(1)Linear-prediction-coefficient-based codec model for very low bit rate speechTo address the problem of coding parameter redundancy in the existing neural-vocoder-based speech coding and decoding schemes,this thesis investigates the linear prediction coefficient-Mel spectrogram conversion model based on Fast Speech2,which only uses the linear prediction coefficients extracted from the encoder to generate the Mel spectrogram required by the neural vocoder,and then obtains high-quality reconstructed speech.Since the parameters of quantization coding are reduced at the encoding end,the coding bit rate is effectively reduced.In this thesis,the Codec2 1.2 kb/s coding algorithm is used as an example.The encoder extracts and quantizes only the linear prediction coefficients,reducing the coding bit rate to 0.675 kb/s.The experimental results show that compared with the Codec2 1.2 kb/s and LPCNet 1.6kb/s codec,the decoded speech quality of this method improves the subjective MUSHRA score by about30 and 5 respectively.(2)Post-processing model for decoding very low bit rate speech based on Mel spectral parametersTo address the high cost of a full replacement of the speech codec system,this thesis studies the decoding post-processing model based on HIFI-GAN,building a decoding speech post-processing model that improves the speech quality without affecting the compatibility of the original codec system.The postprocessing model takes the Mel spectrogram of the decoded speech as input information,and aims to reconstruct the enhanced speech as close to the original speech signal as possible.To reduce the distortion of the reconstructed speech as much as possible,this thesis proposes multi-subband multi-resolution STFT loss for learning the fine-grained information of the frequency components of the speech signal.In this thesis,the above post-processor is added to the back-end of Codec2 1.2 kb/s decoder,and the subjective MUSHRA score of the processed speech quality is improved by about 48 and 13 compared with Codec2 1.2 kb/s and Lyra 3 kb/s codec.
【Key words】 Speech coding at very low bit rate; Deep learning; Linear prediction coefficients; Mel spectrograms;