节点文献
基于深度生成模型的语音带宽扩展算法研究
Research on Speech Bandwidth Extension Algorithm Based on Deep Generative Models
【作者】 王静;
【作者基本信息】 大连理工大学 , 信息与通信工程, 2025, 硕士
【摘要】 语音带宽扩展技术旨在基于数字信号处理理论,通过扩展语音信号的频率范围来改善语音的自然度与清晰度,其广泛应用于语音通信、语音识别、语音合成等领域。近年来,随着深度学习方法的引进,语音带宽扩展技术取得了突破性进展,但在实际应用中仍然存在计算开销大、音质不稳定等问题。为此,本文基于深度生成模型对语音带宽扩展技术进行了深入研究,主要工作总结如下:(1)针对模型参数量大和时频域建模相位信息恢复问题,提出了一种基于时频卷积与因果Transformer块的语音带宽扩展模型。该模型采用生成对抗网络的训练架构,其中生成器用二维深度可分离卷积在时频域提取深度特征,用Transformer模型并行处理不同时间帧信息的依赖关系并引入因果掩码确保系统因果性,用双解码器架构分别映射目标语音频谱的实虚部以获得相位信息;采用多尺度STFT鉴别器与生成器进行对抗训练,在多个尺度上提升频谱相似度。实验结果表明,该模型在减少参数量的同时,在多项指标上表现出色,实现了良好的语音带宽扩展性能。(2)为捕捉语音感知关键特征和提高模型训练稳定性,提出了基于梅尔谱扩展与神经声码器的两阶段语音带宽扩展模型。该模型在第一阶段使用扩散概率模型,以低频带梅尔谱为输入生成宽带梅尔谱。扩散模型训练中,用类似Transformer中的正弦位置编码向量来实现噪声水平嵌入;将低频带梅尔谱作为条件输入,用双向循环扩张卷积增大感受野,为后续上采样提供更多有用信息。在第二阶段,以扩展后的梅尔谱作为输入,利用HiFi-GAN神经声码器合成原始波形,并在HiFi-GAN的多尺度鉴别器与多周期鉴别器架构基础上,引入多尺度STFT鉴别器进行生成器的训练。实验结果表明,该模型扩展后的语音具有良好的人耳听觉感知效果。
【Abstract】 Speech bandwidth extension technology aims to improve the naturalness and clarity of speech by extending the frequency range of speech signals based on digital signal processing theory.It is widely applied in speech communication,speech recognition,speech synthesis,and other fields.In recent years,with the introduction of deep learning methods,speech bandwidth extension technology has made breakthrough progress.However,in practical applications,there are still issues such as high computational overhead and unstable sound quality.To address these challenges,this thesis conducts in-depth research on speech bandwidth extension technology based on deep generative models.The main contributions are summarized as follows:(1)To address the issues of large model parameter size and phase information recovery in time-frequency domain modeling,a speech bandwidth extension model based on time-frequency convolution and causal Transformer blocks is proposed.This model adopts a generative adversarial network training framework,where the generator uses two-dimensional depth separable convolutions to extract deep features in the time-frequency domain.The Transformer model processes the dependencies between different time frames in parallel and introduces causal masking to ensure system causality.A dual-decoder architecture is used to separately map the real and imaginary components of the target speech spectrum to recover phase information.A multi-scale STFT discriminator and generator are used for adversarial training to enhance spectral similarity at multiple scales.Experimental results show that the model achieves excellent performance in voice bandwidth extension while reducing the number of parameters.(2)To capture key perceptual features of speech and improve model training stability,a two-stage speech bandwidth extension model based on Mel-spectrogram extension and neural vocoder is proposed.In the first stage,a diffusion probabilistic model is used to generate wideband Mel-spectrograms from low-frequency Mel-spectrograms as input.During training of the diffusion model,sinusoidal position encoding vectors,similar to those in the Transformer,are used to embed noise levels.The low-frequency Mel-spectrograms are used as conditional input,and bidirectional dilated convolutions are employed to increase the receptive field,providing more useful information for subsequent upsampling.In the second stage,the extended Mel-spectrogram is used as input to synthesize the original waveform using the HiFi-GAN neural vocoder.Based on the multi-scale discriminator and multi-period discriminator architecture of HiFi-GAN,a multi-scale STFT discriminator is introduced to train the generator.Experimental results show that the speech generated by this model exhibits good auditory perception quality.
- 【网络出版投稿人】 大连理工大学 【网络出版年期】2026年 04期
- 【分类号】TN912.3;TP18