节点文献
基于深度学习的复杂背景下的语音增强算法
Speech Enhancement Algorithm Based on Deep Learning in Complex Background
【作者】 涂亮;
【导师】 范晔斌;
【作者基本信息】 华中科技大学 , 计算机技术, 2019, 硕士
【摘要】 传统的语音增强算法通常基于平稳噪声的假设,在复杂的背景下常常失效。基于深度学习的增强算法能很好地抑制非平稳噪声,但在不匹配的噪声环境下也出现性能下降。而提高模型的泛化性能需要大量的数据,这意味着更多的时间和计算资源。与现有技术相比,生成对抗网络(Generative adversarial network,GAN)被用于处理语音增强问题,经过网络架构的优化,可利用适量的训练数据以及常见的计算资源,构建在低信噪比及复杂噪声背景下通用的端到端模型。现有的大多数技术基于傅里叶分析,通常忽略相位信息而直接使用带噪语音的相位重建增强后的语音,这不利于提高低信噪比下的语音质量。利用生成对抗特性,在原始波形水平上操作,尝试利用波形中的细粒度信息(如相位、对齐等),可提高增强语音的质量,同时,由于卷积神经网络共享权值的特性,可实现更高的训练效率和更快地增强过程。优化后的模型相对于原始模型,在客观评估上得到了更为优异的效果。相对于基线DNN模型,具有与其竞争的性能,且更好的泛化性能。主观评估的结果表明,在所给定的特定的现实背景下,GAN得到了更多听众的偏好。
【Abstract】 Traditional speech enhancement algorithms are often based on the assumption of stationary noise and often fail in complex contexts.enhancement algorithm based on deep learning can suppress non-stationary noise well,but it also shows performance degradation in the unmatched environment.Improving generalization performance requires a lot of data,which means more time and computing resources.Compared with the prior art,this paper attempts to use the Generative adversarial network(GAN)to deal with the problems in speech enhancement,and uses an appropriate amount of training data and common computing resources to construct an end-to-end model,which is generally used for low signal-to-noise and complex noise background.Most of the existing techniques are based on Fourier analysis,which usually ignores the phase information and directly reconstructs the enhanced speech using the phase of the noisy speech,this is not conducive to improving the speech quality at low SNR.In this paper,we use the generative adversarial setting and operate at the original waveform level,try to use the fine-grained information in the waveform(such as phase,alignment,etc.)to obtain higher-quality speech.At the same time,the convolutional neural network is used to share weights and biases to achieve faster train and enhancement.Compared with the original model,the optimized model has a more excellent effect on objective evaluation.Relative to the baseline DNN model,it has performance competing with it,and better generalization performance.The results of the subjective assessment indicate that GAN has gained more audience preferences in the given specific context.
【Key words】 Speech enhancement; Deep learning; Generative adversarial network; Complex noise background; Low signal to noise ratio;
- 【网络出版投稿人】 华中科技大学 【网络出版年期】2020年 03期
- 【分类号】TN912.35;TP18
- 【被引频次】1
- 【下载频次】185