节点文献

基于复数SepFormer与多尺度卷积的单通道语音增强方法

Monaural Speech Enhancement Method Based on Complex SepFormer and Multi-scale Convolution

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 李周刘庆华

【Author】 LI Zhou;LIU Qing-hua;Information and Communication College, Guilin University of Electronic and Technology;

【通讯作者】 刘庆华;

【机构】 桂林电子科技大学信息与通信学院

【摘要】 作为信号处理的重要研究方向,语音增强在人际交流、辅助破案和军事领域中扮演着至关重要的角色。针对现有深度学习方法未充分利用相位信息且特征提取单一导致增强语音质量下降的问题,提出了一种基于复数SepFormer与多尺度卷积的单通道语音增强方法(monaural speech enhancement based on complex SepFormer and multi-scale convolution, CS-MSC)。首先,引入了多尺度特征提取模块,改善传统方法特征提取单一的问题,并增强了模型对高频和低频特征的捕获能力,有效提升高频细节增强效果;其次,设计了带有通道注意力机制的跳跃连接,防止信息在深层网络中丢失,缓解深层网络中的梯度消失问题;并基于复数谱,在瓶颈层对振幅和相位的关联性进行建模,克服传统方法忽视相位信息的缺陷;最后,通过在编码器和解码器中添加增强频率轴特征表示的深度卷积提升模型在频率轴上的特征提取能力。实验结果表明,相比于门控卷积循环网络(gated convolution recurrent network, GCRN)、深度复数卷积循环网络(deep complex convolution recurrent network, DCCRN)等同类语音增强网络,本文提出的网络在极低信噪比条件下显著提高了语音信号的质量、可懂度以及信噪比;相较于DCCRN,在VoiceBank-Demand数据集上的语音感知质量和综合质量测度分别提升了31.72%和22.02%,表明该网络能够有效提升语音的可懂度和整体质量,具有较为突出的鲁棒性与泛化能力。

【Abstract】 As a critical research area in signal processing, speech enhancement plays a vital role in interpersonal communication, criminal investigation assistance, and military applications.To address limitations in existing deep learning approaches, including insufficient utilization of phase information and single-scale feature extraction leading to degraded speech quality, a monaural speech enhancement based on complex SepFormer and multi-scale convolution(CS-MSC) was proposed. To resolve the restricted feature diversity in conventional methods, a multi-scale feature extraction module was introduced to enhance the capture of both high-and low-frequency features and improve high-frequency detail reconstruction.Furthermore, skip connections with channel attention mechanisms were designed to mitigate information loss in deep networks and alleviate gradient vanishing. Additionally, complex-valued spectral representations were employed at the bottleneck layer to model amplitude-phase correlations, addressing the common neglect of phase information in conventional approaches.Finally, deepthwise convolution layers were incorporated into the encoder-decoder architecture to strengthen feature characterization along the frequency axis.Experimental results demonstrate that, compared to similar speech enhancement networks such as gated convolution recurrent network(GCRN) and deep complex convolution recurrent network(DCCRN), the proposed network significantly improves speech quality, intelligibility, and signal-to-noise ratio(SNR) under extremely low SNR conditions. Compared to DCCRN, the proposed network achieves improvements of 31.72% in perceptual evaluation of speech quality and 22.02% in composite quality measure on the VoiceBank-DEMAND dataset, effectively enhancing speech intelligibility and overall quality, and demonstrating notable robustness and generalization capabilities.

【基金】 广西自然科学基金(2025GXNSFAA069204);广西创新驱动发展专项(桂科AA21077008)
  • 【文献出处】 科学技术与工程 ,Science Technology and Engineering , 编辑部邮箱 ,2026年03期
  • 【分类号】TN912.35
  • 【下载频次】22
节点文献中: 

本文链接的文献网络图示:

本文的引文网络