节点文献

一种轻量级全频带语音增强网络模型

A Light-Weight Full-Band Speech Enhancement Model

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 胡沁雯侯仲舒乐笑怀卢晶

【Author】 HU Qinwen;HOU Zhongshu;LE Xiaohuai;LU Jing;Key Laboratory of Modern Acoustics, Institute of Acoustics, Nanjing University;

【通讯作者】 卢晶;

【机构】 南京大学声学研究所近代声学教育部重点实验室

【摘要】 基于深度神经网络的全频带语音增强系统面临着计算资源需求高以及语音在各频段分布不平衡的困难。本文提出了一种轻量级全频带网络模型。该模型在双路径卷积循环网络模型的基础上,利用可学习的频谱压缩映射对高频段频谱信息进行有效压缩,同时利用多头注意力机制对频域的全局信息进行建模。实验结果表明本文模型只需0.89×10~6的参数即可实现有效的全频带语音增强,验证了本文模型的有效性。

【Abstract】 Deep neural network based full-band speech enhancement systems face challenges of high demand of computational resources and imbalanced frequency distribution. In this paper, a light-weight fullband model is proposed based on dual path convolutional recurrent network with two dedicated strategies, i. e., a learnable spectral compression mapping for more effective high-band spectral information compression, and the utilization of the multi-head attention mechanism for more effective modeling of the global spectral pattern. Experiments validate the efficacy of the proposed strategies and show that the proposed model achieves competitive performance with only 0.89×10~6 parameters.

【基金】 国家自然科学基金(12274221)
  • 【文献出处】 数据采集与处理 ,Journal of Data Acquisition and Processing , 编辑部邮箱 ,2023年02期
  • 【分类号】TN912.35;TP183
  • 【下载频次】30
节点文献中: 

本文链接的文献网络图示:

本文的引文网络