节点文献

改进粒子滤波跟踪的视听双模态语音识别仿真

Simulation of Audiovisual Bimodal Speech Recognition Based on Improved Particle Filter Tracking

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 岳莉李柯景赵剑

【Author】 YUE Li;LI Ke-jing;ZHAO Jian;College of Computer Science and Technology, Changchun University;

【机构】 长春大学计算机科学技术学院

【摘要】 噪声环境下视听语音不易被识别,为提升语音识别效果,提出改进粒子滤波跟踪的视听双模态语音识别方法。采用谱减法去除噪声数据,完成视听双模态语音的消噪处理;根据人语和唇动信息之间的相关性,采用改进粒子滤波跟踪方法提取视听双模态语音特征信息,构建transformer语音识别模型,将提取的特征信息输入到模型内实施并行训练,实现视听双模态语音的有效识别。实验结果表明,通过对上述方法开展信噪比测试、识别性能测试,验证了上述方法的可行性高、可靠性强。

【Abstract】 In noisy environments, audio-visual speech is not easily recognized. To improve speech recognition performance, an improved particle filter tracking audio-visual bimodal speech recognition method is proposed. Firstly, spectral subtraction was adopted to remove noise data, thus completing the noising removal of audiovisual dual-modal speech. Based on the correlation between human speech and lip movement information, an improved particle filter tracking method was adopted to extract audiovisual dual-modal speech feature information, and then a transformer speech recognition model was constructed. Finally, the extracted information was input into the model for parallel training, thus achieving the effective recognition for audiovisual dual-modal speech. The experimental results show that the proposed method show high feasibility and strong reliability after the signal-to-noise ratio test and recognition performance test.

【基金】 吉林省教育厅科研项目(JJKH20220600KJ)
  • 【文献出处】 计算机仿真 ,Computer Simulation , 编辑部邮箱 ,2024年09期
  • 【分类号】TN912.34;TN713
  • 【下载频次】8
节点文献中: 

本文链接的文献网络图示:

本文的引文网络