节点文献
基于Kinect辅助的机器人带噪语音识别
Automatic Speech Recognition integrating Kinect Sensor for a Robot under Ego Noises
【Author】 WANG Jianrong;GAO Yongchun;ZHANG Ju;WEI Jianguo;DANG Jianwu;School of Computer Science and Technology, Tianjin University;School of Computer Software, Tianjin University;
【机构】 天津大学计算机科学与技术学院; 天津大学软件学院;
【摘要】 音视频信息融合可以提升机器人在噪声环境下的语音识别性能。然而,受说话者的头部旋转、唇部尺寸不一、距摄像头距离不固定以及光照等因素影响,使得唇部信息不能得到有效的全面表征。为此,本文提出了融合机器人与Kinect的多模态系统。该系统采用Kinect获取3D数据和视觉信息,并使用3D数据重构侧唇,以此来补充音视频信息。一系列基于特征融合和决策融合方法的结果表明,本文提出的多模态系统优于基于音视频单流和双流的语音识别效果,能够辅助机器人自身噪声下的语音识别。
【Abstract】 Audio-visual integration is an effective way to improve the performance of automatic speech recognition for the robot under ego noises. However, the head rotation, difference of the lips, camera-subject distance and the lighting variations make the performance of ASR degrades. This paper combined the robot with Kinect sensor, and proposed a novel multi-modal system for the robot. 3D data and visual information were obtained from the Kinect. The profile lips were rebuilt utilizing the 3D data to implements the information of video. Different fusion methods were investigated to incorporate the available multimodal information. A series of experiments under ego noises of the robot were presented to demonstrate the novel multi-modal system is superior to the traditional automatic audio and audio-visual speech recognition,and it can improve the robustness of ASR for a robot.
【Key words】 humanoid robot; ego noises; automatic speech recognition; Kinect multi-sensor; multi-modal system;
- 【会议录名称】 第十三届全国人机语音通讯学术会议(NCMMSC2015)论文集
- 【会议名称】第十三届全国人机语音通讯学术会议(NCMMSC2015)
- 【会议时间】2015-10-25
- 【会议地点】中国天津
- 【分类号】TP242;TN912.34
- 【主办单位】中国中文信息学会语音信息专业委员会