节点文献

融合生理信息的多模态唇读技术研究

Research on Technology of Lip Reading Fused Physiological Information

【作者】 杨帆

【导师】 魏建国;

【作者基本信息】 天津大学 , 计算机科学与技术, 2018, 硕士

【摘要】 作为人与计算机或者其他设备沟通的桥梁,人机交互技术在“智能化”科技和需求的双轮驱动下,经历了单纯从鼠标、键盘的接触式交互到多模态信息非接触式交互的重大变革。作为重要的非接触式交互方式,唇读技术不仅突破了应用场景的限制,在噪声环境下辅助语音识别,且随着三维空间体感传感器的出现,唇读技术有了更广阔的发展前景。对唇部运动信息的全面提取和有效表征直接关系着语义信息的准确表达,唇动特征提取的完备性和代表性直接影响着语义内容的识别以及语义情感的判断。对于唇动特征提取,当前所存在的共同的难点在于:对于人们说话方式的巨大差异,所采用的特征提取方法无法作为一种通用的方法来全面有效地表征唇动信息。为此,本论文旨在研究融合面部肌肉生理信息的多模态唇动识别,研究内容主要包括基于Kinect的多模态数据采集、预处理、面部肌肉模型建立、肌肉模型映射、特征提取和基于DenseNet的训练识别。首先,基于Kinect V2.0,采集了话者唇动过程中的多模态信息,包括音频、彩色图像和深度数据。数据采集完成后,进行了一系列的数据预处理操作。对图像数据,分别进行了人脸检测、唇部定位和数据扩张。对深度数据,纠正了话者录制过程中扭头、歪头、仰头、低头等一系列不自觉的头部运动。然后,论文研究了面部肌肉生理信息,并利用少量参数建立了向量肌的几何模型。根据获取到的1347个面部特征点,将建立的肌肉模型映射到了三维面部模型上。基于建立好的肌肉模型,论文提取了两类特征,分别为几何特征和生理特征,几何特征包括形状特征和角度特征,生理特征包括肌肉长度特征和肌肉位移特征。最后,论文用DenseNet进行了唇读实验,证明了深度信息的加入可以提高唇读系统的识别率,以及论文所提出的生理特征可以增强三维离散点之间的约束,更全面的表征唇动过程。此外,论文对声调和辅音进行研究,发现了仅通过视觉信息来区分声调和辅音的可行性。

【Abstract】 As a bridge between people and computers or other devices,human-computer interaction technology has experienced a significant change from mouse and keyboard to non-contact interaction of multi-modal information under the drive of intelligence technology and demand.As an important non-contact interaction method,lip-reading technology has not only broke through the limitations of application scenarios,assists speech recognition in noisy environments,but also has a broader development prospect with the emergence of three-dimensional sensor.The comprehensive extraction and effective characterization of lip motion information is directly related to the accurate expression of semantic information.The completeness and representation of lip-motion feature extraction directly affect the recognition of semantic content and the judgment of semantic emotion.For lip-motion feature extraction,the common difficulty is that the feature extraction method can’t be used as a general method to comprehensively and effectively represent lip-motion information.So,this paper aims to study multimodal lip reading studies integrating facial muscle physiology information.The research content mainly includes Kinect-based multimodal data acquisition,preprocessing,facial muscle model building,muscle model mapping,feature extraction and DenseNet-based training recognition.First,multi-modal information including audio,color image and depth data were collected based on Kinect V2.0 during the lip movement of the speaker.After that,a series of pre-processing operations were performed on the data.For the image data,face detection,lip positioning,and data augmentation were sequentially performed.For the depth data,a series of unconscious head movements such as turning,hoeing,looking up,bowing,etc.during the recording of the speaker are corrected.Then,the paper studied the facial muscle physiological information and establishes a vector muscle model with a small number of parameters.Based on the acquired 1347 facial feature points,the established muscle model was mapped into the three-dimensional facial model.Based on the established muscle model,the paper extracted two types of features,namely geometric feature and physiological feature.Geometric feature includes shape feature and angle feature,physiological feature includes muscle length feature and muscle displacement feature.Finally,the paper used DenseNet for a lip reading experiment.The discovery proved that the addition of depth information can improve the recognition rate of the lip reading system,and the physiological characteristics proposed in the paper can indeed enhance the constraint between three-dimensional discrete points and more fully characterize the lip movement process.In addition,the paper studied the tones and consonants,and found that it is feasible to distinguish tones and constants by only visual information.

  • 【网络出版投稿人】 天津大学
  • 【网络出版年期】2020年 06期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络