节点文献

基于长短时记忆网络的多模态情感识别和空间标注

Real-time Multimodal Emotion Recognition and Emotion Space Labeling Using LSTM Networks

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 刘菁菁吴晓峰

【Author】 LIU Jingjing;WU Xiaofeng;Image and Intelligent Information Processing Lab,School of Information Science and Technology,Fudan University;

【通讯作者】 吴晓峰;

【机构】 复旦大学信息科学与工程学院图像与智能信息处理实验室

【摘要】 情感计算中音/视频的情感识别对人机交互等领域的深层次认知具有重要应用价值,在现代远程教育中可作为教学过程性实时评估的重要技术之一.为克服单一模态模型识别精度依赖于情感类型这一问题,本文提出一种基于长短时记忆(LSTM)网络的多模态情感识别模型,采用双路LSTM分别模拟人类听觉和视觉处理通路处理语音和面部表情的情感信息,在eNTERFACE’05双模态情感数据集上进行训练和测试,并模拟人脑边缘系统情感区进行决策层加权特征融合,传统情绪六分类标准的准确率可达74.7%.同时,考虑到传统离散情绪六分类法无法进行程度度量,且存在外在表现相似和多情感同时并存的问题,本文提出一种新的多模态情感识别模型的空间标注法,采用模型层特征融合方法将情感分类特征映射到激活度-效价空间(Arousal-Valence Space),从而更好刻画情感的程度,实验结果显示准确率在空间两个维度上分别达到84.1%和86.6%.相比于已有的大多数相关研究,本文提出的模型运算量小,识别精度高,可进行实时在线情感识别.

【Abstract】 Emotion recognition of audio/video have important application value in the field of in deep-level cognition such as human-computer interaction.It can be used as one of the important technologies for real-time assessment of teaching process in modern distance education.In order to overcome the problems of limited emotional information collected by single-modal emotion recognition,this paper proposes a multimodal emotion recognition model based on Long Short-Term Memory(LSTM)networks.The model uses dual-channel LSTM to simulate human hearing and visual processing channels to process speech and facial expression information,and is trained and tested on the eNTERFACE’05 dual-modal emotion dataset.While stimulating the emotional region of human brain’s limbic system,the accuracy rate under weighted decision-level feature fusion is 74.7%.Considering that the traditional 6-categories classification of discrete emotions cannot measure the degree of emotion,and it don’t consider the coexistence of multiple emotions.This paper proposes a new multimodal emotion recognition model for two-dimensional space labeling.The model-level feature fusion method is used to map the emotion classification results to the Arousal-Valence Space.The accuracy rate achieves 84.1%and 86.6%separately in each dimension.Compared with most related studies,the model proposed in this paper has a small amount of calculation and high recognition accuracy,and can be used for real-time online emotion recognition.

  • 【文献出处】 复旦学报(自然科学版) ,Journal of Fudan University(Natural Science) , 编辑部邮箱 ,2020年05期
  • 【分类号】TP18;TN912.3;TP391.41
  • 【被引频次】6
  • 【下载频次】528
节点文献中: 

本文链接的文献网络图示:

本文的引文网络