节点文献

Video expression recognition based on frame-level attention mechanism

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 陈瑞TONG YingZHANG YiyeXU Bo

【Author】 CHEN Rui;TONG Ying;ZHANG Yiye;XU Bo;College of Information & Communication Engineering, Nanjing Institute of Technology;Jiangsu Future Network Innovation Research Institute;

【通讯作者】 TONG Ying;

【机构】 College of Information & Communication Engineering, Nanjing Institute of TechnologyJiangsu Future Network Innovation Research Institute

【摘要】 Facial expression recognition(FER) in video has attracted the increasing interest and many approaches have been made.The crucial problem of classifying a given video sequence into several basic emotions is how to fuse facial features of individual frames.In this paper, a frame-level attention module is integrated into an improved VGG-based frame work and a lightweight facial expression recognition method is proposed.The proposed network takes a sub video cut from an experimental video sequence as its input and generates a fixed-dimension representation.The VGG-based network with an enhanced branch embeds face images into feature vectors.The frame-level attention module learns weights which are used to adaptively aggregate the feature vectors to form a single discriminative video representation.Finally, a regression module outputs the classification results.The experimental results on CK+and AFEW databases show that the recognition rates of the proposed method can achieve the state-of-the-art performance.

【Abstract】 Facial expression recognition(FER) in video has attracted the increasing interest and many approaches have been made.The crucial problem of classifying a given video sequence into several basic emotions is how to fuse facial features of individual frames.In this paper, a frame-level attention module is integrated into an improved VGG-based frame work and a lightweight facial expression recognition method is proposed.The proposed network takes a sub video cut from an experimental video sequence as its input and generates a fixed-dimension representation.The VGG-based network with an enhanced branch embeds face images into feature vectors.The frame-level attention module learns weights which are used to adaptively aggregate the feature vectors to form a single discriminative video representation.Finally, a regression module outputs the classification results.The experimental results on CK+and AFEW databases show that the recognition rates of the proposed method can achieve the state-of-the-art performance.

【基金】 Supported by the Future Network Scientific Research Fund Project of Jiangsu Province (No. FNSRFP2021YB26);the Jiangsu Key R&D Fund on Social Development (No. BE2022789);the Science Foundation of Nanjing Institute of Technology (No. ZKJ202003)
  • 【文献出处】 High Technology Letters ,高技术通讯(英文版) , 编辑部邮箱 ,2023年02期
  • 【分类号】TP391.41
  • 【下载频次】6
节点文献中: 

本文链接的文献网络图示:

本文的引文网络