节点文献

面向复杂场景下多模态情绪识别方法研究

Research on Multimodal Emotion Recognition Methods in Complex Scenarios

【作者】 陈闯;

【导师】 涂铮铮; 孙晓;

【作者基本信息】 安徽大学 , 电子信息(专业学位), 2024, 硕士

【摘要】 情感计算是计算机科学与心理学的交叉研究课题,旨在通过分析视觉和非视觉线索特征推断人类情绪,以便计算机系统更好地理解和响应人类需求。情感计算不仅有助于改善人机交互策略以提升用户体验,更能够在危机筛查和风控管理等方面发挥重要作用,对公共安全维护和社会治理具有重要意义。然而,目前的研究大多局限于受控环境或单一环境的情绪识别,而对于人流量大的自然场景,如机场、学校、商铺等,情绪识别的研究较少。复杂环境中的情绪识别面临诸多挑战,包括采集到的数据低分辨率、存在遮挡等问题,同时非视觉线索特征如对话和生理信号难以通过非接触式设备获取。因此,如何在复杂环境中基于视觉线索实现高效、鲁棒的情绪识别是一个重要的研究课题。基于步态的情绪识别方法在心理学理论的基础上具备实际部署的可行性。步态作为小规模非结构化数据,构建的模型参数较少,有助于实现高效、轻量的情绪识别。然而,现有的方法对于高层情感语义的挖掘能力有限。全场景信息在识别人的情绪中起着重要作用。然而,由于情绪感知的模糊性和数据场景的多样性,现有方法训练的模型缺乏普适性。受大规模视觉语言预训练微调范式的启发,通过从大规模情感数据集中学习丰富的语义和视觉表示,可以提高复杂环境下情绪理解的精准性和泛化能力。然而,如何构建大规模复杂环境情感预训练框架以实现鲁棒的情绪识别是一个研究难点。针对这些挑战,本文提出了三种网络结构,实现复杂环境下情绪识别的高效性和鲁棒性。第一,提出了时空自适应图卷积步态情绪识别网络。从人的情绪认知过程启发,本文设计了一种时空自适应图卷积步态情绪识别网络,解决了目前基于骨骼的步态情绪识别方法的两个问题。在空间建模方面,通过推理出关节点动态图,解决了目前基于骨骼的步态情绪识别方法仅依靠静态图进行模型优化导致图结构信息损失的问题。在时间建模方面,通过自适应的学习关节点多尺度时间特征,解决了目前基于骨骼的步态情绪识别方法对关节点复杂运动模式提取不足或聚合僵硬的问题。第二,提出了增强的时空图卷积步态情绪识别网络。随着研究的深入,本文通过拓展时空自适应图卷积步态情绪识别网络,设计了一种增强的时空图卷积步态情绪识别网络,解决了目前基于骨骼的步态情绪识别方法的两个问题。通过设计帧间差异编码模块获取关节点帧间位移,解决了目前基于骨骼的步态情绪识别方法易于受到与情绪无关的受试者朝向的影响,让网络对与情绪相关的关节点的差异变化更敏感。通过设计通道级的动态图推理方法,解决了工作一中网络学习的图结构信息不足的问题,让网络学习到了更多的与情绪相关的隐式连接关系。第三,提出了跨语义引导的多模态预训练框架。本文设计了一种跨语义引导的多模态预训练框架。通过结合图像空间结构信息、文本图像语义信息和背景前景融合信息,从大规模数据集中学习丰富的语义和视觉表示。之后,在多个下游的情感数据集中微调,本文提出的方法解决了现有方法训练的模型缺乏普适性的问题,实现复杂环境下鲁棒的情绪识别。

【Abstract】 Emotion computing,situated at the intersection of computer science and psychology,aims to decipher human emotions through the analysis of visual and non-visual cues.This enables computer systems to better understand and respond to human needs.While it significantly im-proves user experience in human-computer interaction and plays a vital role in crisis screening and risk management,particularly for public safety and social governance,current research pri-marily focuses on emotion recognition in controlled environments.Limited attention has been given to natural scenes with high human traffic,such as airports,schools,and shops.Emo-tion recognition in such complex environments faces challenges including low-resolution data acquisition and occlusion issues.Additionally,extracting non-visual cues like dialogues and physiological signals using non-contact devices poses difficulties.Therefore,developing effi-cient and robust emotion recognition methods based on visual cues in complex environments remains a crucial research pursuit.The gait-based approach to emotion recognition demonstrates practical feasibility for de-ployment,grounded in psychological theories.Gait,as a small-scale unstructured dataset,requires fewer model parameters,thus facilitating efficient and lightweight emotion recogni-tion.However,current methods have limited capabilities in extracting high-level emotional semantics.Holistic scene information is crucial for discerning human emotions.Yet,due to the ambiguity of emotion perception and the diversity of data scenes,models trained using existing methods lack universality.Drawing inspiration from the paradigm of large-scale vi-sual language pre-training fine-tuning,enhancing semantic and visual representations through learning from extensive emotion datasets can improve the precision and generalization ability of emotion understanding in intricate environments.Nevertheless,establishing a large-scale emotion pre-training framework tailored for complex environments to achieve robust emotion recognition remains a challenging research endeavor.To address these challenges,this paper introduces three network architectures aimed at achieving efficiency and robustness in emotion recognition within complex environments.Firstly,we propose a spatio-temporal adaptive graph convolutional network for gait-based emotion recognition.Inspired by human emotion cognition processes,our design addresses two key issues in current skeleton-based gait emotion recognition methods.To address spatial modeling,we infer dynamic graphs of key points,resolving the problem of information loss in the graph structure caused by current methods that rely solely on static graphs for model optimization.For temporal modeling,we adaptively learn multi-scale temporal features of key points,addressing issues such as insufficient extraction or aggregation stiffness in current methods for capturing complex key point motion patterns.Secondly,an enhanced spatio-temporal graph convolutional network for gait-based emo-tion recognition is proposed.With the advancement of research,this paper extends the spatio-temporal adaptive graph convolutional network and designs an enhanced spatio-temporal graph convolutional network to address two key issues in current skeleton-based gait emotion recog-nition methods.By introducing a frame difference encoding module to capture inter-frame displacements of key points,we mitigate the impact of irrelevant subject orientations on emo-tion recognition,making the network more sensitive to differential changes in key points rele-vant to emotions.Furthermore,through the design of a channel-level dynamic graph inference method,we address the issue of insufficient graph structure information learned by the network in the previous work,enabling the network to capture more implicit connectivity relationships relevant to emotions.Thirdly,a cross-modal guided multimodal pre-training framework is proposed.This paper introduces a cross-modal guided multimodal pre-training framework.By integrating spatial structural information from images,semantic information from text,and fusion information of background-foreground,rich semantic and visual representations are learned from large-scale datasets.Subsequently,fine-tuning on multiple downstream emotion datasets,the pro-posed method addresses the issue of lack of universality in models trained by existing methods,achieving robust emotion recognition in complex environments.

  • 【网络出版投稿人】 安徽大学
  • 【网络出版年期】2025年 10期
  • 【分类号】R318;TP18
节点文献中: 

本文链接的文献网络图示:

本文的引文网络