节点文献
基于跨模态学习的视频情绪识别方法研究
Research on Video Emotion Recognition Via Multimodal Learning
【作者】 王杰;
【导师】 薛均晓;
【作者基本信息】 郑州大学 , 软件工程, 2024, 硕士
【摘要】 在信息化社会中,视频作为信息的重要载体,蕴含了丰富的情感信息。然而,由于视频数据的复杂性和多样性,如何从视频中准确识别情绪成为了一项具有挑战性的任务。传统的视频情绪识别方法往往依赖于单一模态的信息,仅基于视觉特征或音频特征进行情绪识别,往往难以捕捉到视频中情绪信息的全貌。跨模态学习通过整合不同模态的信息,能够更全面地理解和分析视频中的情绪表达。然而,跨模态学习也面临着诸多挑战。首先,不同模态的数据之间存在异构性,如何有效地融合这些异构数据是一个关键问题。其次,视频中的情绪表达具有复杂性和动态性,如何准确地捕捉和解析这些情感信号也是一个挑战。此外,跨模态学习还需要处理大量的数据,如何在保证性能的同时提高算法的效率也是一个需要解决的问题。本研究旨在探索基于跨模态学习的视频情绪识别方法,通过整合视觉和音频等多模态信息,提高情绪识别的准确性和鲁棒性。本文设计了一系列跨模态学习算法,解决了情绪识别、跨模态融合和算法效率提升等关键问题。本文的主要工作包括:(1)从情绪特征的泛用性出发,构建了一个大规模的多模态情绪识别数据集,为情绪识别和舆情分析任务设定了性能基准,并验证了数据集和迁移学习方法的有效性。进一步,设计了一个跨模态情绪识别模型,通过整合图像、光流、音频和文本模态的决策提升了模型性能,相比最优单模态性能提升了约5%的识别性能。最后,通过舆情发展曲线的拟合度,验证了舆情-情绪任务的有效性,并确立了任务的评价标准。(2)以决策融合方式为核心,提出了基于置信度融合的可信情绪识别模型,并探讨了在决策级融合中可信融合对模型识别性能的影响。该融合方法通过置信模块估计各模态决策的置信度,可信融合模块基于置信度实现决策融合。同时,基于不确定性估计实现了对模型识别结果可信性的度量,并确定了可信阈值的设定方法和可信性能的评价标准,该模型以60.14%和82.40%的可信准确率取得了最优可信性能,并且模型识别性能名列前茅。(3)着眼于模型轻量化,提出了基于深度可分离卷积的轻量级情绪识别模型,并在保持模型轻量化的同时,探讨了模型性能表现。该模型通过采用精心设计的分离三维卷积核使得模型参数量下降为可信情绪识别模型的3.48%,并利用知识蒸馏方法学习教师模型在情绪特征提取方面的优秀能力。轻量化模型计算量下降为教师模型的3.27%,凸显了其在实际部署和应用中的优越性。
【Abstract】 Video as a significant carrier of information contains rich emotional content in the information society.However,due to the complexity and diversity of video data,accurately recognizing emotions from videos has become a challenging task.Traditional video emotion recognition methods often rely on information from a single modality,such as visual features or audio features alone,which may fail to capture the full spectrum of emotional information in videos.Multimodal learning,by integrating information from different modalities,allows for a more comprehensive understanding and analysis of emotional expressions in videos.However,multimodal learning also faces several challenges.Firstly,the heterogeneity among different modalities of data poses a key question on how to effectively fuse these diverse data sources.Secondly,the emotional expressions in videos are complex and dynamic,making it challenging to accurately capture and interpret these emotional signals.Additionally,multimodal learning must handle large volumes of data,necessitating solutions to improve algorithmic efficiency while maintaining performance.This research aims to explore video emotion recognition methods based on multimodal learning,enhancing the accuracy and robustness of emotion recognition by integrating multimodal information such as visual and audio data.We have designed a series of multimodal learning algorithms that address key issues like emotion recognition,fusion of multimodal data and enhancement of algorithmic efficiency.The main contributions of this thesis include:(1)Beginning with the generalizability of emotional features and constructs a large-scale multimodal emotion recognition dataset,establishing performance benchmarks for emotion recognition and public opinion analysis tasks,and validating the effectiveness of dataset and transfer learning methods.Furthermore,the thesis develops a multimodal emotion recognition model that enhances performance by integrating decisions from image,optical flow,audio,and text modalities,which improves the recognition performance by about 5%compared to the optimal unimodal performance.Finally,the validity of the opinion-emotion task is confirmed through the fit of public opinion development curves,and evaluation criteria for the task are established.(2)Focusing on decision fusion methods and proposes a trusted emotion recognition model based on confidence fusion,investigating the impact of confidence fusion at the decision level on model recognition performance.The fusion method estimates the confidence of each modal decision through the confidence module,and the trusted fusion module realizes the decision fusion based on the confidence.Additionally,this thesis employs uncertainty estimation to measure the credibility of the model’s recognition results and establishes methods for setting confidence thresholds and evaluating trusted performance,and the model achieves the optimal credibility performance with trusted accuracies of 60.14%and 82.40%,and the model recognition performance is ranked among the best.(3)Focusing on model lightweighting and proposes a lightweight emotion recognition model based on depth wise separable convolutions,simultaneously exploring its performance while maintaining lightweight properties.The model reduces the number of model parameters to 3.48%of trusted emotion recognition model by employing a well-designed separated 3D convolution kernels and leverages knowledge distillation to learn the outstanding emotion feature extraction capabilities of the teacher model.The lightweight model computation is reduced to 3.27%of the teacher model,highlighting its superiority in real-world deployment and application.
【Key words】 Video emotion recognition; Multimodal learning; Transfer learning; Trusted model; Lightweight model;
- 【网络出版投稿人】 郑州大学 【网络出版年期】2026年 06期
- 【分类号】TP391;TN912.3;TP18