节点文献

基于深度学习的视频行为识别方法研究

Research on Video Action Recognition Method Based on Deep Learning

【作者】 王婷;

【导师】 刘光辉;

【作者基本信息】 西安建筑科技大学 , 智能建筑, 2021, 硕士

【摘要】 近年来,视频行为识别已成为计算机视觉领域的热点之一,在视频检索、智能监控、智慧医疗、人机交互等领域有着重要的应用价值。人体行为的复杂性、视频主体和背景的多样性以及视频的视角变化给人体行为的准确识别带来了巨大的挑战。传统机器学习的方法难以适应复杂场景下的行为识别,基于深度学习的卷积神经网络因能获取视频的深层次特征,在视频行为识别中有着广泛的应用前景。本文在总结分析行为识别领域前人研究工作的基础上,做了以下几方面工作:(1)在视频检索和智能监控等领域,行为识别任务通常需要对整个视频内容进行分析编码,针对3D卷积神经网络存在的低层时空特征丢失问题,提出了一种多层时空信息融合的行为识别方法(Multi-Layer Spatiotemporal Information Fusion Network,MLSIFN)。首先,使用基础骨干网络获取视频的高低层时空特征;之后,通过语义特征嵌入融合方式,将高层时空特征包含较多的语义信息引入低层时空特征,增强低层时空特征的语义表达,再利用特征融合模块聚合全局时空信息,提高网络对时空特征的表征能力;最后,通过实验探究了不同的高低层融合策略。实验结果表明,本文提出的特征融合结构在行为识别任务中具有一定的优越性。(2)针对基于卷积神经网络的行为识别方法存在的长时域建模能力不足、光照及视角变化、复杂背景等因素对视频图像有较大干扰时识别效果不佳等问题,提出一种结合时空特征和光流信息的视频行为识别方法(Video Action Recognition Method Combining Spatiotemporal Features and Optical Flow Information,AMCSOF)。首先,采用均匀稀疏采样策略建立全视频段的时域建模,在降低视频帧冗余度的前提下实现长时序信息的充分保留;其次,光流受运动主体差异性和复杂背景等因素的影响较小,并能反映运动主体的方向、速度等信息,基于此,在多层时空信息融合网络MLSIFN的基础上引入光流数据特征,建立光流信息网络并结合空间注意力模型强化光流特征图中的关键信息,通过不同数据模式之间的优势互补,提高网络在不同场景下的鲁棒性;最后,结合提取的时空特征和光流信息,在网络末端进行决策融合。实验结果表明,该模型不仅能够完成长视频段的时域建模,在复杂多变场景中的行为识别精度也有了一定的提升。(3)相较于视频图像,在不涉及对象或场景上下文信息时,骨架数据对行为的表达更加简洁、高效,本文也尝试采用该方式进行行为识别,侧重于解决人体的摔倒动作识别,提出一种基于人体姿态特征的摔倒识别方法(Fall Recognition Method Based on Human Posture Characteristics,FRMPC)。首先,通过Open Pose人体姿态估计算法从视频图像中提取人体骨架图与关键点的坐标信息;其次,通过对老年人摔倒行为进行分析,获取摔倒动作发生时坐标值具备显著性变化的关键点信息并提取人体姿态特征向量;最后,通过训练行为分类网络,完成摔倒动作识别。实验结果表明,所提方法能够适用于独居老人居家日常活动监测。

【Abstract】 In recent years,the video action recognition,which has become one of the hotspots in computer vision community,has important application value in video retrieval,intelligent surveillance,intelligent medical treatment and human-computer interaction and other fields.The complexity of human action,the diversity of video subject and background and the change of video perspective bring great challenges to the accurate recognition of human behavior.Traditional machine learning method is difficult to adapt to the action recognition in complex scenes.Convolutional neural network based on deep learning has a wide application prospect in video action recognition because it can obtain the deep features of video.Based on the summary and analysis of previous research work in the field of action recognition,this paper has done the following work:(1)In the field of video retrieval and intelligent surveillance,the task of action recognition usually needs to analyze and encode the whole video content.Aiming at the problem of low-level spatial-temporal feature loss in 3D convolutional neural network,a Multi-Layer Spatiotemporal Information Fusion Network(MLSIFN)is proposed.Firstly,the basic backbone network is used to obtain the high-level and low-level spatiotemporal features of the video;then,semantic information,contained in the high-level spatiotemporal features,is introduced into the low-level spatiotemporal features through semantic feature embedding and fusion to enhance the semantic expression of the low-level spatiotemporal features,and then the feature fusion module is used to aggregate the global spatiotemporal information to improve the network’s representation ability of spatiotemporal features;finally,different fusion strategies of high-level and low-level are explored through experiments.The experimental results show that the proposed feature fusion structure is superior in the task of action recognition.(2)The action recognition method based on convolutional neural network has the problem that the ability of long time domain modeling is insufficient and the video image will have a poor recognition effect when the change of illumination and view angle and the complex background cause great interference to the video image.In view of these problems,a Video Action Recognition Method Combining Spatiotemporal Features and Optical Flow Information(AMCSOF)is proposed.Firstly,the time domain modeling of the whole video segment is established by using uniform sparse sampling strategy,and the long time sequence information is fully reserved under the premise of reducing the redundancy of video frames;secondly,the optical flow is less affected by the difference of the moving subject and the complex background,and it can also reflect the direction and speed of the moving subject.Based on above work,based on Multi-Layer Spatiotemporal Information Fusion Network(MLSIFN),the optical flow data feature is introduced,the optical flow information network is established and the key information of the optical flow feature map is strengthened by combining the spatial attention model.The robustness of the network in different scenarios is improved through the complementary advantages between different data modes;finally,the decision fusion is carried out at the end of the network combining the extracted spatiotemporal features and optical flow information.The experimental results show that the model can not only complete the time domain modeling of the growing video segment,but also improve the accuracy of action recognition in complex and changeable scenes.(3)Compared with video images,skeleton data is more concise and efficient in the expression of action when it does not involve the object or scene context information.This paper also attempts to use this method for action recognition,focusing on the recognition of human falling action,and proposes a Fall Recognition Method Based on Human Posture Characteristics(FRMPC).Firstly,the coordinate information of human skeleton and key points is extracted from the video image by Open Pose human posture estimation algorithm;secondly,through the analysis of the elderly fall action,the information of key points with significant changes in coordinate values is obtained when the fall action occurs,and the human posture feature vector is extracted;finally,the fall action recognition is completed by training the action classification network.The experimental results show that the proposed method can be used to monitor the daily activities of the elderly living alone.

  • 【分类号】TP18;TP391.41
  • 【被引频次】1
  • 【下载频次】325
  • 攻读期成果
节点文献中: