节点文献
基于深度学习的暴恐视频识别关键技术研究
Research on the Key Technologies of Violent Terror Video Recognition Based on Deep Learning
【作者】 赵凯;
【导师】 肖波;
【作者基本信息】 北京邮电大学 , 电子与通信工程(专业学位), 2019, 硕士
【摘要】 视频识别是计算机视觉领域常见的任务之一。不同于图片识别,视频数据兼备时域、空域信息。如何在运算开销可控前提下,同时利用好视频数据的时空特征,是视频识别任务的难点和取得性能提升的突破点。针对暴恐视频数据,除了正确识别暴恐场景外,准确识别视频中的人员也具备极大现实意义。然而,不可控的视频采集质量为人员识别带来困难。通常针对人员的识别是通过人脸识别来实现。暴恐视频中人脸的角度、扰动等都会带来识别准确率的下降,需要进行特别的人脸优化处理来提高识别精度。为解决上述问题,论文基于深度学技术,对暴恐视频中的场景识别和人脸优化两方面进行了研究探索,主要包括以下几点工作:(1)针对大量实时视频需要快速判别视频场景的问题,提出了一种针对暴恐视频的场景识别模型,以便于及时发现不同类别的暴恐场景。识别场景分为爆炸、枪击、斗殴、人群奔逃、正常五种,模型以平均间隔帧差为帧选输入逻辑,conv-LSTM为时空特征提取单元,以较小的运算开销,实现了较高的暴恐场景识别准确率;(2)针对暴恐视频中人脸质量欠佳的问题,提出了一种基于多帧人脸融合的人脸质量优化方法,以便于提高人脸识别准确度。暴恐视频中人脸图像采集质量不佳是导致识别精度下降的主要原因,多帧人脸融合是解决单次采集不佳问题的有效方法。论文提出了一种针对多帧人脸图像的融合优化方法,并验证了其相对多帧人脸特征融合的明显优势;(3)针对嫌疑人员证件图像存在网纹扰动问题,设计了一种证件照网纹扰动消除方案,以提高与暴恐视频中人脸比对的精度。证件上的人脸图片常作为人脸识别系统留底图片,但是部分证件出于防伪考虑在人脸上绘制的网纹扰动为识别带来困难。论文提出了基于生成对抗网络的扰动消除模型,并以特征监督保障了模型对网纹消除和人脸身份信息保留的平衡,对留底人脸与暴恐视频中人脸的比对性能具备较高保障意义。本文围绕暴恐视频中视频场景识别、人脸识别优化两方面展开研究。视频场景识别中,通过帧选逻辑、conv-LSTM设计了具有运算开销保障、兼顾时空特征的识别模型。人脸识别优化中,通过相同数据集上人脸识别模型识别性能的提升,说明了本论文所提方法对暴恐视频中的现场人脸优化、留底证件照扰动消除的有效性。
【Abstract】 Video recognition is one of the common task scenarios in the field of computer vision.Unlike picture recognition,video data has both time domain and spatial domain information.How to make good use of the spatiotemporal features of video data while controlling the computational overhead is a difficult point for video recognition tasks and a breakthrough point for achieving performance improvement.Against violent terror video data,in addition to correctly identifying the violent scene,it is of great practical significance to accurately identify the personnel identity information in the video.However,uncontrollable video capture quality poses difficulties for personnel identity recognition.Usually the identification of people is achieved through face recognition.The angle of the face,the disturbance,etc.in the violent terror video will bring about a decrease in the recognition accuracy,and special face optimization processing is required to improve recognition accuracy.In order to solve the above problems,based on deep learning technology,this paper has carried out research and exploration on video scene recognition and face recognition optimization in the field of violent terror video recognition,including the following works:(1)Against the need to quickly identify video scenes for a large number of real-time videos,a scene recognition model for violent terror video was proposed to timely discover different categories of terror scenes.Recognition scenes are divided to bomb,shooting,fighting,crowd-running and normal scene.The model uses the average interval frame difference as the frame selection logic,and the conv-LSTM is the spatiotemporal feature extraction unit,which achieves a higher accuracy of violent terror scene recognition with less computational overhead.(2)Against the problem of poor face quality in the video of violent terror,a face correction method based on multi-face fusion is proposed to improve face recognition accuracy.The poor quality of face image acquisition in violent terror video is the main reason for the decline of recognition accuracy.Multi-face fusion is an effective method to solve the problem of single acquisition.The paper proposes a fusion correction method for multi-frame face images,and verifies its obvious advantages relative to multi-frame facial feature fusion.(3)Against the cobwebbing disturbance problem of suspect’s credential photos,a credential photo cobwebbing disturbance elimination model was established to improve the comparison accuracy with faces in violent terror videos.The face photo on the credentials is often used as the criterion picture in face recognition system,but some of the identification photos are difficult to identify directly because of the cobwebbing disturbance for anti-counterfeiting.The paper proposes a disturbance elimination model based on the generative adversarial network,and guarantees the balance of the disturbance elimination and face identity retention by feature supervision,and the performance of comparison between label face and the face in the violent terror video has a higher guarantee.This paper focuses on two aspects of video scene recognition and face recognition optimization in the violent terror video.In video scene recognition,a recognition model with computational overhead guarantee and spatiotemporal feature combining is designed based on frame selection logic and conv-LSTM unit.In the face recognition optimization,the recognition performance of the face recognition model on the same dataset is improved with the participation of optimization model,and therefore the effectiveness of the method proposed in this paper on the optimization of the faces in the violent terror video and the elimination of cobwebbing disturbance on credential photos is illustrated.
【Key words】 deep learning; computer vision; video recognition; face recognition;
- 【网络出版投稿人】 北京邮电大学 【网络出版年期】2019年 09期
- 【分类号】TP391.41;TP181
- 【被引频次】4
- 【下载频次】353