节点文献
面向视频侦查的行人检索关键技术研究
Research on Key Technologies of Pedestrian Retrieval for Video Investigation
【作者】 王晓;
【导师】 陈军;
【作者基本信息】 武汉大学 , 通信与信息系统, 2020, 博士
【摘要】 近年来,大规模视频监控系统已经在我国投入使用,公安部门案件的侦破方式也随之发生了巨大的变革。在实际视频侦查应用中,一线公安人员仍需要调看案发前后及周边监控视频,扩大搜索范围,以便从多个摄像头的视频数据中查找特定嫌疑人的运动轨迹,捕获有用的线索。在现有的案件侦破中,主要通过人工浏览的方式来发现嫌疑人,需要耗费大量的人力物力,效率低下。视频侦查的实际需求推进了行人检索的发展。行人检索,即给定一个监控摄像头视频中的行人对象,在其他摄像头任何时段任何位置的全画面监控视频中查寻该行人的技术,能够帮助一线公安人员快速地锁定嫌疑对象并进行追踪,进而在黄金破案时间内提供有力支持,提高公安部门的案件侦破率,具有重要的应用意义。实际视频侦查主要面向于监控场景,受监控场景复杂性的影响,行人检索使用性能并不理想,很难满足公安部门的需求。行人检索在实际监控场景中面临着以下四个瓶颈:(1)行人活动范围广。行人在日常活动中不受外来因素的影响,可以随意的运动,活动范围较广,会出现在监控图像中的任意位置。给行人候选区域的提取造成了很大的困难,传统的基于全画面滑动窗口获取行人候选区域的方法带来大量的背景冗余信息,效率低下。(2)行人尺度变化大。监控摄像头拍摄的视角较广,往往会涉及行人从距离摄像头较远的位置进入监控画面,再从距离摄像头较近的位置离开监控画面,或者相反。传统基于单一尺度阈值判断的方法在处理这种行人尺度变化大问题上性能显著下降。(3)行人遮挡多样化。在实际视频侦查中,行人会被其他人或者其他物体遮挡,进而处于复杂多样的遮挡形态下。传统的基于固定遮挡模型的方法很难准确的匹配实际多变的遮挡形态,严重影响了行人表达能力。(4)行人视觉差异大。行人在长期活动中,妆容和衣着都有明显的差异,视觉差异较大,传统的行人相似性采用视觉的距离度量来计算,视觉的差异严重影响了距离度量的性能。(1)内部结构协同关联的行人候选区域提取方法已有的研究采用的是面向全画面滑动窗口扫描的策略获取行人候选区域,该方法不仅会带来大量无效背景区域,在特征提取阶段还会带来大量的时间消耗。本文通过图像内部区域在颜色、纹理和轮廓上的差异构建显著信息,结合滑动窗口扫描策略组成便于区分行人和背景的候选区域。将显著信息级联分类器中,快速过滤背景区域,提高了检测速度。实验结果表明,提出的方法相对于基于滑动窗口扫描方法在Caltech数据集上的检测速度提升了2.5倍。(2)尺度空间渐变曲线映射的行人判断方法已有研究采用的是单一尺度阈值的判断方法,在处理大尺度行人时表现出较好的性能,小尺度行人的判别信息较少,其分类分值往往会低于阈值被划分为背景区域,导致严重的漏检。本文通过探索行人分类在尺度空间的映射关系,挖掘行人随尺度渐变的内在规律,建立鲁棒的行人模型,进而准确的判断小尺度的行人图像,实现从单一阈值点的判断到多尺度空面判断的转变。Caltech数据集上的实验表明,提出的方法相对于阈值判断方法在FPPI值上优化了6.4%。(3)自适应组合优化的行人表达方法已有的处理行人遮挡工作关注于常见的几个固定遮挡模型,这些固定形态模型难以精准匹配实际监控场景中变化多样的遮挡形态。本文通过探索卷积神经网络通道和行人内部区域之间的对应关系,学习通道和遮挡形态的激活方式,研究通道指引的遮挡形态表达方法,实现从固定组合到自适应组合优化的转变。CUHK-SYSU数据集上的实验结果表明,提出的方法相对于固定遮挡形态方法在m AP值上提升了6%。(4)多模态动态调整的距离度量方法已有的行人检索采用的是基于行人视觉属性的距离度量函数,能够很好的度量短期活动中视觉差异较小的行人,但是在实际应用中,行人大多处于长期活动中,妆容和衣着有很大的不同,外貌差异较大,基于视觉属性的距离度量函数难以区分。本文在基于外貌的行人检索中加入人脸识别和声纹识别来协同分析行人在长期活动中的相似性度量,为克服人脸和声音在监控中的不稳定性,引入自关注机制,自动的调整三种模态间的动态关系,实现从单模态的相似性度量到多模态相似性度量的转变。CSM-V数据集上的实验结果表明,提出的方法相比于基于视觉的距离度量方法在m AP值提升了11.59%。综上所述,本论文通过挖掘行人检索在复杂监控场景中的受限原因,完成了候选区域有序提取、尺度空间渐变曲线映射判断、自适应组合优化表达和多模态动态调整的距离度量等四方面理论研究,为实际公安部门视频侦查提供了新的方法。
【Abstract】 Large-scale video surveillance systems have been used in China.It has taken place for the police to investigate criminal cases.However,in the process of the practical video investigation,investigators should watch vast amounts of surveillance video around the area of the crime place,before and after the crime time,and search for the same pedestrian,track the pedestrian among multiple camera videos,so that find out and trace the suspect.In the past,video investigation work is also done by manual,which costs a large number of human resources and time,resulting in low efficiency.To the end,the efficiency requirement of the video investigation pushes the development of pedestrian retrieval.Pedestrian retrieval is the technology to judge whether a pedestrian in one camera has ever appeared in the other camera.This technology helps investigators search and track suspects,improve the cases-cracking rate,and must be of great significance.However,the conditions become complicated,and pedestrian retrieval’s effectiveness would be dropped significantly,which cannot fulfill the needs of video investigation applications.There still exist several technological difficulties to be overcome under these four complicated conditions.(1)With the influence of subjective movement in the surveillance,persons may appear at any scale in any location.This arbitrariness brings great difficulties to obtain the positions of persons.It brings a large amount of redundant information due to the specific strategy of the traditional method of getting a person position at any scale based on the sliding window,resulting in low efficiency.(2)The wide-angle view of the surveillance camera often involves pedestrians entering the surveillance image from a long distance away from the camera.After gradually entering the picture,they leave the surveillance image from a distance closer to the camera,or vice versa.Pedestrians will span multiple resolutions during this period,and this type of pedestrian resolution gradient is more common.The traditional method based on single threshold discrimination significantly reduces this kind of pedestrian resolution gradient problem.(3)To elude from the camera,suspects often hide behind other things or persons,leading to a series of occlusion patterns.These suspects are notoriously hard to match due to the substantially various appearance in the intricate occlusion patterns.Existing methods for dealing with occlusion depend on learning several frequent patterns separately,which is not feasible for all occlusion patterns in practice.It brings not only high consumption but also less coverage of patterns in real-world scenarios.(4)In pedestrians’ long-term activities,there is a big difference between makeup and clothing.The visual difference is significant.The traditional pedestrian similarity is calculated by visual distance measurement.The visible difference seriously affects the performance of distance measurement.This paper researches pedestrian retrieval,mainly focusing on selecting candidate regions,pedestrian discrimination,pedestrian representation,and distance metric.The main researches are listed as follows:(1)Candidate region extraction based on internal structure collaborativeExisting studies have adopted a sliding window scanning strategy for a full-frame to obtain pedestrian candidate regions.This method will bring a lot of invalid background regions and bring a lot of time in the feature extraction stage.In this paper,the salient information is constructed by the difference in color,texture,and outline between the image’s internal areas.The salient information is merged into the cascade classifier to quickly filter the background area and quickly improve the detection speed.Experimental results show that the proposed method improves the detection speed on the Caltech data set by 2.5 times compared with the sliding window scanning method.(2)Scale-space curve mapping for pedestrian discrimination.The entry and exit monitoring screens of pedestrians will cross different scales.The existing research uses a single threshold discrimination method,which shows better performance when dealing with large-scale pedestrians.Small-scale pedestrians have less discriminatory information and tend to be below the threshold.They are determined as a background area,leading to dangerous missed detections.This paper explores the mapping relationship of pedestrian discrimination in scale space,excavates the internal laws of pedestrian discrimination that vary with scale,and then accurately judges small-scale pedestrian images.It transfers the single threshold discrimination to scale space surface discrimination.Experiments on the Caltech dataset show that the proposed method is 6.4% better than the threshold discrimination method in the FPPI(False Positive per Image)value.(3)Adaptive combination optimization for pedestrian representation.Existing methods depend on learning several frequent patterns separately for pedestrian occlusions.It is difficult to accurately match the varied occlusion patterns in real-world scenarios.This paper explores the correspondence between each channel of the convolutional neural network and the pedestrian’s internal region,learns the activation patterns of the channel,and constructs the occlusion pattern construction model guided by the channel.It transfers a fixed combination to the optimization of adaptive combinations.The experimental results on the CUHK-SYSU dataset show that the proposed method improves the m AP(mean average precision)value by 6% compared to the fixed occlusion morphology method.(4)Multimodality-based dynamically adjusted distance metric.Existing pedestrian retrieval uses a distance measurement function based on the visual attributes of pedestrians.It can measure pedestrians with small visible differences in short-term activities.However,in practical applications,most pedestrians are in long-term activities.These distance measurement functions based on visual attributes are difficult to distinguish between these persons.This paper adds face recognition and voiceprint recognition to pedestrian retrieval based on appearance to analyze pedestrians’ similarity measure in long-term activities collaboratively.It can adjust the dynamic relationship between the three modalities and transfer single-modality similarity measurement to multi-modality similarity measurement.The experimental results on the CSM-V dataset show that the proposed method improves the m AP(mean average precision)value by 11.59% compared to the visual distance measurement method.In summary,this dissertation overcomes the bottlenecks of video investigation.It completes research on the ordered combination of candidate regions,discrimination of scale-space mapping,adaptive combination optimization representation,and dynamically adjusted distance metric.It provides a new way to practical video investigation.
【Key words】 Video Investigation; Pedestrian Retrieval; Pedestrian Discrimination; Pedestrian Representation; Distance Measure;