节点文献

基于深度学习的行人重识别算法研究

Research on Person Re-Identification Algorithm Based on Deep Learning

【作者】 陈鹏;

【导师】 陈东岳;

【作者基本信息】 东北大学 , 控制工程(专业学位), 2020, 硕士

【摘要】 近年来,随着视频监控摄像头的普及,借助人工智能技术助力视频监控领域的发展成为一个热点问题。行人重识别是智能安防监控领域中非常重要的基础应用研究,在犯罪追踪、安防反恐等方面有广泛的应用需求,而将深度学习与计算机视觉技术相结合是解决行人重识别问题的最有效途径。本文面向行人重识别技术应用研究需求,重点针对摄像头视角和行人姿态变化对行人重识别的性能干扰过大的问题和不同场景间存在的模型域适应能力差、人工标记成本过高的问题开展研究工作,具体内容如下:(1)首先针对视角和行人姿态变换的问题,本文通过使用空间变换网络自适应的调整视角和姿态变换使得行人对齐,再结合密集卷积骨架复用低级和高级特征获得可分辨的表观特征,最后使用三元组损失学习特征空间的度量,使得类内距离拉近,类间距离拉大,避免了类别数量对模型的干扰。结合空间变换网络、密集卷积骨架和三元组损失,实现了一个端到端的鲁棒高效的行人重识别系统。(2)虽然我们的行人重识别系统在一定程度上缓解了视角和姿态变化的问题,但对于运动场景中剧烈的行人姿态变化还是存在着不足;从特征的角度出发,一个分辨性强的表观特征可以克服行人姿态变化带来的噪声。本文提出了一种深度互信息判别网络进行行人特征增强,以获得有分辨性的行人特征。通过对两种基准模型的实验证明了有监督特征增强算法的有效性。(3)对于行人重识别域适应的问题,本文通过新场景中大量的无标注数据进行无监督学习来适应场景。在有监督训练模型作为初始化参数的无监督网络中,最大化图片与特征间的互信息来增强表观特征,使得训练集模型逐渐适应实际应用的新场景数据集,即缓解了无监督域适应的问题。实验结果表明本文提出行人对齐算法和行人特征增强算法整体可行,能够有效解决行人重识别任务中的姿态变化和无监督域适应问题,对于行人重识别方法的深入研究与工业应用具有一定的理论价值和推动作用。

【Abstract】 In recent years,with the popularity of video surveillance system,utilizing artificial intelligence technology to help the development of video surveillance has become a hot issue.With general application requirements in crime tracking,security against terrorism and so on,Person re-identification(ReID)is a crucial fundamental application research in the security monitoring field.Combining deep learning with computer vision technology is the most effective way to solve the problem of ReID.This paper is aimed at the application research requirements of ReID,focusing on the problem that the camera views and pedestrian pose variances have too much interference on ReID and the problem of the limited ability of model domain adaptation between different scenes and the high cost of manual labeling.The details are as follows:For the problem of camera views and pose variances,firstly our paper utilizes the spatial transformation network(STN)to adaptively adjust the variances to align the pedestrians,and then combine the low-level and high-level features of DenseNet architecture to obtain the distinguishable apparent features.Finally,the triplet loss could learn the metric distance in feature space,trying to pull closer the intra-class distance and push away the inter-class distance.The triplet loss avoids the interference of the number of categories on the model.Combining spatial transformation network,DenseNet architecture and triple loss,an end-to-end robust and efficient ReID system is implemented.Although our ReID system alleviates the problems of camera views and pedestrian pose variances to a certain extent,there are still shortcomings for drastic pedestrian posture changes in sports scenes.From the perspective of features,a highly distinguishable feature can overcome the noise caused by the change of pedestrian posture.This paper proposes a deep mutual information discriminative network for pedestrian feature enhancement,therefore achieving distinguishable pedestrian features.Experiments on two benchmark models prove the effectiveness of the supervised feature enhancement algorithm.In this paper,a large number of unlabeled pedestrian images are used to mitigate domain adaptation problems in ReID issue.The supervised training model parameters are adopted as the initial parameters in unsupervised learning.The discriminator network maximizes the mutual information between the input images and features to enhance the feature representation,therefore gradually adapting to the new scene dataset,which alleviates the problem of unsupervised domain adaptation.The experimental results show that the proposed pedestrian alignment algorithm and unsupervised feature enhancement algorithm are feasible,which can effectively solve the pose variance and domain adaptation problem in ReID task.It has certain theoretical value for the in-depth study and industrial application of ReID method.

  • 【网络出版投稿人】 东北大学
  • 【网络出版年期】2024年 01期
  • 【分类号】TP391.41;TP18
节点文献中: 

本文链接的文献网络图示:

本文的引文网络