节点文献

复杂场景下的无人机视觉感知行人多目标跟踪技术研究

Research on Multi-target Tracking Technology of Pedestrians Based on UAV Visual Perception in Complex Scenes

【作者】 王旭;

【导师】 王磊;

【作者基本信息】 电子科技大学 , 电子信息(专业学位), 2025, 硕士

【摘要】 近年来,无人机因体积小、灵活和效率高等优势被广泛应用在军事、民用和商业等领域,无人机视觉多目标跟踪特别是行人目标的跟踪作为无人机获取外部环境信息的重要手段,对于无人机的应用具有重要价值。但是,要完成无人机视角下的精准的行人多目标跟踪存在相似外观、复杂背景、复杂运动和遮挡的问题。为此本文标注了无人机行人多目标跟踪数据集和行人重识别数据集并研究了无人机场景的多目标跟踪技术,最后设计实现了行人多目标跟踪系统。具体工作如下:(1)针对无人机视角下目标存在复杂背景和相似外观的问题,本文基于Res Net50设计了一种外观特征提取算法,引入软注意力模块帮助模型过滤复杂背景的干扰,引入局部特征提取分支通过提取细粒度特征以应对相似外观的干扰。针对现有无人机数据集缺乏的问题,本文拍摄了大量的无人机视频并进行手动标注。最后,实验证明,本文改进的外观特征提取算法在手工标注的行人重识别数据集上的平均精度均值和Rank-1分别达到了88.3%和88.9%,相比于基础的Res Net50模型分别提升了4.7%和4.5%。同时在cuhk03公开数据集上也取得了较好的重识别效果,验证了改进的外观特征提取算法的泛化性。(2)针对目标移动速度过快或尺度变化较大问题,本文改进了交并比设计了搜索窗口方法,通过扩大关联空间以减少错过匹配。同时基于长短期记忆网络设计了目标位置预测网络进行复杂运动的建模,设计了目标位置预测纠正模块用于处理多目标跟中相机抖动导致输入数据出现偏差的问题,设计了带有随机噪声的精准位置预测模块用于模拟复杂运动导致的多目标跟踪出现误匹配的场景。此外本文设计了一种乘积融合方法,以应对因遮挡导致的外观线索和运动线索不可靠问题。最后,本文设计了一种基于置信度分离数据关联框架用以处理低分轨迹和检测。实验证明,本文改进的多目标跟踪算法在手工标注的行人多目标跟踪数据集上达到了71.8%的高阶跟踪精度、82.1%的多目标跟踪精度和85.7的身份F1分数。同时在Sports MOT公开数据集上也取得了较好的多目标跟踪效果,验证了改进的多目标跟踪算法的泛化性。(3)本文设计了一套行人多目标跟踪系统。系统由输入模块、后台处理模块和结果展示模块组成,可以同时展示多目标跟踪和单目标跟踪功能。能直观展示不同多目标跟踪算法的跟踪效果和跟踪性能,验证算法的实用性。

【Abstract】 In recent years,drones have been widely used in military,civilian,and commercial fields due to their advantages such as small size,flexibility,and high efficiency.Visual multi-object tracking(MOT),especially pedestrian tracking,serves as a crucial means for drones to acquire external environmental information,holding significant value for drone applications.However,to achieve precise multi-target pedestrian tracking from the perspective of unmanned aerial vehicles(UAVs),there are problems such as similar appearance,complex background,complex movement and occlusion.To address this,thesis annotates a drone-based pedestrian multi-object tracking dataset and pedestrian re-identification dataset,investigates MOT technologies for drone scenarios,and ultimately designs and implements a pedestrian multi-object tracking system.The specific contributions are as follows:(1)To address the challenges of complex backgrounds and similar appearances of targets in drone-view perspectives,this thesis designs an appearance feature extraction algorithm based on Res Net50.A soft attention module is introduced to help the model filter out interference from complex backgrounds,while a local feature extraction branch is incorporated to capture fine-grained features for mitigating the impact of similar appearances.Regarding the scarcity of existing drone datasets,this thesis captures a large number of drone videos and manually annotates them.Experimental results demonstrate that the improved appearance feature extraction algorithm achieves a mean Average Precision(m AP)of 88.3%and a Rank-1 accuracy of 88.9%on the manually annotated person re-identification dataset,representing improvements of 4.7%and 4.5%,respectively,compared to the baseline Res Net50 model.Additionally,the algorithm achieves competitive re-identification performance on the public cuhk03 dataset,validating the generalization capability of the enhanced appearance feature extraction algorithm.(2)To address issues such as fast-moving targets or significant scale variations,this thesis improves the Intersection over Union(Io U)by design a search window method,which expands the association space to reduce missed matches.Additionally,a target position prediction network based on Long Short-Term Memory(LSTM)is designed to model complex motion patterns.A target position prediction correction module is designed to handle deviations in input data caused by camera jitter in multi-object tracking,and a precise position prediction module with random noise is designed to simulate scenarios where complex motion leads to mismatches in multi-object tracking.Furthermore,this thesis proposes a product fusion method to address unreliable appearance and motion cues due to occlusions.Finally,a confidence-based data association framework is designed to handle low-score trajectories and detections.Experimental results show that the improved multi-object tracking algorithm achieves71.8%Higher Order Tracking Accuracy(HOTA),82.1%multiple Object Tracking Accuracy(MOTA),and 85.7%ID F1 Score(IDF1)on the manually annotated multi-object pedestrian tracking dataset.It also demonstrates strong performance on the public Sports MOT dataset,validating the generalization capability of the improved multi-object tracking algorithm.(3)This thesis designs a comprehensive pedestrian multi-object tracking system.The system consists of an input module,a backend processing module,and a result display module,capable of simultaneously demonstrating multi-object tracking and single-object tracking functionalities.It provides an intuitive visualization of the tracking performance and effectiveness of different multi-object tracking algorithms,verifying their practical utility.

  • 【分类号】V279;TP391.41
节点文献中: