节点文献

面向高效视频目标跟踪的模型学习、压缩与集成

Model Learning,Compression,and Integration for Robust Visual Object Tracking

【作者】 王宁;

【导师】 周文罡; 李厚强;

【作者基本信息】 中国科学技术大学 , 信息与通信工程, 2021, 博士

【摘要】 视频目标跟踪是计算机视觉中的一个基本任务。仅给定初始帧的目标状态,视频目标跟踪要求在后续帧中对该物体进行持续的定位。近年来,基于深度学习的视觉跟踪技术取得了重大进展。然而,该技术也带来了数据标注代价大、模型参数多、跟踪效率低等问题。为了挖掘深度视觉跟踪的潜力,本文在模型学习、压缩和集成等三个方面开展研究。本文的主要贡献和创新点包含以下三个方面:在模型学习方面,本文研究如何减轻训练代价以及挖掘视频的时序信息,提出了无监督跟踪模型学习方法和基于Transformer的视频目标跟踪方法。首先,针对深度跟踪模型学习过程的标注成本高和训练代价大等问题,本文提出了基于前向后向跟踪轨迹一致性的无监督训练框架。无监督学习的动机在于一个鲁棒的跟踪器可以进行双向跟踪。在训练过程中,所提出的算法仅使用无标注视频,通过衡量跟踪器在视频中前向跟踪轨迹和后向跟踪轨迹的一致性来无监督地训练跟踪器。本文学习的无监督跟踪器可以和经典的全监督跟踪器相媲美,并达到了实时的跟踪速度。另外,为了挖掘视频中丰富的时序信息,本文将Transformer结构引入到视觉跟踪领域。Transformer中的编码器、解码器结构将独立的视频帧紧密地桥接起来,便于传递诸如目标时序特征和空间注意力掩膜等丰富的时序信号。结合本文所提出的Transformer结构后,现有的跟踪器获得了稳定的性能提升并达到了领先精度。在模型压缩方面,针对深度跟踪器的模型参数量大、计算复杂度高和跟踪效率低等问题,本文提出了联合模型压缩和知识迁移的深度跟踪器加速框架。本工作使用在图像分类任务中预训练的CNN模型作为教师网络,并将该网络蒸馏成一个轻量级的学生网络用来加速相关滤波算法的特征提取过程。在蒸馏过程中,本文提出保真损失来保证教师网络和学生网络相近的特征表达能力,同时提出相关跟踪损失将学生网络的约束目标从物体识别迁移到视觉跟踪。大量的实验表明,所提出的轻量级学生网路显著地加速了目前领先的深度相关滤波器,使得它们在单块CPU上可以达到实时速度,并几乎保持了原有的跟踪精度。在模型集成方面,本文研究如何集成多个深度跟踪器以实现模型间的优势互补,并提出了基于多线索分析和基于策略选择的两种算法。首先,在多线索分析框架中,本工作构造多个子跟踪模型并行地跟踪目标。通过评估各模型的鲁棒性,每一帧中合适的模型被用于处理当前帧。进一步地,多模型之间的分歧揭示了当前跟踪结果的可靠性,可用于指导子跟踪器的自适应更新以避免模型污染。通过多线索分析策略,跟踪性能得到显著提升。基于策略的切换式集成框架研究如何在不影响效率的情况下发挥集成算法的性能优势。本工作包含智能体网络和多个优势互补的子跟踪模型。通过将模型逐帧切换建模成马尔可夫决策过程,本工作通过强化学习来训练智能体网络。在每一帧中,智能体动态地选择一个合适的子跟踪器用于目标跟踪,而不必执行其他模型,极大地保证了跟踪效率。大量实验证明了该框架的有效性。

【Abstract】 Visual object tracking is a fundamental task in computer vision.Given the initial target state,a visual tracker requires to locate the target object in successive frames.In recent years,deep learning based visual tracking has made significant progress.Nev-ertheless,deep learning techniques also bring the issues such as expensive annotation cost,high model complexity,and limited tracking efficiency.To release the potential of deep visual tracking,this thesis focuses on three aspects including model learning,model compression,and model integration.The main contributions of this thesis are three-fold:In model learning,this thesis investigates how to alleviate the training cost of deep trackers and how to leverage the temporal information resides in the tracking videos.First,to alleviate the expensive data labeling and high training cost in deep visual track-ing,this thesis presents an unsupervised tracking framework based on the forward and backward tracking trajectory analysis.The motivation of unsupervised learning is that a robust tracker should be effective in bidirectional tracking.In the training process,the proposed algorithm measures the consistency between forward and backward trajecto-ries to learn a robust tracker from scratch merely using unlabeled videos.The proposed unsupervised tracker exhibits the baseline accuracy of classic fully supervised trackers while achieving a real-time speed.Next,to exploit the rich temporal information in the video flow,this thesis introduces the transformer architecture to the visual track-ing community.The transformer encoder-decoder structure tightly bridges the isolated video frames to propagate rich temporal cues(e.g.,target features and attention masks)across frames.By virtue of the proposed transformer,existing trackers gain substantial performance improvements and achieve state-of-the-art accuracy.In model compression,to tackle the high computational complexity,huge model parameters,and unsatisfactory tracking efficiency of deep tracking,this thesis proposes to jointly compress and transfer the heavyweight tracking models.This work formulates a CNN model pretrained from the image classification task as a teacher network,and distills this teacher network into a lightweight student network as the feature extractor to speed up correlation filter trackers.In the distillation process,this thesis proposes a fidelity loss to enable the student network to maintain the representation capability of the teacher network,and designs a tracking loss to adapt the objective of the student network from object recognition to visual tracking.Extensive experiments on standard datasets demonstrate that the lightweight student network accelerates the speed of state-of-the-art deep trackers to real-time on a single-core CPU while maintaining almost the same tracking accuracy.In model integration,this thesis investigates how to assemble multiple deep track-ers to achieve the model complementation.This thesis proposes two ensemble algo-rithms including a multi-cue analysis strategy and a policy-based switch framework.Multi-cue analysis framework constructs multiple experts through correlation filter and each of them tracks the target independently.With the proposed robustness evaluation strategy,the suitable expert is selected for tracking in each frame.Furthermore,the di-vergence of multiple experts reveals the reliability of the current tracking,which is quan-tified to update the experts adaptively to keep them from corruption.With the proposed multi-cue analysis,tracking performance is significantly improved.The policy-based switch framework aims to retain the performance advantage of the ensemble frame-work without sacrificing tracking efficiency.This algorithm consists of multiple weak but complementary experts and an agent network.By formulating this expert switch in consecutive frames as a decision-making problem,this approach learns an agent via re-inforcement learning to directly decide which expert to handle the current frame without running others,which greatly ensures the overall tracking efficiency.Extensive exper-iments verify the effectiveness of the proposed method.

节点文献中: