节点文献
面向图像级别标注的弱监督目标检测方法与应用
Weakly Supervised Object Detection Method and Application under Image Level Annotation
【作者】 王振东;
【导师】 宫辰;
【作者基本信息】 南京理工大学 , 计算机科学与技术(专业学位), 2021, 硕士
【摘要】 目标检测不仅是计算机视觉领域最基本的任务之一,也是很多高级应用不可或缺的技术,如自动驾驶、人脸识别、图像检索等。得益于实例级别的精准数据标注,传统的全监督目标检测算法能够达到极高的精度,然而在实际情况下数据标注不仅获取成本昂贵,而且极难确保精准无误,因此研究人员开始将目光转移到面向图像级别标注的弱监督目标检测算法研究上。近年来,弱监督目标检测领域已形成了基于多示例学习的基础框架,研究人员在此基础上通过优化候选框生成、添加自训练模块等多种手段尝试获取更好的检测效果。然而,现有算法仍然难以妥善处理该领域所面临的三大难题,即实例歧义、局部主导和显存消耗,导致算法无法应用于实际任务。因此,本文重点研究基于图像级别标注的效果更鲁棒的弱监督目标检测算法,并尝试实际应用。主要工作包括以下两个方面:1)提出了一种基于类激活图优化的弱监督目标检测算法。该算法可利用弱监督语义分割分支优化弱监督目标检测分支生成的类激活图,再借助优化后的类激活图反过来为弱监督目标检测分支提供更多可靠的前景定位信息,最终通过迭代优化在两个任务上都能取得更好的效果。进一步,在公开数据集上,通过消融实验验证了本文提出算法中弱监督语义分割分支与弱监督目标检测分支相互指导的合理性和有效性,并在困难样本和评估指标上与现有弱监督目标检测算法做出对比,证实了本文提出算法的优越性。2)为将本文提出的弱监督目标检测算法应用于中国航天研究院的重要需求,本文还设计并实现了一个从数据准备到模型训练再到推理预测全流程自动化的算法应用软件。该软件可充分利用本文提出算法帮助研究人员减轻标注负担,通过更多数据提高检测效果,为航天研究中关于空间环境下的卫星环境识别、卫星姿态调整乃至卫星器件检修等实际难题提供低成本解决方案。该软件基于flask框架实现,使用docker镜像部署,具有较高的可移植性和可扩展性。
【Abstract】 Object detection is not only one of the most basic tasks in the field of computer vision,but also an indispensable technology for many advanced applications,such as autonomous driving,face recognition,image retrieval,etc.Benefit from accurate instance-level data annotation,traditional fully-supervised object detection algorithms can achieve extremely high accuracy.However,in actual situations,data annotation is not only expensive to obtain,but also extremely difficult to ensure accuracy.Therefore,researchers began to turn their attention to the research of weakly-supervised object detection algorithms under image-level supervision.In recent years,a basic framework based on multiple instance learning has been formed for weakly-supervised object detection task.On this basis,researchers try to obtain better detection results by optimizing the generation of proposals,adding self-training modules and other methods.However,the existing algorithms are still difficult to properly handle the three major problems in this field(instance ambiguity,part domination and memory consumption),which makes these algorithms unable to be applied to practical tasks.Therefore,this paper focuses on researching more robust weakly-supervised object detection algorithms under image-level supervision and trying to actually apply.The main work includes the following two aspects:1)A weakly-supervised object detection algorithm based on class activation map refinement is proposed.The proposed algorithm can leverage the weakly-supervised semantic segmentation branch to refine the Class Activation Maps(CAMs)generated by the weaklysupervised object detection branch.Then the refined CAMs are utilized to provide more reliable foreground localization cues for the weakly-supervised object detection branch in turn.By such iterative optimizations,better results can be achieved on both tasks.Further,on some public datasets,ablation experiments have been conducted to verify the rationality and effectiveness of the mutual guidance between the weakly-supervised semantic segmentation branch and the weakly-supervised object detection branch in the proposed algorithm.Then the proposed algorithm is compared with existing weakly-supervised object detection algorithms on some difficult samples and some evaluation indicators,which proves the superiority of the algorithm proposed in this paper.2)In order to apply the weakly-supervised object detection algorithm proposed in this paper to the major needs of CSAC,an application software is designed and implemented,which automates the entire process from data preparation to model training to inference.The software can make full use of the algorithm proposed in this paper to help researchers reduce the burden of labeling,improve the detection results through more data,and provide low-cost solutions for some practical problems in aerospace research such as satellite environment recognition,satellite attitude adjustment,and satellite device maintenance in the space environment.The software is implemented based on the flask framework and deployed using docker image,which has relatively high portability and scalability.
【Key words】 Weakly Supervised Learning; Object Detection; Semantic Segmentation; Class Activation Map;
- 【网络出版投稿人】 南京理工大学 【网络出版年期】2024年 05期
- 【分类号】TP391.41