节点文献

基于自适应和多视图尺度不变学习的半监督目标检测方法

A Semi-Supervised Object Detection Method Based on Adaptive and Multi-view Scale-Invariant Learning

【作者】 杨静

【导师】 周华平;

【作者基本信息】 安徽理工大学 , 计算机技术, 2025, 硕士

【摘要】 半监督目标检测旨在利用有限的标注数据提升模型对未标注数据的检测能力,广泛应用于智能监控、自动驾驶等领域。然而目前的研究仍存在一定的局限性:一是特征提取不充分,难以精准捕捉目标细节信息,影响检测性能;二是固定阈值的标签分配策略缺乏灵活性,导致伪标签质量不稳定;三是目标的尺度变化与视角多样性影响模型的泛化能力,降低检测精度。针对上述挑战,本文的主要工作如下:(1)针对传统半监督目标检测中特征提取不充分、标签分配不合理问题,提出一种基于局部特征增强的自适应半监督目标检测方法。首先提出局部特征增强的模块,在关注全局特征的基础上进一步关注图像中的细粒度信息,从而实现特征表示更加丰富准确。其次设计自适应感知置信度阈值模块,基于样本置信度动态调整标签分配,以避免固定阈值策略导致的伪标签质量下降问题。在MSCOCO和VOC数据集上的实验结果,与基线模型相比,平均检测精度(m AP)分别提升了1.5%和1.75%。(2)针对目标边界框定位偏差和尺度变化影响检测精度的问题,提出一种引导标签分配和多视图尺度不变学习的半监督目标检测方法。首先通过预测结果对标签进行分配指导,根据预测值调节不同边界下的正负样本数目,使得标签分配更符合真实数据分布,提升目标框回归准确性。其次设计多视图尺度不变学习方法,融合多个尺度、多个视角的图像特征,提高模型目标尺度变化以及视角变化下的鲁棒性。该方法不仅提高了模型对复杂场景的适应能力,还促进了不同视图特征的融合,从而优化目标检测性能。在MSCOCO和VOC数据集上的实验结果,与基线模型相比,平均检测精度(m AP)分别提升了0.6%和1.64%。图[27]表[10]参[71]

【Abstract】 Semi-supervised target detection aims to improve the detection capability of the unannotated data model by using the limited annotated data,and is widely used in intelligent monitoring,autonomous driving and other fields.However,the current research still has some limitations:first,inadequate feature extraction,it is difficult to accurately capture the target details,which affects the detection performance;second,the flexibility of the label allocation strategy of fixed threshold,which leads to the instability of pseudo label quality;third,the scale change of the target and the diversity of perspective of the target affect the generalization ability of the model to reduce the detection accuracy.For the above challenges,the main work of this paper is as follows:(1)Aiming at the problems of inadequate feature extraction and unreasonable label assignment in traditional semi-supervised target detection,an adaptive semi-supervised target detection method based on secondary local feature enhancement is proposed.Firstly,the module of secondary local feature enhancement is proposed to further pay attention to the fine-grained information in the image based on the global feature,so as to achieve a richer and more accurate feature representation.Secondly,an adaptive label assignment strategy is designed,and the label allocation method is dynamically selected based on the confidence of each sample,which overcomes the problem of reducing the quality of the pseudo label due to the fixed threshold strategy.Experimental results on MSCOCO and VOC data sets show that the average detection accuracy(m AP)is increased by 1.5%and 1.75%,respectively,compared with the baseline model.(2)To address the problem that target bounding box localization bias and scale variation lead to affect the detection accuracy,a semi-supervised target detection method for prediction-guided label assignment and multi-view scale-invariant learning is proposed.First,the label allocation is guided by the prediction results,and the number of positive and negative samples under different boundaries is adjusted according to the predicted value,so that the label allocation is more in line with the real data distribution,and the regression accuracy of the target box is improved.Secondly,the multi-view scale invariant learning method is designed to integrate the image features of multiple scales and multiple viewing angles,and improve the robustness of the model target scale change and the change of the viewing angle.It can not only improve the robustness of the model for complex scenes,but also promote the fusion of features under different views,and optimize the performance of target detection.Experimental results on MSCOCO and VOC data sets show that the average detection accuracy(m AP)is increased by 0.6%and 1.64%,respectively,compared with the baseline model.Figure[27]Table[10]reference[71]

  • 【分类号】TP391.41;TP18
节点文献中: 

本文链接的文献网络图示:

本文的引文网络