节点文献

基于弱监督的细粒度图像分类研究

Research on Weakly Supervised Fine-Grained Visual Classification

【作者】 王瑜;

【导师】 尤新革;

【作者基本信息】 华中科技大学 , 信息与通信工程, 2023, 硕士

【摘要】 细粒度图像分类旨在对属于同一基础类别的图像进行更加细致的子类划分,相关技术被广泛应用于军事目标检测、农林农业、医学诊断、工业检测等领域,具有重要的研究和应用价值。由于细粒度图像具有相似的外观和特征,加之采集过程中存在姿态、遮挡等干扰的影响,数据往往呈现出较小的子类间差异和较大的子类内差异。因此,解决该问题的关键在于如何有效学习图像中的局部判别性特征。早期方法通常依靠人工标注的强监督信息来辅助模型定位局部判别性区域,然而标注数据获取十分昂贵,难以适应现实场景中的需求。相比之下,弱监督细粒度图像分类方法仅利用类别标签就可以实现较为准确的分类,引起了学术界与工业界的广泛关注。本文以弱监督细粒度图像分类为研究对象,针对现有方法存在的不足和缺陷,提出了以下改进:针对现有方法忽视了单张图像因姿态、遮挡等因素导致的判别性信息不完整的问题,本文提出了一种基于多图语义特征融合的细粒度图像分类模型。该模型首先通过随机交换同一子类不同图像之间的局部语义特征,获取更丰富的语义特征组合模式,提升了特征的多样性。同时通过随机交换不同子类图像的背景区域特征,增强了模型在多样化环境中学习局部语义特征的能力。最后使用知识蒸馏指导模型融合多分支、多层次的语义知识,进一步提升了特征的表达能力。三个数据集上的实验结果说明了该模型的有效性。针对现有方法在图像特征编码中存在大量冗余信息造成模型难以聚焦到判别性区域的问题,本文提出了一种基于自注意力机制和信息瓶颈的细粒度图像分类模型。该模型首先根据自注意力机制动态地掩盖无关或者冗余的图像块,以减少背景区域的干扰。其次利用多层特征选择模块筛选判别性特征,提升了模型对多层次细节特征的捕捉能力。并且在整个学习过程中利用信息瓶颈损失压缩特征中的冗余信息,进一步引导模型学习最小充分特征表示。实验结果显示该模型在三个数据集上均取得了优于其他方法的分类精度,同时可视化结果说明了该模型能够更加聚焦关键区域。

【Abstract】 Fine-grained Visual Classification(FGVC)aims to classify images belonging to the same base class into more detailed subclasses,and related techniques are widely applied in military target detection,agriculture and forestry,medical diagnosis and industrial inspection,which have important research and application values.However,fine-grained images often exhibit small inter-subclass differences and large intra-subclass variations due to their similar visual appearance and the presence of pose,occlusion,and other distractions during acquisition.Therefore,the key to solving this problem lies in learning local discriminative features in the images.Early methods often relied on human-annotated strong supervision information to assist models in locating discriminative regions.However,acquiring such labeled data can be prohibitively expensive and difficult to adapt to real-world scenarios.In contrast,weakly supervised FGVC methods can achieve relatively accurate classification using only class labels,attracting widespread attention from academia and industry.This paper focuses on weakly supervised FGVC and proposes several improvements to address the deficiencies of existing methods.The contributions of this paper can be summarized as follows:To address the problem that existing methods ignore incomplete discriminative information of single images due to factors such as pose and occlusion,we propose a model based on Multi-image Semantic Feature Fusion(MSFF).The MSFF first randomly exchanges local semantic features between different images within the same subclass,obtaining a richer combination of semantic features and enhancing feature diversity.Moreover,it randomly exchanges background region features from different subclass images,improving the model’s ability to learn local semantic features in diverse environments.Finally,knowledge distillation is used to guide the model to fuse multibranch and multi-level semantic knowledge,further improving feature representation capability.Experimental results on three datasets demonstrate the effectiveness of the proposed MSFF.To address the problem that existing methods have difficulty focusing on discriminative regions due to the presence of a large amount of redundant information in image feature encoding,we propose a model based on self-attention mechanism and information bottleneck(R2-Trans).The R2-Trans first dynamically masks irrelevant or redundant image patches based on the self-attention mechanism,reducing the interference of background regions.Then,a multi-layer feature selection module is used to filter discriminative features,which improves the model’s ability to capture multi-level fine-grained features.Throughout the learning process,the information bottleneck loss is used to compress redundant information in features,further guiding the model to learn minimal sufficient feature representation.Experimental results show that the proposed R2-Trans achieves better accuracy than other methods on all three datasets,and visualization results illustrate that the R2-Trans is able to focus more on important regions.

  • 【分类号】TP391.41
节点文献中: 

本文链接的文献网络图示:

本文的引文网络