节点文献

基于深度学习的图像特征提取及其应用

Image Feature Extraction and Application Based on Deep Learning

【作者】 袁琳;

【导师】 赵宏伟;

【作者基本信息】 吉林大学 , 计算机应用技术, 2022, 硕士

【摘要】 信息技术和硬件设施的不断发展进步,使得数字图像中包含的数据量不断增加,图像处理及应用开始受到人们重视。尤其是在遥感图像检索、农作物种子分类及检测、路面裂缝检测、医学图像分割等应用领域,图像中所包含的信息的精确分析将对使用者提供很大的帮助。而随着计算机领域的快速发展,人们开始尝试将“如何看见并看懂图像”这一人类感知世界的重要方式赋予给计算机,随之出现了计算机视觉这个概念。而计算机视觉任务在对图像进行理解和分析的过程中,对于图像的特征提取是非常关键的一步。本文从两个计算机视觉中遥感图像检索和路面裂缝检测两个应用领域入手,在当前已有方法的基础上进行算法改进,旨在挖掘卷积神经网络的特征描述能力,进而提高检索、检测等计算机视觉任务的效率和精准度。本文研究内容:1.为了解决特征提取过程中类内样本相似度差异大的问题,提出了一种基于相似性保持损失(Similarity Retention Loss,SRL)的深度度量学习方法。从样本挖掘、网络模型结构、度量损失函数等方面对现有的度量学习方法进行了改进。在样本挖掘方面,重新定义了困难样本和简单样本,并根据样本集合大小和数据集类的空间分布分别挖掘了合适的正样本和负样本。同时,在度量损失函数方面提出了相似性保持损失的概念。根据类别中难易样本的数量情况为挑出来的难样本分配不同的学习系数,通过这种方式学习相同类别样本的空间结构特征。而负样本则根据空间中同一范围内样本的空间分布给予不同系数,以此来保持相似性空间结构特征。最后使用针对遥感图像数据特征进行修改后的微调网络对两个遥感数据集进行了大量综合实验。2.对于特征提取过程中类间差异小导致样本区分困难的问题,结合深度度量学习两种损失函数设计思路(即结构损失和基于结果的损失),提出了全局感知排序损失(Global-aware Ranking Loss,GRL)模型,这是一个基于特征空间和检索候选列表的全局优化模型。并且提出对于每个数据集中的每个类别来说,类内样本之间的相似度是不同的,因此需要学习的样本和样本数也应该是不同的。提出类内空间样本挖掘(Intra-class Space Sample Mining,ISSM),即根据样本的分布情况选择错位样本,而不是像通常那样人为设置样本挖掘的边界阈值。3.针对更加困难的计算机视觉任务,即从复杂背景中自动提取并区分裂缝,提出一种结合Residual注意力机制的Octave U-Net网络模型,解决特征提取过程中前景和背景不平衡的问题。通过结合Octave卷积和Octave转置卷积,并根据网络特点选择Residual注意力模块辅助模型捕获不同类型的图像特征,同时防止U-Net网络随着网络深度的增加丢失底层语义信息。并且提出改进的加权交叉熵损失,该损失既稳定而且解决了类别失衡问题。

【Abstract】 The high-speed growth of computer technology and hardware facilities has promoted the continuous increase of the amount of information that can be mined in digital images,and image processing and applications have begun to attract people’s attention.Especially in remote sensing image retrieval,crop seed classification and detection,road crack detection and other applications,accurate analysis of the information contained in the image will provide users with great help.Because of the development and progress of computer technology,people are trying to let computers learn "how to understand the images they see",which is a unique skill for humans to understand the world.Therefore,a concept,namely computer vision,has emerged.In this task,the image needs to be analyzed,and the most necessary is vector extraction of samples.This paper starts from the two application fields of remote sensing image retrieval and pavement crack detection in computer vision,and improves the algorithm on the basis of the current existing methods.Efficiency and accuracy of vision tasks.The research content of this paper:1.In view of the large difference in sample similarity in different categories,a deep metric learning method based on Similarity Retention Loss(SRL)is proposed.The existing metric learning methods are improved from the aspects of sample mining,network model structure,metric loss function and so on.In terms of sample mining,difficult samples and simple samples are redefined,and appropriate positive samples and negative samples are mined respectively according to the size of the sample set and the spatial distribution of the dataset classes.At the same time,the concept of similarity preserving loss is proposed.According to the number of difficult and easy samples in the category,different learning coefficients are assigned to the selected difficult samples,and the spatial structure characteristics of the samples of the same category are learned in this way.The negative samples are given different coefficients according to the spatial distribution of samples within a uniform range in space,so as to maintain the similarity spatial structure characteristics.2.For the problem of difficulty in distinguishing samples due to small differences between classes in the feature extraction process,combined with the two loss function design ideas of deep metric learning(ie,structural loss and result-based loss),a globalaware ranking loss(Global-aware Ranking Loss)is proposed.Loss,GRL)model,which is a global optimization model based on feature space and retrieval candidate list.And it is proposed that for each class in each dataset,the similarity between the samples within the class is different,so the samples and the number of samples that need to be learned should also be different.Intra-class Space Sample Mining(ISSM)is proposed,which selects dislocation samples according to the distribution of samples,instead of artificially setting the boundary threshold of sample mining as usual.3.For the more difficult computer vision task,namely automatically extracting and distinguishing cracks from complex backgrounds,an Octave U-Net network model combined with Residual attention mechanism is proposed,so that it can still be done when the foreground and background are not balanced.Combining Octave convolution and Octave transposed convolution,and selecting Residual attention module auxiliary model according to network characteristics to capture different types of image features,at the same time,it can prevent U-Net network from losing underlying semantic information with the increase of network depth.And an improved weighted crossentropy loss is proposed,which is both stable and solves the class imbalance problem.

  • 【网络出版投稿人】 吉林大学
  • 【网络出版年期】2023年 01期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络