节点文献

基于对比学习的场景图像识别与分割技术研究

Research on Scene Image Recognition and Segmentation Based on Contrastive Learning

【作者】 李朝阳;

【导师】 武港山;

【作者基本信息】 南京大学 , 计算机技术(专业学位), 2021, 硕士

【摘要】 场景识别和语义分割作为计算机视觉领域的两个重要的基础任务,一直以来都是研究人员关注的重点。场景识别需要准确识别场景内多个物体的语义,并考虑场景内不同物体之间的关系以及整体的环境,才能对场景图像进行精确分类。语义分割则不仅需要准确理解场景内容的语义信息,还需要对每个物体甚至像素级的信息进行空间上的定位。因此场景识别和语义分割具有较高的挑战性。近年来,受益于大规模数据集和高性能GPU资源的发展,场景识别和语义分割均取得了显著的成效。在结合多模态数据进行场景识别的任务上,研究人员大多采用多个分支网络单独训练不同模态的模型,然后对多个分支的特征进行融合或者采用模态翻译的方式融合多模态数据。由于缺少对应的初始化模型以及多模态数据之间的数据分布差异等问题,这类方法往往只能利用多模态数据的部分有效信息。在语义分割任务上,目前主流的方法都使用空间金字塔结构的方式进行语义分割,这种方式虽然可以获得有用的多尺度信息,但是难以学习重要的整体结构信息。考虑到对比学习在自监督学习领域中的突出性能,将对比学习与场景识别和语义分割结合后,可以充分利用对比学习提高主干网络区分相关样本和无关样本的能力,从而提高场景识别和语义分割的性能。本文针对场景识别和语义分割的任务特性,分别设计了不同的对比学习框架以及使用不同的模态数据进行实验。具体来说,本文的主要贡献如下:●在基于对比学习的场景识别任务中,本文设计了一种基于对比学习的自监督特征学习框架,用于训练场景识别任务的初始化模型,利用对比学习提升主干网络区分相关样本和无关样本的能力。本文分别将RGB图像和Depth图像作为模型的输入模态和目标模态,并通过模态翻译操作减少RGB图像和Depth图像之间的数据分布差异。本文还使用了一个生成对抗模块,用于约束生成模态和目标模态的高级语义特征的内部分布趋于一致,进一步增强主干网络对相关样本的编码能力。考虑到场景识别任务更关注高级语义信息,本文使用深层卷积网络提取的高级语义特征构建对比学习的正负样本。●在基于对比学习的语义分割任务中,本文将对比学习和语义分割结合形成一个多任务模型。由于彩色化的语义分割标签图像具有清晰的整体结构信息,本文将RGB图像和彩色化的语义分割标签图像分别作为输入模态和目标模态,同样使用模态翻译操作减少RGB图像和彩色化的语义分割标签图像之间的数据分布差异。考虑到场景的整体结构信息对语义分割的重要性,本文使用浅层卷积网络提取的整体结构特征构建对比学习的正负样本。本文将提出的两种基于对比学习的技术分别在对应的公开数据集上进行了详尽的实验。实验结果表明,本文提出的两种技术均获得了明显高于基线的突出性能。

【Abstract】 Scene recognition and semantic segmentation as two important basic tasks in the field of computer vision have been the focus of researchers.Scene recognition needs to accurately identify the semantics of multiple objects in the scene,and consider the relationship between different objects in the scene and the overall environment in order to accurately classify the scene images.Semantic segmentation requires not only accurate understanding of the semantic information of the scene content,but also spatial positioning of each object and even pixel-level information.Therefore,scene recognition and semantic segmentation are very challenging.In recent years,benefiting from the development of large datasets and high performance GPU resources,both scene recognition and semantic segmentation have achieved remarkable results.In the task of combining multi-modal data for scene recognition,researchers mostly use multiple branch networks to train models of different modality separately,and then fuse the features of multiple branches or use modality translation to fuse the multi-modal data.Due to the lack of the corresponding initialization model and the difference of the data distribution among the multi-modal data,such methods can only use part of the valid information of the multi-modal data.In terms of semantic segmentation task,the current mainstream methods all use spatial pyramid structure for semantic segmentation.Although this method can obtain useful multi-scale information,it is difficult to learn important overall structure information.Considering the outstanding performance of contrastive learning in self-supervised learning,the combination of contrastive learning with scene recognition and semantic segmentation can make full use of contrastive learning to improve the ability of backbone network to distinguish relevant samples from irrelevant samples,so as to improve the performance of scene recognition and semantic segmentation.Aiming at the task characteristics of scene recognition and semantic segmentation,this paper designs different contrastive learning frameworks and uses different modality to conduct experiments.Specifically,the main contributions of this paper are as follows:● In the task of scene recognition based on contrastive learning,a self-supervised feature learning framework is designed to train the initialization model of the scene recognition task,and contrastive learning is used to improve the ability of the backbone network to distinguish between relevant samples and irrelevant samples.In this paper,RGB image and Depth image are taken as the input modality and target modality of the model respectively,and the difference of data distribution between RGB image and Depth image is reduced through modality translation.This paper also uses a generative adversarial module to constrain the high-level semantic features of the generated modality and the target modality to further enhance the ability of the backbone network to encode relevant samples.Considering that the scene recognition task pays more attention to high-level semantic information,this paper uses the high-level semantic features extracted from deep convolutional networks to construct positive and negative samples for contrastive learning.●In the task of semantic segmentation based on contrastive learning,this paper combines contrastive learning and semantic segmentation to form a multi-task model.In this paper,the RGB image and the colorized semantic segmentation label are taken as the input modality and the target modality respectively because the colorized semantic segmentation label has clear overall structure information.The modality translation operation is also used to reduce the data distribution difference between the RGB image and colorized semantic segmentation label.Considering the importance of the overall structure information of the scene for semantic segmentation,this paper uses the overall structure features extracted from shallow convolutional networks to construct positive and negative samples for contrastive learning.The two techniques based on contrastive learning proposed in this paper are conducted detailed experiments on the corresponding open datasets respectively.The experimental results show that the two techniques presented in this paper achieve outstanding performance significantly higher than the baseline.

  • 【网络出版投稿人】 南京大学
  • 【网络出版年期】2022年 05期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络