节点文献

基于注意力机制的指称图像分割方法研究

Research on Referring Image Segmentation Method with Attention Mechanism

【作者】 刘芳;

【导师】 尹宝才;

【作者基本信息】 大连理工大学 , 计算机技术(专业学位), 2023, 硕士

【摘要】 随着计算机视觉和自然语言处理技术的飞速发展,指称图像分割(Referring Image Segmentation,RIS)逐渐成为跨媒体研究领域的关键课题之一,并在交互式图像编辑和人机交互等多种任务中得到广泛应用。指称图像分割旨在根据自然语言描述找到并分割出图像中对应的目标区域。尽管现有的指称图像分割方法已经取得了不错的效果,但仍存在许多挑战,如仍然不能有效地融合文本和图像特征,以及对高成本真值分割掩码的高度依赖,这些问题限制了指称图像分割方法的发展和应用。为解决这些问题,本文提出了两种指称图像分割方法:一种基于Transformer的指称图像分割方法和一种文本监督下的弱监督指称图像分割方法。为了解决语言和视觉信息的联合表征问题,本文提出了一个基于Transformer的指称图像分割网络。该方法设计了一个混和视觉特征提取模块,可以用来提取多尺度的视觉特征,并在编码器融合阶段整合不同粒度的多模态特征。同时,跨模态特征融合注意模块以序列化的方式结合全局上下文和局部细节特征。此外,该方法还设计了一个跨层信息整合模块,使网络更好地探索编码器中相邻层特征之间的关系。在四个指称图像分割数据集的实验中,该模型也表现出了先进的性能。针对全监督指称图像分割方法中依赖高成本像素级标注的问题,本文提出了一个文本监督下的弱监督指称图像分割方法。该方法的第一阶段利用文本描述生成伪标签,第二阶段将伪标签作为监督信息训练新的指称图像分割网络。首先,该方法学习从一组文本描述中识别出与图像相关的正确文本,通过寻找对正确描述产生高响应的像素来定位目标对象,生成的响应图作为初始分割结果。为了促进网络从预训练模型中学习更多的知识,该方法提出了一个双边提示注意模块,用于协调两个模态特征之间的差异,并提高对目标的定位准确度。为了进一步提高响应图的定位和分割准确度,该方法提出了一种矫正损失来对生成响应图中的前景和背景区域进行矫正。在多个标准数据集上的实验结果表明,即使只使用文本作为监督信号,该方法仍能取得优异的定位和分割性能。最后,本文对两种提出的指称图像分割方法进行了梳理和总结,分析了它们各自的优缺点,并对该研究领域的未来发展方向进行了展望。

【Abstract】 With the rapid development of computer vision and natural language processing technologies,Referring Image Segmentation(RIS)has gradually become one of the key topics in cross-media research,and has been widely used in various tasks such as interactive image editing and human-computer interaction.The goal of RIS is to find and segment the corresponding target area in an image based on natural language descriptions.Although existing RIS methods have achieved good results,there are still many challenges such as the inability to effectively fuse text and image features,and the high dependence on expensive pixel-level segmentation masks,which limit the development and application of RIS methods.To address these challenges,this paper proposes two RIS methods: a Transformer-based RIS method and a weakly-supervised RIS method with text supervision.To address the joint representation problem of language and visual information,this paper proposes a Transformer-based RIS network.The method designs a hybrid visual feature extraction module to extract multi-scale visual features,and integrates different granularity multi-modal features in the encoder fusion stage.Meanwhile,the cross-modal feature fusion attention module combines global context and local detail features in a serialized manner.In addition,the method designs a cross-layer information integration module to better explore the relationship between adjacent layer features in the encoder.In experiments on four RIS datasets,the model also demonstrates advanced performance.To address the problem of high-cost pixel-level annotation in fully-supervised RIS methods,this paper proposes a weakly-supervised RIS method with text supervision.The first stage of this method generates pseudo-labels using text descriptions,and the second stage trains a new RIS network using the pseudo-labels as supervision information.Firstly,the method learns to identify the correct text related to the image from a set of text descriptions,locates the target object by finding pixels that produce high responses to the correct description,and generates the response map as the initial segmentation result.To facilitate the network to learn more knowledge from pre-trained models,the method proposes a bilateral prompt attention module to coordinate the differences between the two modal features and improve the accuracy of target location.To further improve the localization and segmentation accuracy of the response map,the method proposes a correction loss to correct the foreground and background regions in the generated response map.Experimental results on multiple standard datasets show that even using only text as the supervision signal,this method can still achieve excellent localization and segmentation performance.Finally,this paper summarizes and analyzes the two proposed RIS methods,their respective advantages and disadvantages,and prospects for the future development direction of this research field.

  • 【分类号】TP391.41
节点文献中: 

本文链接的文献网络图示:

本文的引文网络