节点文献
基于知识图谱的遥感影像智能推荐研究
Research on Remote Sensing Images Intelligent Recommendation Based on Knowledge Graph
【作者】 王飞;
【导师】 张永军;
【作者基本信息】 武汉大学 , 摄影测量与遥感, 2023, 博士
【摘要】 随着对地观测技术的快速发展,遥感影像的数量呈现爆炸式增长。遥感影像在国家战略决策、社会经济发展、人们日常生活中扮演着举足轻重的角色。人类探索空间的外延以及遥感影像应用范围的扩展使得各行业用户对实现遥感影像高效、精确的共享服务提出了更高的要求。然而由于信息知识的匮乏,人们难以从海量遥感影像数据集中精准获取目标影像,面临着“数据成果找不到”的尴尬局面。因此,如何从遥感影像中提取知识,并服务于用户对影像的精准化、个性化获取需求,成为实现遥感影像全面共享和大众化的痛点和难点。现有的遥感影像共享模式要求用户总能提供与目标影像具有强一致性的关键字或者影像作为检索源,而当用户缺乏相关数据时,检索结果则难以满足他们的需求,是一种检索精度不稳定、用户体验感较差的被动式服务方法。带有文本反馈的图像检索允许用户对检索结果进行修改,推荐算法允许用户不提供输入而是根据用户的历史行为进行主动的数据分发服务。本文从数据形成知识、知识应用于服务两个过程出发,围绕遥感影像知识图谱构建、知识图谱应用于带有文本反馈的影像检索和影像的个性化推荐展开研究。具体研究内容包括以下四个方面:(1)构建了遥感影像知识图谱。设计了自顶向下的遥感知识图谱构建流程,该方法在领域知识的模式层约束下,从多种类型的数据中提取知识。然后在该构建流程下,以遥感影像为数据源,以影像的语义分割图为基础构建包含影像特征、时间特征和空间特征的场景图,并通过实体链接形成仅包含文本实体的单模态遥感影像知识图谱以及同时包含文本实体和影像实体的多模态遥感影像知识图谱。通过连接遥感影像信息形成知识网络,实现影像从数据到知识的转变。(2)针对带有文本反馈的遥感影像检索,提出了知识感知的多模态信息全局融合网络。场景图是对影像的结构化表示,有助于对场景内容的再次表达;场景图之间的差异反映了影像之间对象级的差异,与修改文本的描述尺度一致。基于此,本文引入场景图作为输入的一部分。首先融合场景图和影像的多级特征,以生成同时具有局部性和全局性表达的场景特征。然后舍弃场景特征的风格信息,使用多模态全局内容模块根据文本的内容信息修改场景特征的内容信息,该模块摒弃或简化了解耦的多模态非局部模块中复杂的矩阵计算。最后使用仿射变换根据文本的风格信息为场景特征重新引入风格信息。该方法使用场景图增强了影像之间的差异化表达,并通过优化矩阵计算减少了特征修改过程中的参数。在三个数据集上的大量实验,证明了本文方法的有效性。(3)针对遥感影像的个性化推荐,提出了知识图谱感知的轻量级图卷积神经网络。标准图卷积神经网络中的特征变换矩阵和非线性激活函数对基于图卷积神经网络的推荐算法的性能具有抑制作用,本文通过消融实验验证了它们的联合表达会给基于图卷积神经网络的知识图谱感知的推荐算法带来负面影响。然后使用舍弃了这两个组件的轻量级图卷积层分别聚合用户-遥感影像交互图和遥感影像知识图谱中的信息。在聚合知识图谱中的信息时,设计了考虑关系对用户和中心节点重要性的注意力机制,生成邻居节点的重要性。该方法充分挖掘了两种数据源中的信息,并且没有引入额外的可训练参数。在两个数据集上的实验结果证明了本文方法在推荐场景和冷启动场景中的有效性。(4)为了提高推荐算法对多模态数据的应用能力,提出了多模态知识图谱感知的深层图注意力网络。现有的算法未充分利用遥感影像的视觉特征,而且受限于过平滑问题只能提取浅层的协同信号。本文使用多模态遥感影像知识图谱加强多模态节点的特征表达。然后通过实验验证了只提取浅层协同信号时,模型对高阶影像的推荐能力较差,因此引入初始残差连接缓解过平滑问题,聚合更高阶协同信号。在特征聚合过程中,设计了关系注意力机制评价邻居节点的重要性。最后在训练阶段,使用额外的知识推理损失函数辅助模型的优化。该方法充分利用了影像的多模态特征和图中的高阶协同信号。实验结果表明本文方法在推荐场景、高阶推荐场景和冷启动场景中均表现出优异的性能。综上所述,本文以遥感影像知识图谱为重要辅助信息,针对遥感影像的被动检索和主动推荐,分别提出了知识感知的多模态信息全局融合网络、知识图谱感知的轻量级图卷积神经网络和多模态知识图谱感知的深层图注意力网络。大量实验证明了本文提出方法的有效性,对实现遥感影像的共享具有一定的现实意义。
【Abstract】 With the rapid development of earth observation technology,the number of remote sensing images has shown explosive growth.Remote sensing images play a pivotal role in national strategic decision-making,social and economic development,and people’s daily life.The outreach of human exploration space and the expansion of remote sensing image applications have led to higher requirements from users in various industries for efficient and accurate sharing services of remote sensing images.However,due to the lack of information knowledge,it is difficult for people to accurately obtain target images from the massive remote sensing image dataset,and they face the embarrassing situation of "data results cannot be found".Therefore,how to extract knowledge from remote sensing images and serve to users’ demand for the accurate and personalized acquisition of images has become a pain point and difficulty for the comprehensive sharing and popularization of remote sensing images.Existing remote sensing image-sharing pattern require users to always provide keywords or images with strong consistency with the target image as the retrieval source.When users lack relevant data,the retrieval results can hardly meet their needs,which is a passive service method with unstable retrieval accuracy and poor user experience.Image retrieval with text feedback allows users to modify the retrieval results,and recommendation algorithms allow active data distribution services based on users’ historical behavior instead of providing input.This paper starts from the two processes of data forming knowledge and knowledge applying to services and focuses on the construction of a remote sensing image knowledge graph,the application of knowledge graph to image retrieval with text feedback,and personalized recommendation of images.The specific research includes the following four aspects.(1)A remote sensing image knowledge graph is constructed.An up-down remote sensing knowledge graph construction process is designed,which extracts knowledge from various type of data under the constraint of the schema layer of domain knowledge.Then,under this construction process,with remote sensing images as the data source,the scene graph containing image features,temporal features,and spatial features is constructed based on the semantic segmentation results of the image,and a unimodal remote sensing image knowledge graph containing only text entities and a multi-modal remote sensing image knowledge graph containing both text entities and image entities are formed by entity linking.By connecting remote sensing image information into a knowledge network,the transformation from data to knowledge is realized.(2)For remote sensing image retrieval with text feedback,a Knowledge-aware Multi-modal Information Global Fusion Network is proposed.Scene graphs are structured representations of images,which help to re-express the scene content;the differences between scene graphs reflect the object-level differences between images and are consistent with the description scale of the modified text.Based on this,this paper introduces scene graphs as part of the input.The multi-level features of scene graphs and images are first fused to generate scene features with both local and global representations.Then discard the style information of the scene features and modify the content information of the scene features based on the content information of the text using the Multi-modal Global Content block,which discards or simplifies the complex matrix computation in the Disentangled Multi-modal Non-Local block.Finally,the affine transformation is used to reintroduce style information for the scene features based on the style information of the text.The method enhances the differentiated representation between images using scene graphs and reduces the parameters in the feature modification process by optimizing the matrix computation.Extensive experiments on three datasets demonstrate the effectiveness of our method and its ability to maintain an advanced performance despite incomplete scene graph construction.(3)For personalized recommendation of remote sensing images,a Knowledge graph-aware Light Graph Convolutional Network is proposed.The feature transformation matrix and nonlinear activation function in the standard graph convolutional network have a suppressive effect on the performance of graph convolutional network-based recommendation algorithms,in this paper,we verify through ablation experiments that the joint representation of them negatively affects the performance of graph convolutional network-based knowledge graph-aware recommendation algorithms.Then a light graph convolutional layer with these two components discarded is used to aggregate the information in the user-remote sensing image interaction graph and the remote sensing image knowledge graph,respectively.When aggregating the information in the knowledge graph,an attention mechanism that considers the importance of the relationship between the user and the central node is designed to generate the importance of the neighbor nodes.The method fully exploits the information in both data sources and does not introduce additional trainable parameters.Experimental results on two datasets demonstrate the effectiveness of our method in both recommendation scenarios and cold-start scenarios.(4)In order to improve the application capability of the recommendation algorithm to multi-modal data,a Multi-modal Knowledge graph-aware Deep Graph Attention Network is proposed.Existing methods do not fully utilize the visual features of remote sensing images and are limited by the over-smoothing problem that only shallow collaborative signals can be extracted.In this paper,we use a multi-modal remote sensing image knowledge graph to enhance the feature representation of multi-modal nodes.Then,it is experimentally verified that the model has poor recommendation ability for higher-order images when only shallow collaborative signals are extracted,so the initial residual connection is introduced to alleviate the over-smoothing problem and aggregate higher-order collaborative signals.During feature aggregation,a relational attention mechanism is designed to evaluate the importance of neighboring nodes.Finally,in the training phase,an additional knowledge reasoning loss function is used to assist in the optimization of the model.The method makes full use of the multi-modal features of the images and the higher-order collaborative signals in the graph.Experimental results show that our method shows superior performance in recommendation scenarios,higher-order recommendation scenarios,and cold-start scenarios.In summary,for passive retrieval and active recommendation of remote sensing images,this paper proposes Knowledge-aware Multi-modal Information Global Fusion Network,Knowledge graph-aware Light Graph Convolutional Network and Multi-modal Knowledge graph-aware Deep Graph Attention Network,respectively,using remote sensing image knowledge graph as important auxiliary information.Extensive experiments have proved the effectiveness of the method proposed in this paper,which has certain practical significance for realizing the sharing of remote sensing images.
【Key words】 remote sensing images; knowledge graph; multi-modal knowledge graph; image retrieval with text feedback; recommendation;
- 【网络出版投稿人】 武汉大学 【网络出版年期】2025年 08期
- 【分类号】P237