节点文献
面向文物领域的知识图谱构建技术研究
Research on Technologies of Knowledge Graph Construction in Cultural Relics
【作者】 张敏;
【导师】 耿国华;
【作者基本信息】 西北大学 , 软件工程, 2021, 博士
【摘要】 博物馆作为文物保护与传承的载体,承载着人类数千年的文明。在蓬勃发展的物联网和人工智能等技术的驱动下,智慧化建设成为博物馆界关注的热点。然而,由于文物资源种类多数量大,以及互联网相关文物数据的多源异构性,使得博物馆智慧化建设中对文物信息资源管理与利用存在以下两个问题:文物信息资源缺乏有效组织和文物数据之间缺乏有效联系。文物知识图谱挖掘文物事实,利用文物间潜在的联系形成三元组,构建文物知识库,实现文物的有效组织,为文物资源的融合与共享提供基础。同时,文物知识图谱对延展文物知识,丰富文物展陈方式,促进智能问答、语义搜索和智慧导览项目的开发,提升博物馆智慧化服务具有重要意义,其研究吸引了大量研究者关注。文物知识图谱的构建虽已涌现出诸多相关研究,但仍面临以下挑战:(1)文物实体抽取任务中,有监督的方法需要大量标注数据,但构建大规模文物标注数据费时耗力,另外,中文文物实体数据构词具特殊性。(2)文物关系抽取任务中,文物数据存在多重关系,同时文物领域文本关键词具稀疏性。(3)文物实体对齐任务中,百科类网站文物数据具多源异构性,现有的仅从单一的字符或词语层面获取实体相似度的实体对齐方法精确率相对较低。(4)文物知识图谱补全任务中,文物实体间存在隐含关系,以及文物领域具隐含关系的标注三元组相对缺乏。本文针对以上挑战,开展了面向文物领域的知识图谱构建技术研究,主要探讨了文物实体抽取、关系抽取、实体对齐、知识图谱补全等问题,为文物知识图谱构建提供理论和技术支持。本文主要工作和贡献如下:(1)提出一种基于自训练算法的半监督文物实体抽取方法。首先,为解决文物文本构词特殊性问题,使用ELMo语言模型生成词表示,动态提取实体上下文特征;其次,为实现全局最优的标签序列预测,利用Bi LSTM和CRF模型实现特征提取和实体标注;最后,为提高模型的性能,设计一种基于双重标注样本选择策略的自训练算法,通过双重标注选取高置信度的样本。实验结果表明,本文提出的方法利用50%的标注数据在文物实体抽取任务上取得了较好的效果。(2)提出一种基于词注意力机制的胶囊网络文物关系抽取方法。首先,为同时获取语义和语序信息,融合字、词嵌入以及词性和词语位置信息作为模型的输入,以有效捕获语义和语序特征;其次,为解决文物文本关键词稀疏性的问题,设计一种基于词注意力机制的动态路由算法,赋予信息词较高权重,迭代修正连接强度来解决关键词稀疏问题;最后,为解决实体间多重关系问题,利用转换矩阵对胶囊实例化参数预测。实验结果表明,本文提出的方法有效实现了文物领域多重关系的提取。(3)提出一种基于多特征相似度的文物实体对齐方法。首先,针对百科网站文物数据的多源异构性,提取实体属性、实体摘要和实体全文特征,并计算其相似度,分别从字符、词语和句子层面获取实体特征;其次,为了提高实体对齐的精确率,融合实体属性、实体摘要和实体全文特征相似度构建文物实体对齐模型;最后,通过阈值判断两个实体是否对齐。实验结果表明,本文提出的方法在三类文物实体的对齐任务中的精确率分别提高了2.11%,4.98%和4.18%。(4)提出一种融合实体类型的BERT文物知识图谱补全方法。首先,为有效获取实体丰富的语义信息,融合实体类型这一外部知识,使模型消除违反类型约束的反例,实现文本语义增强表示;其次,为有效识别隐含关系,解决关系稀疏性问题,使用多头注意力机制获取文本特征;最后,使用大量无标注数据预训练BERT模型,少量标注文物三元组对模型微调,解决文物领域标注三元组缺乏问题。实验表明,本文提出的方法使用35%标注数据在文物知识图谱补全任务中取得的结果优于对比方法。
【Abstract】 Museums are carriers for the protection and inheritance of cultural relics,and carry the ancient and civilized history.Driven by the booming internet of things,artificial intelligence,wisdom museum has attracted much attention in the museum field.However,due to many kinds and large quantities of cultural relics,and multi-source heterogeneity of cultural relic data on the Internet,there are two problems in the management and utilization of cultural relic information resources: there are lacks of effective organization in cultural heritage resources and effective correlation between the cultural relics.The cultural relic knowledge graph extracts the knowledge of cultural relics,forms a triple by using the potential connections among cultural relics,and constructs the knowledge base of cultural relics to realize the effective organization of cultural relics and provide the foundation for the integration and sharing of cultural relics resources.At the same time,the cultural relic knowledge graph is of great significance for extending cultural relic knowledge,enriching cultural relic display methods,promoting the development of intelligent question-and-answer,semantic search and wisdom tour projects,and improving museum intelligent services.The research of the cultural relic knowledge graph has attracted the attention of a large number of researchers.Although many studies have been done,the following challenges remain in building a high-quality cultural relic knowledge graph.(1)Cultural relic entity extraction,the supervised method requires a large amount of labeled data,but constructing large-scale labeled cultural relic entities is very laborious and time consuming.Besides,the word formation of Chinese cultural relic entity has particularity.(2)Cultural relic relation extraction,there are overlapped relations in cultural relic data,and keywords are sparse.(3)Cultural relic entity alignment,the cultural relic data in encyclopedia websites are multi-source and heterogeneity,and the precision of existing entity alignment methods which only obtain entity similarity from character or word level is relatively low.(4)Knowledge graph completion for cultural relics,there are implicit relations among cultural relic entities,and the labeled triples with implicit relation are scarce.To address the above challenges,this dissertation studies cultural relic entity extraction,relation extraction,entity alignment,and knowledge graph completion to provide support for CRKG construction.The achievements of this dissertation aresummarized as follows:(1)An entity extraction approach for cultural relics based on self-training semi-supervised is proposed.First,ELMo is used to extract the entity context features to solve the problem of word-formation particularity.Then,Bi LSTM and CRF models are constructed for feature extraction and tag prediction respectively to achieve the global optimal tag sequence prediction.Finally,a sample selection strategy of double-labeled for self-training is designed to promote the confidence of sample selection in semi-supervised pretraining,which selectes samples with high confidence by secondary labeling.Experimental results demonstrate that our approach achieves better performance with 50% labeled data in the entity extraction task for cultural relics.(2)A relation extraction approach for cultural relics based on capsule networks with wordattention synamic routing is proposed.First,character and word embedding incorporate part of speech and position to obtain semantic and word order features.Then,a dynamic routing algorithm based on word attention mechanism is designed,the information words are given higher weight and the connection strength is iteratively corrected to solve the problem of keyword sparsity.Finally,the instantiation parameters of advanced capsules are predicted by transformation matrix to realize overlapped relation extraction.Experimental results demonstrate that our approach effectively realizes the overlapped relation extraction.(3)An entity alignment approach for cultural relics based on multi-feature similarity is proposed.First,the entity attribute,entity abstract and entity context features are extracted respectively,and their similarities are calculated to obtain entity features from character,word and sentence levels.Then,the entity alignment model is constructed by integrating the entity attribute,entity abstract,and entity context feature similarities.Finally,the threshold is used to determine whether the two entities are aligned.Experimental results demonstrate that the precision of our approach is 2.11%,4.98% and 4.18% higher than that of comparison models in the entity alignment task,respectively.(4)A knowledge graph completion approach for cultural relics based on the BERT with entity-type information is proposed.First,the entity type as external knowledge is integrated to enhance the representation of text semantics to obtain the rich semantic information of entity effectively and eliminate counterexamples.Then,the multi-head attention mechanism is used to dynamically obtain text features to identify implicit relations effectively,and solve the sparsity of relations.Finally,a small number of labeled triples are used to fine-tune the model to solve the lack of labeled triples.Experimental results demonstrate that our approach utilizes35% labeled data to achieve good results in triple classification,link prediction and relation prediction task for cultural relics.
【Key words】 Smart museum; cultural relic knowledge graph; entity extraction; relation extraction; entity alignment; knowledge graph completion;