节点文献
多源信息融合的知识图谱构建方法研究与应用
Research and Application of Knowledge Graph Construction Method Based on Multi-source Information Fusion
【作者】 王野;
【导师】 马永军;
【作者基本信息】 天津科技大学 , 软件工程, 2023, 硕士
【摘要】 随着互联网和物联网等技术的飞速发展,数据量大幅增长,这些海量数据中蕴含着丰富的知识和信息,同时数据来源呈现多源化趋势,如食品营养领域的数据包含了丰富的食品营养知识,且具有明显的多源异构性。多源信息融合的知识图谱能够把来自多个数据源的相关知识进行融合,从而提高数据的可用性和价值,推动知识的发现和应用。如何从海量的多源数据中准确地自动抽取出实体类型和实体关系,对于构建知识图谱具有重要意义。本文面向食品营养领域进行多源信息融合的知识图谱构建与应用,主要研究内容和取得的成果有:(1)本文构建了多源信息融合的食品营养领域数据集,分别针对命名实体识别和关系抽取两部分工作进行标注,构建了多源信息融合的食品营养领域词典。(2)针对中文命名实体识别中融合词典信息准确率提升不足的问题,使用在模型内部融合词典信息的策略,提出一种基于LNBC(LE-NEZHA-Bi LSTM-CRF)模型的中文命名实体识别方法。首先通过词典树匹配到所有潜在的词,然后采用NEZHA模型进行融合嵌入表示,将训练得到的字词融合向量输入到Bi LSTM网络中进行特征提取,最后采用CRF层来减少错误标签输出的概率。(3)提出融合对抗训练的全局指针网络关系抽取模型RFG(Ro BERTa-FGMGlobal Pointer)模型,有效解决了食品营养领域文本中存在多种实体和多层次实体关系的问题,提高了模型的鲁棒性和泛化能力。首先利用Ro BERTa作为编码器,学习深层语义特征,然后使用FGM对抗学习算法添加扰动,增强泛化能力,最后将Global Pointer作为解码器,处理关系分类。(4)本文利用图数据库Neo4j对知识图谱进行存储和展示,并基于该图谱设计开发了食品营养信息查询平台。实验表明,本文提出的实体和关系抽取模型在公开数据集和自建食品营养数据集上的表现良好,开发的食品营养信息查询平台运行稳定,有助于提高公众对于食品营养和健康管理方面的意识。
【Abstract】 With the rapid development of Internet and Internet of things technology,the amount of data has increased substantially.These massive data contain rich knowledge and information.At the same time,the data sources show a trend of multi-source.Knowledge graph of multi-source information fusion can fuse related knowledge from multiple data sources,so as to improve the availability and value of data,and promote the discovery and application of knowledge.How to extract entity type and entity relationship accurately and automatically from massive multi-source data is of great significance for the construction of knowledge graph.In this paper,multi-source information fusion knowledge graph is constructed and applied in the field of food nutrition.The main research contents and achievements are as follows:(1)This paper constructs a food nutrition dataset with multi-source information fusion,labels the two parts of named entity recognition and relation extraction respectively,and constructs a food nutrition dictionary with multi-source information fusion.(2)Aiming at the problem of insufficient integration of dictionary information to improve the accuracy of Chinese named entity recognition,a Chinese named entity recognition method based on LNBC(LE-NEZHA-Bi LSTM-CRF)model is proposed by using the strategy of dictionary information fusion.Firstly,all possible words were matched through the dictionary tree,and then NEZHA model was used for fusion embedding representation,and the trained word fusion vector was input into Bi LSTM network for feature extraction.Finally,the CRF layer is used to reduce the probability of incorrect label output.(3)A global pointer network relation extraction model RFG(Ro BERTa-FGM-Global Pointer)model integrating confrontation training is proposed,which effectively solves the problem of multiple entities and multi-level entity relationships in food and nutrition texts,and improves the robustness and generalization ability of the model.Firstly,Ro BERTa is used as an encoder to learn the deep semantic features,then FGM adversarial learning algorithm is used to add perturbations to enhance generalization,and finally Global Pointer is used as a decoder to handle relational classification.(4)In this paper,the graph database Neo4 j is used to store and display the knowledge graph,and a food nutrition information query platform is constructed based on the graph.Experiments show that the entity and relationship extraction model proposed in this paper performs well on the open data set and the self-built food nutrition data set,and the food nutrition information query platform developed runs stably,which helps to improve the public’s awareness of food nutrition and health management.
【Key words】 multi-source information; food nutrition; domain dictionary; entity recognition; relationship extraction;
- 【网络出版投稿人】 天津科技大学 【网络出版年期】2025年 03期
- 【分类号】TP391.1