节点文献

基于深度学习的中文林业知识图谱的构建研究

Research on Construction of Chinese Forestry Knowledge Graph Based on Deep Learning

【作者】 李想;

【导师】 岳琪;

【作者基本信息】 东北林业大学 , 管理科学与工程, 2021, 硕士

【摘要】 随着互联网的飞速发展,数据资源不断积累,中文文本信息呈指数式增长,数据价值仍未被充分挖掘,尤其是在林业方面。进入现代以来,林业管理任务日益复杂,产生了大量冗杂的林业知识,当前迫切需要一种快捷高效的林业信息管理方法。近年来,研究人员开始探索将知识图谱应用于林业领域,知识图谱具有强大的语义处理和开放互联能力,有助于从冗杂的数据中快速提取有效信息,构建林业知识图谱可以融合碎片化林业文本数据,解决目前林业网络中存在的知识散乱、无序和关联性不强等问题。构建中文林业知识图谱是整合林业知识和管理林业信息的一种新方法,本文主要工作内容如下。(1)论文围绕中文林业知识图谱的构建进行了系统的研究,通过调研国内外知识图谱构建技术,林业现状,以及子任务命名实体识别和实体关系抽取的现状,分析相关工作的不足。针对林业现状混乱和传统知识图谱的构建依赖于统计机器学习和专家知识的问题,提出基于深度学习的中文林业知识图谱构建方法。(2)优化融合BERT和双向RNN的方法用于林业实体识别和实体关系抽取任务,提高林业信息抽取的准确率。基于全词Mask的BERT模型预处理文本,可以自动提取出序列中丰富的词级和语义特征,解决目前信息抽取中词向量化处理不能充分利用先验知识和不能充分联系上下文的问题。双向RNN的变体(BiLSTM,BiGRU)对句子进行双向分析建模,加强上下文语义联系,充分利用文本结构信息,提高林业知识抽取的效率。(3)为验证模型的有效性和通用性,在不同数据集上进行对比实验。在常见数据集上,命名实体识别BERT-BiLSTM-CRF模型的F1值达到96%,实体关系抽取BERT-BiGRU-Dual Attention模型的F1值也提升到85%,性能良好;在自建林业数据集上,基于BERT-BiGRU-Dual Attention的实体关系抽取实验的查全率和准确率分别可以达到75%和80%以上,基于BERT-BiLSTM-CRF命名实体识别实验的准确率达到90%。(4)通过对林业知识图谱的构建研究,提出一种自下而上构造林业知识图谱的方法,融合Django开发框架以及林业实体命名识别、林业实体关系抽取等关键算法构建林业知识图谱系统。

【Abstract】 With the rapid development of the Internet,data resources continue to accumulate,Chinese text information is increasing exponentially,and the value of data is still not fully explored,especially in forestry.Since entering modern times,forestry management tasks have become more and more complex,resulting in a large amount of redundant forestry knowledge,and there is an urgent need for a fast and efficient forestry information management method.In recent years,researchers have begun to explore the application of knowledge graph to the forestry field.Knowledge graph has powerful semantic processing and open interconnection capabilities,which help to quickly extract effective information from redundant data.Building a forestry knowledge graph can fuse fragmented forestry text data,and solve the problems of scattered,disordered,and weakly related knowledge in the current forestry network.Constructing a Chinese forestry knowledge graph is a new method to integrate forestry knowledge and manage forestry information resources.The main contents are as follows.(1)The dissertation conducted systematic research on the construction of Chinese forestry knowledge graph and analyzed the deficiencies of related work by investigating domestic and foreign knowledge graph construction technology,forestry status,and the status of subtask named entity recognition and entity relation extraction.In view of the confusion of forestry status and the construction of traditional knowledge graph relying on statistical machine learning and expert knowledge,a method for constructing Chinese forestry knowledge graph based on deep learning is proposed.(2)The method of optimizing the fusion of BERT and bidirectional RNN is used for forestry entity recognition and entity relation extraction tasks to improve the accuracy of forestry information extraction.The BERT model preprocessing text based on the whole word Mask can automatically extract the rich word-level and semantic features in the sequence,and solve the problem that the current word vectorization processing in information extraction cannot make full use of prior knowledge and contextual semantics.A variant of the bidirectional RNN(BiLSTM,BiGRU)performs bidirectional analysis and modeling of sentences,strengthens contextual semantic connection,and makes full use of text structure information,to improve the efficiency of forestry knowledge extraction.(3)In order to verify the validity and versatility of the model,comparative experiments are carried out on different datasets.On the common dataset,the F1-score of the named entity recognition BERT-BiLSTM-CRF model reached 96%,and the Fl-score of the entity relation extraction BERT-BiGRU-Dual Attention model also increased to 85%,with good performance;on the self-built forestry dataset,the recall and accuracy of the entity relation extraction experiment of BERT-BiGRU-Dual Attention can reach 75%and 80%respectively,and the accuracy of the named entity recognition experiment based on BERT-BiLSTM-CRF reached 90%.(4)Through the research on the construction of forestry knowledge graph,a bottom-up method of constructing forestry knowledge graph is proposed,and the forestry knowledge graph system is constructed by integrating the Django development framework and key algorithms such as forestry entity named recognition and forestry entity relation extraction.

  • 【分类号】F307.2;G353.1
  • 【被引频次】1
  • 【下载频次】470
  • 攻读期成果
节点文献中: 

本文链接的文献网络图示:

本文的引文网络