节点文献

基于深度学习的时钟领域命名实体识别方法研究

Research on Named Entity Recognition in Clock Domain Based on Deep Learning

【作者】 王慧;

【导师】 唐焕玲; 史亚萍;

【作者基本信息】 山东工商学院 , 工程硕士(专业学位), 2021, 硕士

【摘要】 目前时钟系统已经应用到国防、经济、金融、工业等各个领域,且时钟系统的复杂性越来越高,这就使时钟系统在方案设计、生产制造、售后服务等方面存在诸多的行业问题,而知识图谱技术在时钟领域的应用有利于解决行业痛点问题。通过对时钟领域中的结构化和非结构化的文本进行分析,构建时钟领域知识图谱,能够协助时钟领域专业人员在面向不同的用户时,智能化地提供解决方案,既能多方面的满足用户的需求,降低成本,还能为后续的知识发现提供基础的支持。其中,命名实体识别是构建知识图谱中基础乃至关键的一步。因此,研究如何提升时钟领域文本的实体识别效果,具有重要的意义。针对目前时钟领域的问题,本文重点对时钟领域文本和深度学习网络模型进行了深入探讨,开展了如下创新型研究。(1)针对时钟领域缺乏标注数据集的问题,定义时钟领域实体类别,设计辅助标注平台,构建高质量的时钟领域标注数据集(clock-dataset)。利用互信息和左右邻接熵的新词发现算法来发现新词,分析时钟领域专业术语,定义时钟领域实体类别。选择适合本任务的实体标注策略和规范,设计辅助标注平台,从而高效的构建高质量的clock-dataset,为后续命名实体识别奠定基础。(2)针对时钟领域标注样本数量少的问题,提出一种BERT-LCRF的时钟领域命名实体识别模型。利用预训练语言模型BERT进行时钟领域文本的特征提取,然后利用线性链条件随机场(Linear-CRF)方法进行序列标注。对比实验结果表明,该模型能够充分学习时钟领域的特征信息,提升序列标注精度,进而提升时钟领域的命名实体识别效果。(3)设计与实现时钟领域实体识别系统,提供接口供企业调用使用。该平台实现了数据预处理、实体类别定义、辅助标注、模型的训练、测试和评估的功能。不仅能够满足时钟领域专业人员对数据分析、构建标注数据集和实体识别的需求,还为后续知识图谱的构建打下坚实的基础,充分证明本文所提模型的实用性和有效性。综上所述,本文提出的方法能够进一步的提升时钟领域命名实体识别任务的效果,解决目前时钟领域所遇到的问题,为时钟领域的实体识别技术提供一种可行性的方法,最终为知识图谱的构建打下坚实的基础。

【Abstract】 At present,the clock system has been applied to national defense,economy,finance,industry,communication and other fields,and the complexity of the clock system is getting higher and higher,which brings many industry problems to the design of the clock system scheme,production and manufacturing,after-sales service,after-sales question and answer,and the application of knowledge graph technology in the field of clock is conducive to solving the problem of industry pain points.By analyzing structured and unstructured texts in the clock field,building a knowledge graph in the clock field can assist clock professionals in providing intelligent solutions when facing different users,which can satisfy users in many aspects.Needs,reduce costs,and provide basic support for subsequent knowledge discovery.Among them,named entity recognition is a basic and even key step in building a knowledge graph.Therefore,it is of great significance to study how to improve the entity recognition effect of text in the clock domain.In response to the current problems in the clock domain,this article focuses on the clock domain text and deep learning network model for in-depth discussion,and carried out the following innovative research.(1)In view of the lack of label data sets in the clock field,define the entity categories of the clock domain,design the auxiliary labeling platform,and build a high-quality clockdataset data set.New word discovery algorithms using mutual information and left and right adjacent entropy are used to find new words,analyze the professional terms in the field of clock,and define the categories of entities in the field of clock.Select entity labeling policies and specifications that are appropriate for this task,design an auxiliary labeling platform,and efficiently build high-quality clock-dataset to lay the foundation for subsequent naming entity identification.(2)In view of the problem of entity nesting and small number of label samples,a BERT-LCRF clock domain named entity recognition model is proposed.The pre-trained language model BERT is used to extract the characteristics of the text of the clock field,and then the linear chain conditional random field(Linear-CRF)method is used for sequence labeling.The comparative experimental results show that the model can fully study the characteristic information in the field of clock,improve the accuracy of sequence labeling,and then improve the recognition effect of named entities in the field of clock.(3)Design and implement the clock domain entity recognition system,provide an interface for enterprise call use.The platform realizes the functions of data pre-processing,entity category definition,auxiliary labeling,model training,testing and evaluation.It can not only meet the needs of clock professionals for data analysis,construction of label data sets and entity recognition,but also lay a solid foundation for the construction of subsequent knowledge graph,which fully proves the practicality and validity of the model proposed in this paper.In summary,the method proposed in this paper can further enhance the effect of the task of naming entities in the clock domain,solve the problems encountered in the field of clock,provide a feasible method for the entity recognition technology in the field of clock,and finally lay a solid foundation for the construction of knowledge graph.

  • 【分类号】TP391.1;TP18;TH714
  • 【被引频次】1
  • 【下载频次】128
  • 攻读期成果
节点文献中: 

本文链接的文献网络图示:

本文的引文网络