节点文献

融合attention机制的BI-LSTM-CRF中文分词模型

BI-LSTM-CRF Chinese Word Segmentation Model with Attention Mechanism

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 黄丹丹郭玉翠

【Author】 HUANG Dan-dan;GUO Yu-cui;School of Science,Beijing University of Posts and Telecommunications;

【机构】 北京邮电大学理学院

【摘要】 中文的词语不同于英文单词,没有空格作为自然分界符,因此,为了使机器能够识别中文的词语需要进行分词操作。深度学习在中文分词任务上的研究与应用已经有了一些突破性成果,本文在已有工作的基础上,提出融合Bi-LSTM-CRF模型与attention机制的方法,并且引入去噪机制对字向量表示进行过滤,此外为改进单向LSTM对后文依赖性不足的缺点引入了贡献率?对BI-LSTM的输出权重矩阵进行调节,以提升分词效果。使用改进后的模型对一些公开数据集进行了实验。实验结果表明,改进的attention-BI-LSTM-CRF模型以及训练方法可以有效地解决中文自然语言处理中的分词、词性标注等问题,并较以前的模型有更优秀的性能。

【Abstract】 In English words,spaces are used as natural delimiters between words,and there are no such clear delimiters between Chinese words.Therefore,deep learning models and methods that obtain good results in English natural language processing cannot be directly applied.Deep learning has achieved breakthrough results in the field of natural language processing in English.Based on the existing work,this paper proposes a method to integrate the Bi-LSTM-CRF model and the attention mechanism,and introduces a denoising mechanism to filter the word vector representation.In addition,the contribution rate ? of the unidirectional LSTM is reduced.The output weight matrix of the BI-LSTM is adjusted to improve the word segmentation effect.We conducted experiments using the public data set in the above model.Experimental results show that the improved attention-BI-LSTM-CRF model and training method can effectively solve the problem of word segmentation and part of speech tagging in Chinese natural language processing,and can obtain good performance.

  • 【文献出处】 软件 ,Computer Engineering & Software , 编辑部邮箱 ,2018年10期
  • 【分类号】TP391.1
  • 【被引频次】23
  • 【下载频次】628
节点文献中: 

本文链接的文献网络图示:

本文的引文网络