节点文献
基于BiLSTM-CRF和深度强化学习的命名实体识别方法
Approaches to Named Entity Recognition Based on BiLSTM-CRF and Deep Reinforcement Learning
【作者】 高翔;
【导师】 高志强;
【作者基本信息】 东南大学 , 软件工程, 2021, 硕士
【摘要】 命名实体识别在自然语言处理领域中具有重要意义,目前主要采用深度学习的方法,如BiLSTM-CRF模型。虽然该模型中的BiLSTM结构可以捕获句子中双向较长距离依赖,但存在以下问题:1)文档级标签一致性指文档中某一特定单词序列的不同出现往往具有相同的实体类别,它是命名实体识别的有效指示,而BiLSTM-CRF模型仅针对句子进行序列标注操作,对文档级标签一致性的利用不够充分;2)仅通过简单地调整超参数或是改变网络结构,模型性能很容易达到瓶颈,其识别结果的质量难以进一步提升。针对以上问题,本文的主要工作如下:(1)针对BiLSTM-CRF模型对文档级标签一致性利用不充分的问题,本文在BiLSTM-CRF模型的基础上添加键值记忆网络(Key-Value Memory Network,KVMN)结构,形成BiLSTM-KVMN-CRF模型。KVMN保存一篇文档经过BiLSTM网络得到的所有隐藏状态向量和相应的标签嵌入向量,在CRF层进行解码前,通过多头注意力机制提取单词在文档中其他出现的上下文信息和标签嵌入,生成文档级上下文表示向量和文档级标签嵌入向量,与当前单词的隐藏状态向量进行融合。实验结果表明,与经典的BiLSTM-CRF模型相比,BiLSTM-KVMN-CRF模型在数据集Co NLL-2003上的F1值为91.48%,提高了0.28%,在数据集Onto Notes5.0上的F1值为87.42%,提高了0.43%。另外,本文还进行了消融实验,分析KVMN中的上下文表示向量和标签嵌入向量对BiLSTM-KVMN-CRF模型的贡献,结果表明两者都对模型性能有提升作用,并且两者同时使用时对于模型的提升效果大于分别单独使用时的提升效果之和。(2)针对BiLSTM-KVMN-CRF模型存在性能瓶颈的问题,本文在BiLSTM-KVMNCRF模型的基础上添加基于深度强化学习的标签修正过程,形成BiLSTM-KVMN-CRFDRL模型。基于深度强化学习的Agent作为标签修正器,设置标签修正阈值,将BiLSTMKVMN-CRF模型的标注结果中不确定度大于该阈值的标签进行修正处理。实验结果表明,与BiLSTM-KVMN-CRF模型相比,BiLSTM-KVMN-CRF-DRL模型在数据集Co NLL-2003上的F1值为92.35%,提高了0.87%,在数据集Onto Notes5.0上的F1值为88.05%,提高了0.63%。另外,不同标签修正阈值对模型性能影响的实验表明,设置合适的标签修正阈值,能有效避免对正确标签进行错误的修正处理。
【Abstract】 Named entity recognition is crucial in natural language processing.Currently,deep learning methods are mainly used,such as BiLSTM-CRF model.Although BiLSTM structure in this model is capable of capturing bidirectional long-distance dependency in sentences,following problems exist: 1)Document-level label consistency is an effective indicator that different occurrences of a particular token sequence are very likely to have the same entity types in a document.But BiLSTM-CRF model itself has sequential nature,labeling only at the sentence level.The constraint prevents the full utilization of document-level label consistency.2)Only by simply tuning the hyperparameters or changing the network structure,the performance of the model can easily reach the bottleneck,and the quality of the recognition results cannot be further improved.To deal with above problem,the main work of this thesis is as follows:(1)Aiming at the problem of unable to take full advantage of document-level label consistency,we add key-value memory network(KVMN)structure to the BiLSTM-CRF model and propose BiLSTM-KVMN-CRF model.KVMN is adopted to memorize all the hidden state vectors and their corresponding label embeddings of the entire document,which the BiLSTM networks yield by taking a document as input.Before CRF decoding,using the multi-head attention mechanism to extract the context information and label embeddings of other occurrences of the word in the document,generating document-level context representation and document-level label embedding,and merge it with the hidden state vector of the current word.The experimental results show that,compared with classic BiLSTM-CRF model,the F1 value of BiLSTM-KVMN-CRF model on Co NLL-2003 dataset is 91.48%,increased by 0.28%;on Onto Notes5.0 dataset,the F1 value is 87.42%,increased by 0.43%.In addition,we also conduct ablation experiments to analyze the contribution of the context representation and the label embedding in KVMN to the BiLSTM-KVMN-CRF model.The results show that both can improve the performance of the model.And when both are used at the same time,the improving effect is greater than the sum of the improving effects used separately.(2)Aiming at the problem of existing performance bottlenecks in the BiLSTM-KVMNCRF model,we add label correction based on deep reinforcement learning to BiLSTM-KVMNCRF model and propose BiLSTM-KVMN-CRF-DRL model.Agent based on deep reinforcement learning is used as a label corrector,setting label correction threshold,to correct the labels with uncertainty greater than threshold in the labeling results of the BiLSTM-KVMNCRF model.The experimental results show that,compared with BiLSTM-KVMN-CRF model,the F1 value of BiLSTM-KVMN-CRF-DRL model on Co NLL-2003 dataset is 92.35%,increased by 0.87%;on Onto Notes5.0 dataset,the F1 value is 88.05%,increased by 0.63%.In addition,experiments on the impact of different label correction thresholds on model performance show that setting a suitable label correction threshold can effectively avoid incorrect correction processing of the correct label.
【Key words】 Named Entity Recognition; BiLSTM; Key-Value Memory Network; Deep Reinforcement Learning;