节点文献

基于多模型融合的汉语介词短语识别

Chinese Prepositional Phrase Recognition Based on Combined Multiple Models

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 刘彤黄德根张聪

【Author】 LIU Tong;HUANG Degen;ZHANG Cong;School of Computer Science,Dalian University of Technology;

【机构】 大连理工大学计算机学院

【摘要】 该文提出了一种多模型融合的介词短语识别方法,不仅能识别并列型介词短语,而且提高了嵌套型介词短语的识别精度。首先,利用简单名词短语识别模型识别出语料中的短语信息并进行融合,简化语料,降低介词短语内部复杂性;其次,用CRF模型识别嵌套的内层介词短语,即若存在嵌套则识别嵌套的内层,若无嵌套则识别该介词短语;最后,将初始语料中识别出来的内层介词短语进行分词融合并修改其特征信息,重新训练外层介词短语识别模型进行识别。在内外层介词短语自动识别后,利用双重错误校正系统对识别的介词短语进行校正。在2000年《人民日报》语料中的7 028个介词短语进行五倍交叉实验,结果表明,该方法识别的介词短语的正确率、召回率、F值分别为94.11%、94.02%、94.06%,比基于简单名词短语的介词短语识别方法(baseline)分别提高了1.09%、1.07%、1.08%,有效提高了介词短语识别的性能。

【Abstract】 A method of prepositional phrase recognition based on fusion of multiple models is proposed to deal with coordinate prepositional phrases and improve the performance of nested prepositional phrase recognition.First,a simple noun phrase recognition model is used to identify and merge the phrases in the corpus in order to simplify corpus and reduce internal complexity of prepositional phrases,Then,the CRF model is used to identify the inner layer of the nested prepositions phrases,i.e.if the preposition phrases is nested,recognize the inner layer,otherwise,recognize the whole preposition phrase,Finally,merge the recognized inner prepositional phrases in the corpus and modify the feature information in order to train a new model for outer prepositional phrase recognition.In addition,after the recognition of both inner and outer prepositional phrases,a double error correction system is used to correct the recognized phrases.Five-fold cross validation is conducted on the corpus of People’s Daily of 2000 including7 028 prepositional phrases,and the results achieve 94.11% in precision,94.02% in recall,and 94.06% in Fmeasure,outperformaing the baseline by 1.09%,1.07%,1.08%.

【基金】 国家自然科学基金(61672127,61672126)
  • 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2017年06期
  • 【分类号】TP391.1
  • 【被引频次】5
  • 【下载频次】127
节点文献中: