节点文献

融合语言特性的越南语兼类词消歧

Vietnamese Multi-category Words Disambiguation Combined with Language Features

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 郭剑毅赵晨刘艳超毛存礼余正涛

【Author】 Guo Jianyi;Zhao Chen;Liu Yanchao;Mao Cunli;Yu Zhengtao;School of Information Engineering and Automation, Kunming University of Science and Technology;Yunnan Key Laboratory of Artificial Intelligence, Kunming University of Science and Technology;

【机构】 昆明理工大学信息工程与自动化学院昆明理工大学云南省人工智能重点实验室

【摘要】 兼类词歧义直接影响词性标注的准确率。本文针对越南语兼类词歧义问题提出一种融合语言特性的越南语兼类词消歧方法。通过构建越南语兼类词词典和兼类词语料库,分析越南语的语言特征和兼类词特点,选取有效的特征集;然后利用条件随机场能添加任意特征等优点,在使用词和词性上下文信息的同时,引入句法成分和指示词特征,得到消歧模型。最后在兼类词语料上实验,准确率达到了87.23%。实验表明本文所提出的越南语兼类词消歧方法有效可行,可以提高词性标注正确率。

【Abstract】 Multi-category words disambiguation directly affects the part of speech(POS)tagging accuracy.This paper proposed a statistical disambiguation method combined with linguistic characteristics of Vietnamese multi-category words. First,the paper builds Vietnamese multi-category words dictionary and Vietnamese multi-category words corpus,and selects effective feature sets for multi-category words by analyzing of Vietnamese language and multi-category words. Secondly,the paper takes into account the advantages of adding any features of CRFs model,introduces the syntactic and lexical features excepting the features of words and POS,and then builds up the disambiguation model. Finally,testing is carried out on the real multi-category category words corpus,and the accuracy is 87.23%. Experimental results show that the proposed Vietnamese multi-category words disambiguation model is effective and feasible,which can improve the correct rate of POS tagging.

【基金】 国家自然科学基金(61262041,61562052,61662041)资助项目,国家自然科学基金重点(61732005)资助项目
  • 【文献出处】 数据采集与处理 ,Journal of Data Acquisition and Processing , 编辑部邮箱 ,2019年04期
  • 【分类号】TP391.1;H44
  • 【被引频次】2
  • 【下载频次】99
节点文献中: 

本文链接的文献网络图示:

本文的引文网络