节点文献
基于信息增益改进贝叶斯词义消歧模型
Word Sense Disambiguation Based on Bayes Model and Information Gain
【Author】 Bin Deng~1 Zhengtao Yu~(1,2)Lu Han~1 Wengang Che~(1,2)Jianyi Guo~(1,2) 1 The School of Information Engineering and Automation,Kunming University of Science and Technology,Kunming,P.R.China 650051 2 The Institute of Intelligent Information Processing,Computer Technology Application Key Laboratory of Yunnan Province,Kunming,P.R.China 650051
【机构】 昆明理工大学信息工程与自动化学院; 云南省计算机技术应用重点实验室智能信息处理研究所;
【摘要】 词义消歧是自然语言处理的关键问题。本文通过信息增益的方法,统计出歧义词上下文各个位置对岐义词词义的影响,以此为基础,选取影响岐义词前后6个位置词构建词义消歧特征向量,采用贝叶斯算法,通过信息增益为特征向量12维特征赋予不同的权重值,从而改进了贝叶斯消歧模型。采用知网义项来描叙岐义词词义,对10个汉语常用歧义词进行消歧测试实验,结果证明该方法有效,其中封闭测试正确率达95.72%,开放测试正确率达85.71%。
【Abstract】 Word sense disambiguation has always been a key problem in Natural Language Processing.In the paper, We use the method of Information Gain to calculate the weight of different position’s context,which affect to ambiguous words.And take this as the foundation.We select the ahead and back six position’s context of ambiguous words to construct the feature vectors.The feature vectors are endued with different value of weight in Bayesian Model. Thus,the Bayesian Model is improved.We use the sense of the HowNet to describe the meaning of ambiguous words. The average accuracy rate oftbe experiments of 10 Chinese ambiguous words was 95.72% in close test and the average accuracy rate was 85.71% in open test.The results showed that the method was proposed in this paper were very effective.
【Key words】 Natural Language Processing(NLP); Word Sense Disambiguation(WSD); Information Gain; weight of context position; Bayesian Model;
- 【会议录名称】 第四届全国信息检索与内容安全学术会议论文集(上)
- 【会议名称】第四届全国信息检索与内容安全学术会议
- 【会议时间】2008-11
- 【会议地点】中国北京
- 【分类号】TP391.1
- 【主办单位】中国中文信息学会信息检索与内容安全专业委员会