节点文献
基于最大熵方法的汉语词性标注
A Chinese Part of Speech Tagging Method Based on Maximum Entropy Principle
【摘要】 最大熵模型的应用研究在自然语言处理领域中受到关注 ,文中利用语料库中词性标注的上下文信息建立基于最大熵方法的汉语词性系统。研究的重点在于其特征的选取 ,因为汉语不同于其它语言 ,有其特殊性 ,所以特征的选取上与英语有差别。实验结果证明该模型是有效的 ,词性标注正确率达到 97.34%。
【Abstract】 A lot of researches have been made on the application of the maximum entropy modeling in the natural language processing during recent years. This paper presents a new Chinese part of speech tagging method based on maximum entropy principle because Chinese is quite different from many other languages. The feature selection is the key point in this system which is distinct from the one used in English. Experiment results have shown that the part of speech tagging accuracy ratio of this system is up to 97.34%.
【基金】 国家自然科学基金资助项目 (699750 0 8) ;国家 973规划资助项目 (G1 9980 30 50 7)
- 【文献出处】 计算机应用 ,Computer Applications , 编辑部邮箱 ,2004年01期
- 【分类号】TP391.1
- 【被引频次】32
- 【下载频次】500