节点文献

基于最大熵方法的汉语词性标注

A Chinese Part of Speech Tagging Method Based on Maximum Entropy Principle

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 林红苑春法郭树军

【Author】 LIN Hong~(1),YUAN Chun-fa~(2),GUO Shu-jun (1.Hebei Meteorological Observatory,Hebei Meteorological Bureau,Shijiazhuang Hebei 050021,China; 2.Department of Computer Science and Technology,Tsinghua University,Beijing 100084,China)

【机构】 河北省气象局省气象台清华大学计算机科学与技术系河北省气象局省气象台 河北石家庄050021北京100084河北石家庄050021

【摘要】 最大熵模型的应用研究在自然语言处理领域中受到关注 ,文中利用语料库中词性标注的上下文信息建立基于最大熵方法的汉语词性系统。研究的重点在于其特征的选取 ,因为汉语不同于其它语言 ,有其特殊性 ,所以特征的选取上与英语有差别。实验结果证明该模型是有效的 ,词性标注正确率达到 97.34%。

【Abstract】 A lot of researches have been made on the application of the maximum entropy modeling in the natural language processing during recent years. This paper presents a new Chinese part of speech tagging method based on maximum entropy principle because Chinese is quite different from many other languages. The feature selection is the key point in this system which is distinct from the one used in English. Experiment results have shown that the part of speech tagging accuracy ratio of this system is up to 97.34%.

【基金】 国家自然科学基金资助项目 (699750 0 8) ;国家 973规划资助项目 (G1 9980 30 50 7)
  • 【文献出处】 计算机应用 ,Computer Applications , 编辑部邮箱 ,2004年01期
  • 【分类号】TP391.1
  • 【被引频次】32
  • 【下载频次】500
节点文献中: