节点文献
统计和规则相结合的中文机构名称识别
Identification of Chinese Organization Names Based on Statistics and Rules
【Author】 Zhang Yanli Huang Degen Zhang Lijing Yang Yuansheng (Department of Computer Science and Technology, Dalian University of Technology 116024)
【机构】 大连理工大学计算机系;
【摘要】 中文机构名称是专名的一种,量大且层出不穷,因而大多不能收入词典,这便给自然语言处理,尤其是机器翻译和机器理解带来很大困扰.本文将统计和规则两种方法结合起来,建立了中文机构名称的识别模型.系统闭式精确率和召回率分别达92.5%和92%,开式精确率和召回率分别达88.5%和76.6%.
【Abstract】 Chinese organization names are one kind of proper noun, most of them can’t be stored in the dictionary. Identification of Chinese organization names is very important to improve the accuracy of automatic word segmentation and the intelligibility of machine translation. In this paper, we present one model based on statistics and rules, in which we use the conception of the reliability for the word segment and some appropriate rules to identify Chinese organization names. The preliminary experiment shows that the precision and recall rate respectively reach 92.5% and 92% by close test, while the precision and recall rate are 88.5% and 76.6% by open test.
【Key words】 Chinese organization names; Uni-gram Frequency; Bi-gram Frequency;
- 【会议录名称】 自然语言理解与机器翻译——全国第六届计算语言学联合学术会议论文集
- 【会议名称】全国第六届计算语言学联合学术会议
- 【会议时间】2001-08
- 【会议地点】中国山西
- 【分类号】H085
- 【主办单位】山西大学计算机系