节点文献
基于k-近似的汉语词类自动判定
Part-of-Speech Identification for Unknown Chinese Words Based on k-Nearest Neighbors Strategy
【摘要】 生词处理在面向大规模真实文本的自然语言处理各项应用中占有重要位置 .词类自动判定就是对词类未知的生词由机器自动赋予一个合适的词类标记 .文中提出了一种基于 k-近似的词类自动判定算法 ,并在一个 1亿字汉语语料库及一个 6 0万字经过人工分词和词类标注的汉语熟语料库的支持下 ,构造了相应实验 .实验结果初步显示 ,本算法对汉语开放词类——名词、动词、形容词的词类自动判定平均正确率分别为 99.2 1%、84.73%、70 .6 7% ,基本上能够满足工程实现的需要
【Abstract】 Unknown word processing plays an important role in many natural language application systems aiming at large scale unrestricted texts. The task of part of speech identification is to automatically assign a part of speech tag to an unknown word with empty part of speech information. A part of speech identification algorithm based on k- nearest neighbors strategy is presented in this paper. The preliminary experiment, supported by a Chinese corpus of 100M characters and a part of speech annotated corpus of 0.6M characters, shows that the average accuracy rates of the algorithm can reach 99.21%, 84.73%, 70.67% for Chinese words of nouns, verbs and adjectives respectively.
【Key words】 part of speech identification; unknown word processing; Chinese information processing; natural language processing; artificial intelligence;
- 【文献出处】 计算机学报 ,CHINESE JOURNAL OF COMPUTERS , 编辑部邮箱 ,2000年02期
- 【分类号】TP391.43
- 【被引频次】26
- 【下载频次】338