节点文献

基于k-近似的汉语词类自动判定

Part-of-Speech Identification for Unknown Chinese Words Based on k-Nearest Neighbors Strategy

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 孙茂松左正平邹嘉彦

【Author】 SUN Mao Song 1) ZUO Zheng Ping 1) TSOU B K 2) 1) (The State Key Laboratory of Intelligent Technology & Systems, Tsinghua University, Beijing 100084) 2) (Language Information Sciences Research Center, City University o

【机构】 清华大学智能技术与系统国家重点实验室!北京100084香港城市大学语言资讯科学研究中心!香港

【摘要】 生词处理在面向大规模真实文本的自然语言处理各项应用中占有重要位置 .词类自动判定就是对词类未知的生词由机器自动赋予一个合适的词类标记 .文中提出了一种基于 k-近似的词类自动判定算法 ,并在一个 1亿字汉语语料库及一个 6 0万字经过人工分词和词类标注的汉语熟语料库的支持下 ,构造了相应实验 .实验结果初步显示 ,本算法对汉语开放词类——名词、动词、形容词的词类自动判定平均正确率分别为 99.2 1%、84.73%、70 .6 7% ,基本上能够满足工程实现的需要

【Abstract】 Unknown word processing plays an important role in many natural language application systems aiming at large scale unrestricted texts. The task of part of speech identification is to automatically assign a part of speech tag to an unknown word with empty part of speech information. A part of speech identification algorithm based on k- nearest neighbors strategy is presented in this paper. The preliminary experiment, supported by a Chinese corpus of 100M characters and a part of speech annotated corpus of 0.6M characters, shows that the average accuracy rates of the algorithm can reach 99.21%, 84.73%, 70.67% for Chinese words of nouns, verbs and adjectives respectively.

【基金】 国家自然科学基金!( 6970 5 0 0 5 )
  • 【文献出处】 计算机学报 ,CHINESE JOURNAL OF COMPUTERS , 编辑部邮箱 ,2000年02期
  • 【分类号】TP391.43
  • 【被引频次】26
  • 【下载频次】338
节点文献中: 

本文链接的文献网络图示:

本文的引文网络