节点文献

带后缀“者”的派生词识别

Recognition of Words With The Suffix"者"

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 冯敏萱杨翠兰陈小荷

【Author】 Feng Minxuan, Yang Cuilan, Chen Xiaohe School of Chinese Language and Literature, Nanjing Normal Univ., Nanjing 210097

【机构】 南京师范大学文学院

【摘要】 我们通过对1200万字语料的统计得出,派生词约占词条总数的8.66%,构成派生词的词缀共有188个。其中,后缀“者”所构成的派生词词条数最多,构词成分最为复杂。我们采用基本词表、词例知识规则并结合词语的搭配、共现频率的混合策略对带后缀“者”的派生词进行了自动识别,封闭测试的精确率为93.06%,开放测试的精确率为82.40%。

【Abstract】 Through the statistic of corpus of 1.2 million characters, the authors concluded that derivatives take 8.66% of the total word items. The amount of affixes that can be used to form derivatives is 188. Among them, the numbers of the word items consist of the suffix "者" are the largest one, furthermore, their comprising factors are the most complicated. Using the combined policy which including the basic lexicon, the knowledge of Lexicalism, the matching of words, and the frequency of the words together appearance, the authors performed the automatic recognition of the derivatives with the suffix "者". The accuracy of the closed test is 93.06%, while the accuracy of the open test is 82.40%.

【关键词】 派生词后缀自动识别
【Key words】 derivativessuffixesautomaticrecognition
  • 【会议录名称】 全国第八届计算语言学联合学术会议(JSCL-2005)论文集
  • 【会议名称】全国第八届计算语言学联合学术会议(JSCL-2005)
  • 【会议时间】2005-08
  • 【会议地点】中国南京
  • 【分类号】TP391.43
  • 【主办单位】南京师范大学、清华大学智能技术与系统国家重点实验室
节点文献中: 

本文链接的文献网络图示:

本文的引文网络