节点文献

基于语义树的中文词语相似度计算与分析

Chinese Word Similarity Computing Based on Semantic Tree

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 张亮尹存燕陈家骏

【Author】 ZHANG Liang1,2,YIN Cunyan1,CHEN Jiajun1(1.State Key Laboratory for Novel Software Technology,Nanjing University,Nanjing,Jiangsu 210093,China;2.Jiangsu Police Institute,Nanjing,Jiangsu 210000,China)

【机构】 南京大学计算机软件新技术国家重点实验室江苏警官学院公安科技系

【摘要】 词语相似度的分析与计算是自然语言处理关键技术之一,对句法分析、机器翻译、信息检索等能提供很好的帮助。基于语义资源Hownet的中文词语相似度计算是近年来的研究热点,但大多数的研究都是对中国科学院计算技术研究所刘群提出的计算方法的改进和完善。该文充分分析和利用新版Hownet(2007)的概念架构和语义多维表达形式,从概念的主类义原、主类义原框架以及概念特性描述三个方面综合分析词语相似度,并在计算中区分语义特征相似度和句法特征相似度。实验结果理想,与人的直观判断基本一致。

【Abstract】 Word similarity analysis and computing is one of the key technologies in natural language processing.It can offer substantial help to parsing,machine translation and information retrieval etc.Recently Chinese word similarity computing based on Hownet has become a hot research issue,though most of which are improvements or modifications to what was proposed in(Liu,2002).Based on new Hownet(2007) with its concept frame and the multi-dimension semantic expression form,this paper proposes a new method to analyze and compute Chinese word similarity from three dimensions: the main sememe,the main sememe frame and the concept characteristic description.This method also distinguishes the semantic similarity and the syntax similarity in computation.Experiment shows that the method produces a good performance.

【基金】 国家863高技术发展研究计划资助项目(2006AA010109);国家自然科学基金资助项目(60673043)
  • 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2010年06期
  • 【分类号】TP391.1
  • 【被引频次】78
  • 【下载频次】1088
节点文献中: 

本文链接的文献网络图示:

本文的引文网络