节点文献

特征词提取中同义处理的新方法

A New Method for Synonymous Processing in Feature Word Extraction of Text Categorization

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 邹娟周经野邓成高南莎

【Author】 ZOU Juan~1,ZHOU Jing-ye~1,DENG Cheng~1,GAO Nan-sha~2(1.Information Engineering College of Xiangtan University,Xiangtan,Hunan 411105,China;2.Software institute of Dongnan University,Nanjing,Jiangsu 210000,China)

【机构】 湘潭大学信息工程学院东南大学软件学院 湖南湘潭411105湖南湘潭411105江苏南京210000

【摘要】 本文利用文本分类中文本的特点提出了一种基于模糊集的同义词处理的新方法。本方法充分考虑不同文本类型中同义(近义)词之间的差别,在训练中自动计算不同类型文本中特征词对其对应的同义概念的隶属度,从而实现了用模糊集来定义同义概念;然后应用同义概念来提取文本中的特征值。另外,本系统还利用模糊集来处理多义词的问题。文中给出了系统的处理算法。比较试验的结果表明该方法提高了分类的正确率,效果是令人满意的。整个系统达到了较高的自动化水平和较强的可移植性。

【Abstract】 A new method for synonymous processing in feature word extraction of text categorization is proposed in this paper.Fully considering the difference among synonyms in texts of different types,this method can calculate the membership degrees of feature words in their common synonymous concept automatically while training,so that we can define synonymous concepts with rough sets.Then we use synonymous concepts to extract feature values in texts.In addition,we process the polysemous problem using rough sets.The algorithms of the system are presented in the paper.And the results of the comparing tests show that our method improve the correct rates of text categorization effectively and the system is more automatic and more portable.

【基金】 湖南省自然科学基金资助项目(02JJY2092)
  • 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2005年06期
  • 【分类号】TP391.11
  • 【被引频次】29
  • 【下载频次】407
节点文献中: 

本文链接的文献网络图示:

本文的引文网络