节点文献

一种基于特征重要度的文本分类特征加权方法

A Feature Weighting Scheme for Text Categorization Based on Feature Importance

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 刘赫刘大有裴志利高滢

【Author】 Liu He1,2, Liu Dayou1,2, Pei Zhili3, and Gao Ying1,2 1(College of Computer Science and Technology, Jilin University, Changchun 130012) 2(Ministry of Education Key Laboratory of Symbolic Computation and Knowledge Engineering, Jilin University, Changchun 130012) 3(College of Computer Science and Technology, Inner Mongolia University for Nationalities, Tongliao, Inner Mongolia 028043)

【机构】 吉林大学计算机科学与技术学院吉林大学符号计算与知识工程教育部重点实验室内蒙古民族大学计算机科学与技术学院

【摘要】 针对文本分类中的特征加权问题,提出了一种基于特征重要度的特征加权方法.该方法基于实数粗糙集理论,通过定义特征重要度,将特征对分类的决策信息引入到特征权重中.然后,在标准文本数据集Reuters-21578 Top10和WebKB上进行了实验.结果表明,该方法能改善样本空间的分布状态,使同类样本更加紧凑,异类样本更加松散,从而简化从样本到类别的映射关系.最后,使用Nave Bayes,kNN和SVM分类器在上述数据集上对该方法进行了实验.结果表明,该方法能提高分类的准确率、召回率和F1值.

【Abstract】 Text categorization is one of the key research fields in text mining. Feature weighting is an important problem in text categorization. For computing feature weights, a feature weighting scheme for text categorization is proposed. In this scheme, the feature importance is defined based on the real rough set theory. By this concept, decision-making information of a feature for categorization is introduced into the weight of this feature. Then, the experiments are performed on two international and standard text datasets, namely, Reuters-21578 Top10 and WebKB. Through the computation of the total within-class scatter and between-class scatter in Fisher linear discriminant, it is verified that the proposed scheme can decrease the total within-class scatter and increase the between-class scatter; that is to say, the scheme can make samples in the same class more compact and those in different classes looser for the two datasets. Thereby, the proposed scheme can improve the space distribution of samples and simplify the mapping relation from samples to classes. Finally, the proposed scheme is evaluated on the two datasets by Nave Bayes, kNN and SVM classifiers. The experimental results show that the scheme can enhance the precision, recall and the value of F1 for categorization.

【基金】 国家自然科学基金重大项目(60496321);国家自然科学基金项目(60773099,60573073);国家“八六三”高技术研究发展计划基金项目(2006AA10Z245,2006AA10A309);吉林省科技发展计划基金重大项目(20020303);吉林省科技发展计划基金项目(20030523);欧盟项目TH/Asia Link/010(111084)~~
  • 【文献出处】 计算机研究与发展 ,Journal of Computer Research and Development , 编辑部邮箱 ,2009年10期
  • 【分类号】TP391.1
  • 【被引频次】83
  • 【下载频次】1142
节点文献中: 

本文链接的文献网络图示:

本文的引文网络