节点文献
文本分类的特征提取方法比较与改进
Comparison and Improvments of Feature Extraction Methods for Text Categorization
【摘要】 文本的特征提取是文本分类过程中的一个重要环节,它的好坏将直接影响文本分类的准确率。该文介绍了词条的χ2统计方法(CHI)、词条与类别的互信息(MI)、信息增益(IG)、词条的期望交叉熵(CE)等文本特征提取方法,并对其取词策略进行了改进。为了对这些特征提取方法进行系统地比较,选择了三种代表性的分类器对《读卖新闻》文本数据库进行了分类实验。实验结果表明χ2统计方法具有最好的准确率,各种改进的特征提取方法都能提高文本分类的准确率。
【Abstract】 Feature extraction technology is an essential part of text categorization, which affects directly the precision of categorization. This paper introduces four popular feature extraction methods, i.e. a χ2-test (CHI), mutual information (MI), information gain (IG), and cross entropy(CE), and proposes corresponding improvements on extracting character. In order to compare these methods comprehensively, we perform simulations on Yomiuri News Corpus using three typical classification algorithms. The experimental results show that the modified feature extraction methods can improve the precision of categorization. In addition, a χ2-test method obtains the best classification precision.
【Key words】 Feature selection; Text categorization; Mutual information; Support vector machine;
- 【文献出处】 计算机仿真 ,Computer Simulation , 编辑部邮箱 ,2006年03期
- 【分类号】TP391.1
- 【被引频次】109
- 【下载频次】1311