节点文献

文本分类的特征提取方法比较与改进

Comparison and Improvments of Feature Extraction Methods for Text Categorization

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 申红吕宝粮内山将夫井佐原均

【Author】 SHEN Hong~1, LU Bao-liang~1, Utiyama Masao~2, Isahara Hitoshi~2(1.Department of Computer Science and Engineering, Shanghai Jiaotong University, Shanghai 200030, China;2.Computational Linguistics Group, National Institute of Information and Communications Technology, Kyoto 610-0289, Japan)

【机构】 上海交通大学计算机科学与工程系国立信息与通讯技术研究所计算语言实验室国立信息与通讯技术研究所计算语言实验室 上海200030上海200030日本京都府619-0289

【摘要】 文本的特征提取是文本分类过程中的一个重要环节,它的好坏将直接影响文本分类的准确率。该文介绍了词条的χ2统计方法(CHI)、词条与类别的互信息(MI)、信息增益(IG)、词条的期望交叉熵(CE)等文本特征提取方法,并对其取词策略进行了改进。为了对这些特征提取方法进行系统地比较,选择了三种代表性的分类器对《读卖新闻》文本数据库进行了分类实验。实验结果表明χ2统计方法具有最好的准确率,各种改进的特征提取方法都能提高文本分类的准确率。

【Abstract】 Feature extraction technology is an essential part of text categorization, which affects directly the precision of categorization. This paper introduces four popular feature extraction methods, i.e. a χ2-test (CHI), mutual information (MI), information gain (IG), and cross entropy(CE), and proposes corresponding improvements on extracting character. In order to compare these methods comprehensively, we perform simulations on Yomiuri News Corpus using three typical classification algorithms. The experimental results show that the modified feature extraction methods can improve the precision of categorization. In addition, a χ2-test method obtains the best classification precision.

【基金】 国家自然科学基金资助(60375022)
  • 【文献出处】 计算机仿真 ,Computer Simulation , 编辑部邮箱 ,2006年03期
  • 【分类号】TP391.1
  • 【被引频次】109
  • 【下载频次】1311
节点文献中: 

本文链接的文献网络图示:

本文的引文网络