节点文献

基于密度的kNN文本分类器训练样本裁剪方法

A Density-Based Method for Reducing the Amount of Training Data in kNN Text Classification

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 李荣陆胡运发

【Author】 LI Rong Lu and HU Yun Fa (Department of Computing and Information Technology, Fudan University, Shanghai 200433)

【机构】 复旦大学计算机与信息技术系复旦大学计算机与信息技术系 上海200433上海200433

【摘要】 随着WWW的迅猛发展 ,文本分类成为处理和组织大量文档数据的关键技术 kNN方法作为一种简单、有效、非参数的分类方法 ,在文本分类中得到广泛的应用 但是这种方法计算量大 ,而且训练样本的分布不均匀会造成分类准确率的下降 针对kNN方法存在的这两个问题 ,提出了一种基于密度的kNN分类器训练样本裁剪方法 ,这种方法不仅降低了kNN方法的计算量 ,而且使训练样本的分布密度趋于均匀 ,减少了边界点处测试样本的误判 实验结果显示 ,这种方法具有很好的性能

【Abstract】 With the rapid development of World Wide Web, text classification has become the key technology in organizing and processing large amount of document data As a simple, effective and nonparametric classification method, k NN method is widely used in document classification But k NN classifier not only has large computational demands, but also may decrease the precision of classification because of the uneven density of training data In this paper, a density based method for reducing the amount of training data is presented, which solves two problems mentioned above It not only reduces the computational demands of k NN classifier, but also makes the density of training data even and decreases the wrong classification between the edge of classes The experiment also shows that it has good performance

【关键词】 文本分类kNN快速分类
【Key words】 text classificationk nearest neighborfast classification
【基金】 国家自然科学基金项目 (60 173 0 2 7)
  • 【文献出处】 计算机研究与发展 ,Journal of Computer Research and Development , 编辑部邮箱 ,2004年04期
  • 【分类号】TP393.09
  • 【被引频次】333
  • 【下载频次】1448
节点文献中: 

本文链接的文献网络图示:

本文的引文网络