节点文献

词聚类在文本分类中的应用

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 朱慕华陈文亮朱靖波

【机构】 东北大学自然语言处理实验室

【摘要】 现有的文本分类方法需要较大的训练语料,在训练语料足够大的前提下可取得不错的效果,训练语料的规模直接影响分类的效果。然而,要大规模人工进行语料标注是一个难题。本文将k-means聚类算法引入到文本分类中,首先在无标注语料上进行词聚类,然后将聚类结果作为文本特征来代替词特征。通过这种方法,利用无标注的训练语料来改善训练语料不足的情况下文本分类的效果。实验结果表明,采用这种方法,在同等训练语料的情况下,分类性能确实有所提高。

【Abstract】 A variety of techniques for supervised learning algorithms have demonstrated reasonable performance for text categorization. The performance is affected by the size of training corpus. Creating these sets of labeled data is tedious and expensive, because labeled documents should be labeled by hand. This paper proposes an approach that we use k-means clustering algorithm for text categorization. We cluster the words from unlabeled corpus, and use these learned clusters as the features for text categorization. The experimental results show that the proposed approach can improve the performance using unlabeled corpus.

【关键词】 文本分类k-means聚类背景语料
【Key words】 text categorizationk-meansclusteringbakground corpus
【基金】 国家教育部科学技术研究重点项目(104065);国家自然科学基金和微软亚洲研究院联合(60203019)
  • 【会议录名称】 第二届全国学生计算语言学研讨会论文集
  • 【会议名称】第二届全国学生计算语言学研讨会
  • 【会议时间】2004-08
  • 【分类号】TP391.4
  • 【主办单位】中国中文信息学会
节点文献中: 

本文链接的文献网络图示:

本文的引文网络