节点文献
优化初始值的K均值中文文本聚类
K-means Chinese Document Clustering with Optimized Initial Centers
【摘要】 文本聚类是中文文本挖掘中的一种重要分析方法。K均值聚类算法是目前最为常用的文本聚类算法之一。但此算法在处理高维、稀疏数据集等问题时存在一些不足,且对初始聚类中心敏感。本文针对这些不足,提出了用特征词向量空间模型来降低向量的维数;并提出一种新的优化初始聚类中心的算法,即根据文章的特征词选择有代表性的初始聚类中心。实验表明特征词向量空间模型和优化初始聚类中心的算法能降低计算复杂度,增强结果的稳定性,并产生质量较高的聚类结果。
【Abstract】 Document clustering is an important analysis method in Chinese document mining.K-means clustering algorithm is one of the most popular document clustering algorithms.However,this algorithm has some disadvantage in processing high dimension and sparse data.Also the clustering results depend on the initial centers.This paper presents the keyword vector space model in order to reduce the dimension of the vector.A new method of optimizing initial centers is adopted.That is,to choose representative initial centers according to
- 【文献出处】 微计算机信息 ,Microcomputer Information , 编辑部邮箱 ,2009年21期
- 【分类号】TP391.1
- 【被引频次】13
- 【下载频次】338