节点文献
一种新的基于层次和K-means方法的聚类算法
A Novel Clustering Algorithm Based on Hierarchical and K-means Clustering
【Author】 Li Wenchao,Zhou Yong,Xia Shixiong School of Computer Science & Technology,China University of Mining & Technology,Xuzhou,Jiangsu 221008,P.R.China
【机构】 中国矿业大学计算机科学与技术学院;
【摘要】 业界提出的传统H-K(Hierarchical K-means)聚类算法虽有效解决了K-means算法初始化中心选择的经验性和随机性,但昂贵的计算复杂度使其被难以广泛应用。本文提出一种新的基于层次和K-means的聚类算法,具有较优的计算复杂度。首先引入轮廓系数的概念,从而确定事先未知分类信息的数据集中所包含的最优聚类数Kopt;然后通过凝聚层次聚类的方法获得数据集的分布,确定初始聚类中心;最后利用K-means方法完成聚类。IRIS测试数据集的实验结果验证了该算法的有效性。
【Abstract】 Although the apriority and randomicity to initiate clustering centers of K-means have been solved by traditional hi-erarchical k-means clustering algorithm,the algorithm is difficult to be applied widespread popularly owing to its high com-putational complexity.So a novel clustering algorithm based on hierarchical and K-means clustering,which has good compu-tational complexity,is proposed in this paper.Firstly,the concept of silhouette coefficient is introduced and the optimal clus-tering number Kopt included in data set of unknown class information is decided.Then the distribution of data set is gotten through hierarchical clustering and clustering center is decided.Finally,the clustering is completed through K-means cluster-ing.The efficiencies of the algorithm is validated through the test of IRIS testing data set.
【Key words】 Clustering; K-means; Hierarchical clustering; Silhouette coefficient;
- 【会议录名称】 第二十六届中国控制会议论文集
- 【会议名称】第二十六届中国控制会议
- 【会议时间】2007-07-26
- 【会议地点】中国湖南张家界
- 【分类号】TP311.13
- 【主办单位】中国自动化学会控制理论专业委员会(Technical Committee on Control Theory,Chinese Association of Automation)