节点文献

基于遗传算法的k-means聚类挖掘方法的研究

Research of K-Means Clustering in Data Mining Based on Genetic Algorithm

【作者】 赵艳丽

【导师】 魏权利;

【作者基本信息】 青岛科技大学 , 控制理论与控制工程, 2009, 硕士

【摘要】 数据挖掘是随着信息技术不断发展而形成的一门新学科,是信息处理和数据库技术领域的一个新兴的研究热点。数据挖掘的任务是从海量数据中发现隐含的有用知识,为科学决策提供支持。聚类分析是数据挖掘的一个非常重要的研究分支。聚类是一种无监督的分类方法,目标是在没有任何先验知识的情况下,将数据集划分成不同的类,使得相同类中的对象尽可能相似,不同类中的对象尽可能相异。k-means算法作为聚类分析中的经典算法现已被广泛应用在商务、市场分析、生物学、文本分类等领域。然而,k-means算法具有对初始值敏感、易陷入局部极小值等缺点。针对这些缺陷,本文结合遗传算法的思想,提出了一种基于遗传算法和k-means算法的混合聚类方法,并通过仿真实验验证算法的有效性。本文工作主要体现在以下几个方面:首先,详细介绍了聚类分析技术,对现有的聚类算法进行了分类,分析了这些算法的优缺点,并在此基础上,重点研究了k-means算法。其次,全面介绍了数据挖掘中的一个重要算法——遗传算法。对遗传算法的特点、基本要素、工作流程等进行了详细描述。再次,基于遗传算法和k-means算法的特点,提出了一种改进的遗传k-means聚类算法,并从编码方法、适应度函数的构造、选择算子、交叉算子和变异算子的设计、k-means优化操作等方面对提出的算法进行了详细描述。最后,为了测试本文提出的聚类算法的性能,本文用k-means算法和改进的算法进行了三组实验,并对两种算法的聚类结果进行比较,实验结果表明本文算法能够有效地解决聚类问题。

【Abstract】 Data mining is a new subject formed with the development of the information technology and is a new research point in the information and database technology. The purpose of data mining is to discovery hidden and useful knowledge which can support the science decision from huge amounts of data.Cluster analysis is one of the important themes in data mining. Clustering is an unsupervised classifying method, the goal of clustering is to partition data set into such clusters that objects within a cluster have high similarity in comparison to one another, but are very dissimilar to objects in other clusters without any prior knowledge. As a classical method of clustering analysis, k-means has been widely used in commerce, market analysis, biology, text classification and so on. However k-means has two severe defects—sensitive to initial data and easy to get into a local optimum. On this condition, combining the idea of genetic algorithm, a hybrid algorithm of clustering which is based on genetic algorithm and k-means algorithm is proposed, the performance of the hybrid algorithm is tested.The main research work of the paper includes:Firstly, clustering analysis technology is introduced in details, most existing clustering algorithms are classified, and their advantages and disadvantages are analyzed. On this basis, k-means method is choosen as research target.Secondly, an important method—genetic algorithm in data mining is introduced, and the characteristic, basic element, applied flow of it are described in details.Thirdly, based on the characteristics of genetic algorithm and k-means method, a new clustering method of k-means based on improved genetic algorithm is proposed. The proposed algorithm is described in details from coding method, fitness function, selection operators, crossover operators, mutation operators, k-means operators and other aspects.Finally, for testing the performance of the proposed algorithm, the paper gives three simulation experiments. Simulation results show that comparing with k-means method, the proposed algorithm can get a better clustering result.

  • 【分类号】TP311.13
  • 【被引频次】30
  • 【下载频次】1097
  • 攻读期成果
节点文献中: 

本文链接的文献网络图示:

本文的引文网络