节点文献
基于遗传算法的聚类挖掘研究
Research of Clustering Method in Data Mining Based on Genetic Algorithm
【作者】 苏守宝;
【导师】 倪志伟;
【作者基本信息】 安徽大学 , 计算机及应用, 2004, 硕士
【摘要】 聚类分析也是数据挖掘研究中的一个非常活跃的研究课题。聚类就是把一个没有类别标记的样本集按某种准则划分成若干类,使类内样本的相似性尽可能大,而类间样本相似性尽量小,是一种无监督的分类方法。聚类分析虽已广泛地应用于模式识别、数据挖掘、计算机视觉、模糊控制、决策分析和预测等许多领域,但它在理论和方法上仍不完善,甚至还有严重的不足之处。对聚类算法的进一步优化研究将不仅有助于算法理论的完善,更有助于算法的推广和应用。 该课题研究的目标就是在数据挖掘的背景下,从理论、算法和应用的角度对聚类分析技术在如下三个方面进行了探索性的研究并取得了一定的成果。 1) 现有聚类算法的分类研究。从聚类准则、聚类的表示、算法框架等不同角度来考察并区分这些算法,然后从混合聚类方法、增量聚类、自动化和可视化等技术方面对现有算法加以比较分析评价。分析了现有算法的优缺点,以利于进一步改进;通过对现有算法的性能评述,有利于数据挖掘用户能够针对特定的数据集选择正确的算法,以获得最优化的结果和性能;也可以为现有算法分类比较的进一步研究以及建立聚类基准奠定基础。 2) 聚类分析的遗传算法研究。传统的基于聚类准则的聚类算法本质上是一种局部搜索算法,它们采用了一种迭代的爬山技术来寻找最优解,存在着对初始化敏感和容易陷入局部极小的致命缺点。遗传算法(GA)是一种通过模拟自然进化过程搜索最优解的方法,其显著特点是隐含并行性和对全局信息的有效利用能力。文中讨论了聚类分析的遗传操作改进方法,首次提出了基于佳点集GA的聚类算法GAmeans,降低了传统聚类算法对初始化的要求,具有收敛快、较强的稳健性和可避免早熟的特点。提出了一种混合聚类方法HgaMeans,实验比较测试表明它具有更好的聚类质量和综合性能。 3) 增量算法探讨。增量式聚类方法有适应大规模、动态数据、降低内存需求、可实现并行处理和增量更新等诸多优点,而且时空复杂性较小。增量算法的要求是聚类特征一般是可加的、非迭代的,该文提出了一种基于密度的网格聚类算法GDCLUS,并在此基础上提出了增量式算法IGDCLUS。该算法可发现任意形状的聚类,适用于数据的批量更新,具有高的效率且容易实现。目前的增量算法尚未完全解决对数据顺序的敏感性,高效、自适应性和交互性地动态地增量式聚类算法有待进一步研究,数据挖掘中的聚类技术仍面临着许多问题和挑战。
【Abstract】 Clustering analysis is one of most heated research topic of the day. Data clustering, a unsupervised classifying method, is the process of grouping together similar multi-dimensional data vectors into a number of clusters or bins. Clustering technique have been applied to a wide range of problems, including pattern cognition, data mining, decision-analyzing and prediction, etc., yet it is imperfect both theoretically and methodologically, even severe fault. Optimizing deeply clustering algorithms will not only help to perfect its theory, but also help to its popularization and application.This thesis aimed at studying following three aspects of clustering analysis from its theory, algorithms and applications in data mining.Firstly, classification of popular clustering algorithms is studied. Most existing clustering algorithms are classified and inter-compared from three different viewpoints, namely clustering criteria, cluster representation, and algorithm framework, and analysed and evaluated with hybrid methods, incremental algorithms, automation and visualization. It can make for existing algorithms to be improved by analysing their advantages and disadvantages, and for users to choose a right algorithm for a specified dataset in order to receive a optimization clustering results. It is also the basis of further classifying popular algorithm and establishment of clustering benchmark.Secondly, genetic algorithm(GA)-based clustering method is researched. Conventional clustering criteria-based algorithms is a kind of local search method by using iterative mountain climbing technique to find optimization solution, which has two severe defects-sensitive to initial data and easy as can get into local minimum. GA is a computational models of the human evolution, with implicit parallelism and capacity of using effectively global information. This thesis presented a modified genetic operators in clustering analysis, and firstly introduced good point set-based clustering algorithm-GAmeans, which characterized by inferior sensitivity to initial, robustness, and removable premature, and also firstly presented a hybrid method with GA and GAmeans. Experiment show that the hybrid method with general performances can find better clustering results.Finally, this thesis explored incremental algorithm, which featured normally in addable and non-iterative with some advantages, such as applicable to large and dynamic database, lower demand for memory, implementation of parallel processing and incremental update. This paper introduced an incremental grid density-based clustering algorithm-IGDCLUS, which can find high effectively arbitrary shape clusters, and is applicable in periodically incremental environment. However, existing algorithms is still sensitive to data order. Higheffective, self-adaptive, interactively dynamic, incremental clustering algorithm should be studied. Clustering technique in data mining will yet be faced with many problems and challenges.
【Key words】 data mining; genetic algorithm; clustering; incremental algorithm; good-point sets.;
- 【网络出版投稿人】 安徽大学 【网络出版年期】2004年 03期
- 【分类号】TP311.13
- 【被引频次】8
- 【下载频次】702