节点文献
基于马太效应的聚类分析方法的研究
Research on Merton Effect-based Clustering Analysis Method
【作者】 肖文;
【导师】 鞠时光;
【作者基本信息】 江苏大学 , 计算机应用技术, 2007, 硕士
【摘要】 随着计算机软硬件技术的发展和应用水平的提高,人类社会产生数据和获取数据的能力迅速增长,导致我们淹没在数据的汪洋大海中却饥渴于知识。人们迫切需要一种能够自动地将数据快速转换成知识的技术和工具,于是数据挖掘技术应运而生。聚类分析是数据挖掘研究领域中一个非常活跃的研究课题,是一种搜索簇的无监督学习过程,它的应用极为广泛。目前已开发出基于划分、层次、密度、网格和模型等多种聚类方法,其中最基本也是最重要的是凝聚层次聚类法,研究表明凝聚层次聚类算法能够产生高质量的簇。本文研究凝聚层次聚类算法,做了如下主要工作:首先介绍当前主要聚类算法及存在的问题。针对聚类分析中数据类型复杂多样,在对各种类型数据间的邻近性度量研究后,提出混合类型变量邻近性度量方法,对所有类型变量一次处理。它考虑了数据的标准化、变量加权、非对称属性和属性值遗漏等情况。针对目前常用簇间邻性度量方法存在的不足,提出基于马太效应的MEICD(Merton Effect-based Inter-cluster Distance)距离作为簇间邻近性度量。实验结果表明将簇间邻近度看成簇间距离以及簇大小等因素的多元的函数,可以提高聚类质量。针对目前层次聚类算法中对簇个数的设定存在困难,特别对于包含高维对象的数据集更是如此,设计出MHCA(MEICD-based Hierarchical ClusteringAlgorithm)层次聚类算法。它利用聚类过程中得到的合并向量和描述函数,以可视化的方法,从全局的观点识别出自然簇的个数,不需要额外的外部参数。该算法能处理混合类型变量、处理任意形状和大小的簇,对具有噪声的数据集也能得到较好的结果,并且具有较好的可解释性。最后将MHCA聚类算法应用到长江电气集团的电子商务智能决策支持系统中,在原有系统中插入了客户聚类模块。基于客户的购买心理有一种从众现象,将点击流数据与后台内部数据结合起来进行智能分析,实现了对客户的聚类,对决策者和客户具有指导意义。
【Abstract】 With the development and application of Computer hardware and software, the ability of human society to produce and obtain data is rapidly growing. So we are drowning in data yet starving for knowledge. People urgently need tools to convert data into knowledge automatically and rapidly. Then data mining came into being.Clustring analysis is an active research project in dada mining, which is an unsupervised learning and has been widely used. There have been many methods proposed for clustering, such as division method, hierarchical method, density based method, grid based method and model based method, but the most basic and important one is the agglomerative hierarchical clustering method, which has been proved by lots of study that can generate high-quality cluster.As for the study of agglomerative hierarchical clustering method, following work is accomplished.Firstly the present main clustering algorithm and their weakness are introduced. After studying the proximity measure for different type data, a mixed-type data’s proximity measure method is proposed. It considers the standardization of data, variable’s weight, asymmetrical attribute and attribute’s value omission and so on.For the drawback of the present inter-cluster proximity measure, MEICD (Merton Effect - based Inter-cluster Distance) is proposed. Experiment showed that the inter-cluster proximity measure which considers the inter-cluster distance and cluster size can improve the clustering quality.For the difficulty to set the cluster number of given data set, particularly hard for high-dimensional data set, MHCA (MEICD-based Hierarchical Clustering Algorithm) is designed. It identifies the natural clusters visually and globally by using the association vector and descriptive function obtained in the clustering process, without resorting to external parameters. This algorithm can deal with mixed-attribute data and can recognize clusters with arbitrary shape and size even if there are outliers.Lastly the MHCA is applied to the e-business intelligent decision support system of Changjiang electronic Group Corp. A clustering module is inserted into the system. Based on the customer’s herd purchase psychological phenomena, click-stream data and inner data of corporations is used to cluster the client. The clustering result can be used to guide the decision-maker and the client.
【Key words】 unsupervised learning; hierarchical clustering; proximity measure; Merton Effect; association vector; descriptive function;
- 【网络出版投稿人】 江苏大学 【网络出版年期】2008年 08期
- 【分类号】TP311.13
- 【被引频次】1
- 【下载频次】276