节点文献
混合属性数据聚类算法及其应用
Clustering Algorithm for Mixed Type Data and Its Application
【作者】 张扬;
【导师】 何增有;
【作者基本信息】 大连理工大学 , 软件工程(专业学位), 2013, 硕士
【摘要】 目前已有的聚类算法大部分局限于处理连续属性或是分类属性的数据,然而在实际应用中,许多数据集是由连续属性数据和分类属性数据共同组成的,仅适用于单种数据类型的聚类算法就不能满足需求。因此,对混合了分类属性数据和连续属性数据的聚类算法的研究,具有重要的理论意义和实际价值。本文的主要研究工作包括以下几个方面:(1)首先介绍无监督离散化算法和k-ANMI聚类算法,然后提出基于一种无监督离散化的混合数据聚类算法,在UCI数据集上的实验结果表明,提出的无监督离散化的混合数据聚类算法聚类混合类型数据是非常有效的。(2)有监督离散化算法CAIM的介绍,然后提出基于有监督离散化的混合数据聚类算法,在UCI混合数据集上的实验结果表明,提出的算法优于k-prototypes算法,UCI连续数据集上的实验证明,提出的基于有监督离散化的连续数据聚类算法对比k-means算法具有更好聚类效果。(3)介绍基于质谱技术的蛋白质鉴定以及蛋白质推断问题,然后提出如何应用本文的聚类算法解决蛋白质推断问题,并给出解决方案,通过真实的蛋白质数据验证算法在蛋白质推断应用中的可行性和有效性。
【Abstract】 Until now, most of the existing clustering algorithms have been limited to deal with the data which contains either numerical attributes or categorical attributes. However, a lot of the practical databases and large datasets contain not only numerical data but also categorical ones. It’s necessary to handle both of them at the same time. Thus, it is of great theoretical and practical significance to develop a clustering algorithm which can deal with numerical data and categorical data simultaneously.The main research work of this paper can be summarized as follows:(1) It firstly introduces the unsupervised discretization algorithms and then proposes a new clustering algorithm for mixed type data which is based on unsupervised discretization algorithms. The experimental results on UCI dataset indicate that this new clustering algorithm is very effective for the mixed type data.(2) Introduce the supervised discretization algorithm, CAIM, and on the basis of CAIM, the paper proposes supervised discretization clustering algorithm. The experimental results on UCI mixed dataset show that the proposed algorithm is superior to k-prototypes algorithm. Moreover, for UCI numerical dataset, this algorithm outperforms k-means.(3) Introduce the mass spectrometry-based protein identification and protein inference problem.Then, the clustering algorithms proposed in this paper are applied to solve the protein inference problem. Through running on two proteomics datasets, the inference performance of these clustering algorithms is verified.
【Key words】 Mixed data type; Clustering; Discretization; Protein inference;
- 【网络出版投稿人】 大连理工大学 【网络出版年期】2013年 08期
- 【分类号】TP311.13
- 【被引频次】1
- 【下载频次】452