节点文献
聚类分析及其在基因表达数据中的应用研究
Cluser Analysis and Its Application in Gene Expression Data
【作者】 邓庆山;
【导师】 刘青;
【作者基本信息】 华中科技大学 , 计算机应用技术, 2004, 硕士
【摘要】 基因微阵列技术使得人们可以同时监测成千上万个基因的表达水平。目前对基因表达数据进行分析的各种方法中,聚类分析方法应用得最多。常用于基因表达数据分析的聚类方法有很多。与聚类相关的问题的有数据预处理、相似性度量、聚类有效性等。针对基因表达数据的特点,考虑到一些聚类算法的优缺点,重点介绍了两种基因表达数据聚类分析模型。一种模型是基于有效性测度谢白尼指数的基因表达数据的模糊聚类分析。它采用了一种适合于模糊聚类的聚类有效性测度谢白尼指数来衡量相同聚类个数情况下不同的聚类结果,并以谢白尼指数为标准来决定该基因表达数据集应划分为多少个聚类。将该种模型运用于公开的白血病基因表达数据集进行实验,实验表明该方法能自动获取基因表达数据的聚类数,并得到较高的分类准确率。另外一种模型是基于有效性测度的自组织映射与k均值方法相结合的一种基因表达数据聚类模型。考虑到自组织映射网络结点的聚类边界并不明显,它采用K平均值方法解决这个问题。另外,由于自组织映射聚类结果受到结点初始值和样本学习顺序的影响,每次聚类结果并不完全一致,本模型采用一种适用于普通聚类的聚类有效性测度斯路艾特指数来衡量不同的聚类结果。将该种模型运用于公开的白血病基因表达数据集和结肠基因表达数据集,取得了比较理想的实验结果。
【Abstract】 Gene microarray technique makes it possible to observe thousands of genes simultaneously. Currently, cluster methods are used most frequently among the methods applied to the analysis of gene expression data. There are lots of cluster methods applied to the analysis of gene expression data. The problems relating to cluster include data preprocessing, similarity measure and cluster validity. Considering the specialty of gene expression data and the characteristic of some cluster algorithms, two cluster models of gene expression data analysis are introduced.One model is fuzzy cluster analysis of gene expression data based on a cluster validity measure named Xie-Beni index. The model used Xie-Beni index, a validity measure applicable to fuzzy cluster, to measure the validity of different cluster results under the same cluster number and the results under different cluster numbers. We applied the model to analyze the expression data set of leukaemia. The experimental result proved that this model can get cluster numbers automatically and a high accuracy of classification.The other is cluster analysis of gene expression data associating SOM with k-means based on a cluster validity named Silhouette index. the clustering boundaries of nodes are not clear in the SOM results, this model applied k-means clustering to the results of SOM results. Besides, the clustering results are not consistent each time due to the influence of the initial value of nodes and learning order of samples. This model applied Silhouette index, a validity measure applicable to hard cluster, to measure the validity of different clustering results. We applied the model to the analysis of gene expression data of leukaemia and colon. Good experimental results were gained.
【Key words】 Cluster; Fuzzy C-means; Self-Organizing map; Cluster validity; Bioinformatics;
- 【网络出版投稿人】 华中科技大学 【网络出版年期】2005年 03期
- 【分类号】TP311.13
- 【被引频次】8
- 【下载频次】497