节点文献

数据预处理和初始化方法对K-均值聚类的影响

Effects of Data Preprocessing and Intialization on K-means Clustering

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 杨春梅万柏坤丁北生

【Author】 Yang Chunmei Wan Baikun Ding Beisheng(College of Precision Instrument and Opto-Electronics Engineering, Tianjin University, Tianjin 300072, China)

【机构】 天津大学精密仪器与光电子工程学院

【摘要】 基于酵母二次迁移实验中表达谱相似的五类基因表达数据,研究了不同相似性度量准则、数据预处理方法及质心初始化方式对K-均值聚类效果的影响。结果表明:若对基因表达数据进行K-均值聚类分析,最好采用能反映数据结构特征的向量对质心进行初始化。若随机初始化质心,则采用取相对表达水平的预处理方式,以欧几里德距离(Euclidean distance)作为相似性测量准则,可以获得最佳的聚类结果;在欧氏距离准则下,标准化处理因可能破坏原始数据的幅度特征,而导致聚类结果变坏。若以Pearson相关系数为相似性准则则不同的数据预处理方式对结果无显著影响。

【Abstract】 Based on the five groups of genes expressed similarly during yeast diauxic shift, we studied the effects of different measuring metrics, data preprocessing and centroids initialization on K-means clustering. The results illustrate that the best centroids initialization in K-means clustering is to select vectors characterized the structure of the dataset. However, if the centroids are initialized randomly, clustering on the relative expression ratio under Euclidean distance metrics can obtain the best results. With Euclidean distance, normalization of the dataset only leads to worse results, for amplitude character of the dataset maybe destroyed. Meanwhile, different data preprocessing ways make unclear differences under Pearson correlation coefficient metrics.

  • 【会议录名称】 中国仪器仪表学会第五届青年学术会议论文集
  • 【会议名称】中国仪器仪表学会第五届青年学术会议
  • 【会议时间】2003
  • 【分类号】TP274.2
  • 【主办单位】中国仪器仪表学会
节点文献中: 

本文链接的文献网络图示:

本文的引文网络