节点文献

一种基于语义引力及密度分布的聚类算法

A Clustering Algorithm Based on Semantic Gravitation and Density Distribution

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 李政涛夏树倩王大玲冯时张一飞

【Author】 LI Zheng-Tao~1,XIA Shu-Qian~1,WANG Da-Ling~(1,2),FENG Shi~1,ZHANG Yi-Fei~(1,2) 1 School of Information Science and Engineering,Northeastern University,Shenyang 110819 2 Key Laboratory of Medical Image Computing(Northeastern University),Ministry of Education

【机构】 东北大学信息科学与工程学院医学影像计算教育部重点实验室(东北大学)

【摘要】 由于传统的相似性度量计算方法在数据聚类、特别是高维数据聚类过程中存在的问题,基于数据重力的相似度计算方法被引入聚类过程。针对此类方法在表达类间相似关系方面存在的不足,本文提出一种新的基于语义引力及密度分布的聚类算法。一方面,将物理学中的质量和引力等概念引入到聚类分析中,将语义引力作为数据间相似性的度量方法,不但充分考虑了数据间的几何距离可分性,而且强调了数据间属性的相关性,使其对不规则分布的样本也有较好的聚类效果;另一方面,将基于划分的聚类与基于密度的聚类方法相结合并予以改进,通过对对象密度的计算,以密度较大的对象为聚类中心进行聚类,从而降低了由于初始聚类中心选择偏差造成的影响,保证了更好的精度。实验结果表明本文提出的算法具有更准确的聚类结果,特别是在文本这样的高维、稀疏的数据中更是如此。

【Abstract】 Traditional similarity measure methods have troubles in data clustering,especially in high-dimensional data clustering process,so the similarity measure method based on data gravity is introduced into clustering process.This paper proposes an innovational clustering algorithm based on semantic gravitation and density distribution for repairing lacks of the method in describing similarity between clusters.On one hand,by introducing some concepts of physics,such as mass and gravitation,into cluster analysis,this algorithm takes semantic gravitation as similarity measure method between data.This idea doesn’t only take the reparability of geometric distance between data into account,but also emphasizes the interdependency of attributes between data, making preferable clustering results available even when it comes to irregular distributed samples.On the other hand,by combining and improving partitioning methods and density-based methods,after computing objects density,this algorithm uses objects with larger density as clustering centers during the clustering process. Thereby,this approach reduces the impact caused by the selection bias of initial cluster centers and ensures better accuracy.Experiment results shows that,compared with similarity measure method based on data gravity and traditional clustering algorithm,the algorithm proposed in this paper is able to reach more accurate clustering results,especially for high-dimensional and sparse data,such as text.

【关键词】 聚类语义引力密度分布
【Key words】 ClusteringSemantic GravitationDensity Distribution
【基金】 国家自然科学基金60973019
  • 【会议录名称】 第六届全国信息检索学术会议论文集
  • 【会议名称】第六届全国信息检索学术会议
  • 【会议时间】2010-08-12
  • 【会议地点】中国黑龙江牡丹江
  • 【分类号】TP391.1
  • 【主办单位】中国中文信息学会信息检索与内容安全专业委员会
节点文献中: 

本文链接的文献网络图示:

本文的引文网络