节点文献
基于最优密度估计的密度峰值聚类算法
Density peaks clustering algorithm based on optimal density estimation
【摘要】 针对密度峰值聚类算法(clustering by fast search and find of density peaks,DPC)聚类无特定形状的实际数据集时聚类精度欠佳的问题,提出一种最优化密度估计的密度峰聚值类算法。使用最优Oracle逼近(Oracle approximating shrinkage,AS)计算出最优协方差矩阵,利用最优协方差矩阵构造马氏距离,通过最优协方差矩阵提高DPC对数据相似度的区分能力,在此基础上结合K近邻算法,实现数据样本密度最优估计,利用最优密度估计提高DPC对实际数据集的聚类精度。在人工数据集和UCI真实数据集上进行仿真实验,实验结果表明,改进DPC算法的思路是可行的。
【Abstract】 To address the issue of bad performance on non-specific shape real world datasets of clustering by fast search and find of density peak(DPC)algorithm,a method based on optimal density estimation was proposed.The Oracle approximating shrinkage(OAS)was used to estimate the covariance matrix which was used to calculate the Mahalanobis distance,and the similarity distinguishing ability was improved using the covariance matrix.The real world datasets clustering accuracy was improved by optimal density estimation calculated based on the K-nearest neighbor algorithm.The results of experiment on artificial datasets and UCI datasets indicate that the idea of the proposed algorithm is feasible.
【Key words】 clustering by fast search and find of density peaks; K-nearest neighbor; covariance matrix; Oracle approximating shrinkage; optimal density estimate;
- 【文献出处】 计算机工程与设计 ,Computer Engineering and Design , 编辑部邮箱 ,2020年07期
- 【分类号】TP311.13
- 【被引频次】7
- 【下载频次】649