节点文献
K-means聚类算法的优化研究
Optimization Research of K-means Clustering Algorithm
【摘要】 计算机网络技术日益成熟完善,大数据时代日新月异。每时每刻产生的大量数据信息也逐渐复杂,因此数据分析也越来越受到广大学者的关注。而聚类分析是数据挖掘中划分和分组的重要方法之一,在生物、医疗、机器学习、市场营销等诸多领域都得到广泛应用。K-means聚类算法作为一种可扩展性、高效性的聚类算法,在解决大规模数据聚类处理问题时有着独到的优势。基于此,该研究提出了一种针对K-means聚类算法肘点法中的先验性K值不准确的K值纠正优化算法。该算法结合了ISO-DATA算法原理,实现了在先验性K值的一定范围内根据数据集的特征利用多参数实现K值更准确的选择。
【Abstract】 Computer network technology is becoming more and more mature, and the era of big data is changing. A large amount of data information generated at all times is gradually complex, so data analysis has attracted more and more attention. Clustering analysis is one of the important methods of partitioning and grouping in data mining, which is widely used in many fields such as biology, medicine, machine learning, marketing and so on. As a scalable and efficient clustering algorithm, K-means clustering algorithm has unique advantages in solving big data clustering problems. Based on the above, this study proposes a K value correction optimization algorithm for the inaccurate prior K value in the elbow method of K-means clustering algorithm. The algorithm combines the principle of ISO-DATA algorithm, and realizes more accurate selection of K value by using multiple parameters according to the characteristics of the data set within a certain range of the prior K value.
【Key words】 cluster analysis; K-means algorithm; ISO-DATA algorithm; elbow point method; parameter selection;
- 【文献出处】 软件 ,Software , 编辑部邮箱 ,2023年10期
- 【分类号】TP311.13
- 【下载频次】59