节点文献

改进K均值聚类的不平衡数据欠采样算法

Improved Unbalanced Data Undersampling Algorithm For K-means Clustering

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 于艳丽江开忠王珂盛静文

【Author】 YU Yan-li;JIAN Kai-zhong;WANG Ke;SHENG Jing-wen;School of Mathematics,Physics & Statistics,Shanghai University of Engineering Science;School of Electrical and Electronic Engineering,Shanghai University of Engineering Science;

【机构】 上海工程技术大学数理与统计学院上海工程技术大学电子与电气工程学院

【摘要】 传统欠采样方法在处理不平衡数据问题时只考虑多数类样本的绝对位置而忽略了其相对位置,从而使产生的平衡数据集存在边界模糊问题。提出一种改进K均值聚类的不平衡数据欠采样算法(UD-PK)。该算法首先利用改进的PSO算法迭代寻找全局最优解作为K-means聚类所需初始值,然后通过K-means进行聚类,再按照每个类别中多数类与少数类的比例定义所取多数类样本个数,并根据多数类样本与簇心距离择优选择参与平衡数据集构造。在UCI数据集上的对比试验表明,该算法在少数类准确率上较一些经典算法有很大提升。

【Abstract】 The traditional undersampling method only considers the problem that the absolute position of most class samples ignores its relative position when dealing with the unbalanced data problem,so that the resulting balanced data set has boundary blurring problems. This paper proposes an improved unbalanced data undersampling algorithm for K-means clustering(UD-PK). The algorithm first uses the improved PSO algorithm to iteratively find the global optimal solution as the initial value needed for K-means clustering,and clusters by K-means;then according to the ratio of most classes to minority classes in each category the number of samples taken from the majority of the class is defined to participate in the construction of the balanced data set according to the selection of the majority class sample and the cluster center distance. The comparison experiments on the UCI dataset show that the proposed algorithm has a great improvement in the accuracy of a few classes compared with some classical algorithms.

  • 【分类号】TP181
  • 【被引频次】11
  • 【下载频次】212
节点文献中: 

本文链接的文献网络图示:

本文的引文网络