节点文献

一种改进型的不平衡数据欠采样算法

Improved Under-sampling Algorithm for Imbalanced Data

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 魏力张育平

【Author】 WEI Li;ZHANG Yu-ping;School of Computer Science and Technology,Nanjing University of Aeronautics and Astronautics;

【机构】 南京航空航天大学计算机科学与技术学院

【摘要】 不平衡数据集经常出现于很多应用领域,如果直接使用这种数据集进行分类,会对算法的学习过程造成干扰.而传统的欠采样方案会严重丢失多数类样本的信息.为解决这一问题,通过结合NearMiss算法和K-Means聚类在处理不平衡数据时的优点,提出了CBNM(Clustering-Based NearMiss)算法,该算法通过计算簇中心点的NearMiss距离,赋予该点选择权重.实验通过选择十组UCI数据集,验证在本算法中三类NearMiss算法的优劣,并与NearMiss-2算法进行比较.实验结果表明,CBNM算法在F-Measure和G-Mean上有显著提升,对分类效果的改进明显.

【Abstract】 The data in real-world applications often are imbalanced class distribution,which has attracted growing attention from both academic and industry. It will interfere with algorithm’s learning process if all the data are used to be the training data. Traditional undersampling is a controversial method in dealing with class-imbalance problem because many majority class examples are ignored. To overcome this deficiency,a Clustering-Based Near Miss(CBNM) algorithm was proposed combining the advantages of NearMiss algorithm and K-Means in processing data. CBNM gives a weight to the cluster center by calculating the Near-Miss distance. Through utilizing UCI data sets,the experiment verifies the performance of the three NearMiss algorithms in CBNM,which is compared with NearMiss-2 algorithm. The experimental results show that the CBNM algorithm has significant improvement in F-Measure and GMean,and the improvement of the classification effect is also obvious.

【关键词】 不平衡数据欠采样聚类
【Key words】 imbalanced dataunder-samplingclustering
  • 【文献出处】 小型微型计算机系统 ,Journal of Chinese Computer Systems , 编辑部邮箱 ,2019年05期
  • 【分类号】TP181;TP311.13
  • 【被引频次】45
  • 【下载频次】545
节点文献中: 

本文链接的文献网络图示:

本文的引文网络