节点文献

基于自适应聚类过采样的软件缺陷预测研究

Software Defect Prediction Based on Adaptive Clustering Oversampling

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 贾燕华李英梅

【Author】 Jia Yanhua;Li Yingmei;Harbin Normal University;

【通讯作者】 李英梅;

【机构】 哈尔滨师范大学

【摘要】 软件缺陷预测是软件质量保障领域的热点研究课题,缺陷预测模型的性能与训练数据质量有着密切关系.针对软件缺陷预测中数据类不平衡问题,该文提出一种结合局部密度和K-Means++聚类的自适应判断过采样方法(local density adaptive oversampling based on K-Means++, LDKMAS).该方法首先利用K-Means++聚类算法为少数类样本聚类,获得多个子簇;其次计算各子簇中样本的局部密度,并合计为子簇密度;最后根据子簇密度自适应确定各子簇的过采样量,插值合成新样本直至数据集平衡.将LDKMAS算法与其他经典的过采样方法进行对比实验,用不同指标评价预测效果.实验表明,该文算法的软件缺陷预测效果更为出色,展现了较之于其他采样方法在软件缺陷预测不平衡数据处理上的优势.

【Abstract】 Software defect prediction is a hot research topic in the field of software quality assurance. The performance of defect prediction models is closely related to training data quality. Aiming at the problem of data class imbalance in software defect prediction, in this paper, a local density adaptive oversampling based on K-Means++, LDKMAS is proposed. Firstly, K-Means++ algorithm is used to cluster defective samples. Secondly, the local density of samples in each sub cluster is calculated, and the sub cluster density is merged; Finally, the oversampling amount of each sub cluster is adaptively determined according to the sub cluster density, then new samples are interpolated until the whole data set is balanced. Compared with original data set and other class imbalance oversampling methods, the experiment shows that the software defect prediction effect of this algorithm is better on two evaluation indicators, showing the advantages of unbalanced data processing in software defect prediction.

【基金】 黑龙江省省属高校基本科研业务费科研项目(2020-KYYWF-0358);黑龙江省教育教学改革项目(SJGY20210457);哈尔滨师范大学教改重点项目(XJGZ2021001)
  • 【文献出处】 哈尔滨师范大学自然科学学报 ,Natural Science Journal of Harbin Normal University , 编辑部邮箱 ,2023年02期
  • 【分类号】TP311.53;TP311.13
  • 【下载频次】15
节点文献中: 

本文链接的文献网络图示:

本文的引文网络