节点文献
面向不平衡数据集分类模型的优化研究
Research on Optimization of Classification Model for Imbalanced Data Set
【摘要】 为提高不平衡数据集的分类效率,建立一种分类模型,从样本采样和分类算法两方面进行优化。对决策边界的少类样本进行循环过采样生成新样本集,并与决策边界外合成的少类样本集合并,提高样本的重要度。针对传统ε-支持向量机(ε-SVM)在对不平衡数据集分类时超平面偏移的问题,引入正负惩罚系数和混合核函数,并利用客观的熵值法选取惩罚系数,提高分类算法的性能。实验结果表明,与标准的SVM算法相比,该分类模型在不平衡数据集分类上F-measure值平均提高18.1%,具有较好的分类效果。
【Abstract】 In order to improve the classification efficiency of unbalanced data sets,this paper proposes a classification model. The sample sampling and classification algorithm are optimized. A new sample set is generated by cyclic sampling of the few samples of the decision boundary,combined with the small sample sets synthesized outside the boundary of the decision-making,then the importance of the sample is improved. Aiming at the problem of hyperplane offset in classification of imbalanced data sets by traditional ε-Support Vector Machine( ε-SVM),the positive and negative penalty coefficients and the mixed kernel function are introduced. The objective entropy value method is used to select the penalty coefficients and the performance of the classification algorithm is improved. Experimental results show that compared with the standard SVM algorithm,the classification is better in the classification of imbalanced data sets,the average F-measure value is increased by 18.1%,and the better classification results are achieved.
【Key words】 text categorization; imbalanced data set; data mining; sample resampling; entropy method;
- 【文献出处】 计算机工程 ,Computer Engineering , 编辑部邮箱 ,2018年04期
- 【分类号】TP18;TP311.13
- 【被引频次】24
- 【下载频次】369