节点文献

基于特征选择与集成学习的高维不平衡数据分类算法研究

Research on High-dimensional Unbalanced Data Classification Algorithm Based on Feature Selection and Ensemble Learning

【作者】 张志强;

【导师】 陈佐; 文吉刚;

【作者基本信息】 湖南大学 , 计算机技术(专业学位), 2021, 硕士

【摘要】 随着大数据的兴起,海量的非结构化数据呈现高维和不平衡特性,现有的算法无法有效处理。尽管部分研究针对高维特征或者数据的不平衡性进行了改进,但这些方法只适合应用在某些特定领域,泛化性差,如何提高模型的分类性能和泛化能力成为该领域的重点研究方向。本文结合数据的特点,从混合采样、特征选择和集成学习三个层面加以改进。(1)针对数据集中出现的重复、异常以及冗余等干扰样本,设计了一个编辑最近邻噪声清除算法ENNC(Edited Neighborhood Noise Cleaning),针对数据分布不平衡的特点,提出基于边界邻域划分和K-means聚类改进的混合采样算法HBNPK(improved Hybrid-sampling based on Boundary Neighborhood Partition and K-means clustering),通过利用基于边界邻域划分改进的过采样算法OBNP(improved Oversampling algorithm based on Boundary Neighborhood Partition)对边界最难分类的少数类样本进行过采样,合成新的少数类样本,然后采用基于Kmeans聚类改进的欠采样算法UKC(improved Under-sampling algorithm based on K-means Clustering)对远离聚类中心的多数类边缘样本进行欠采样。(2)进一步分析特征的冗余性和相关性,通过引入特征的互补性,设计了一种基于互补性的FCBF特征选择算法FFSC(FCBF Feature Selection algorithm based on Complementarity),利用MIC系数和FCBF算法过滤掉不相关特征和冗余特征,然后根据特征互补性高低,利用C4.5分类器的分类效果对特征子集进行评估,选择最优特征子集。(3)经过混合采样和特征选择之后,得到预处理后的样本,为了进一步得到分类准确率高、适应性强、鲁棒性好的分类模型,构建Stacking两层模型框架,选择支持向量机、决策树、随机森林、自适应提升作为基模型层的分类器,这些学习器之间差异性大并且单个学习器分类性能好,选择运行速度较快、分类效果极强的XGBoost算法作为元模型层分类器。最后,本文通过对比实验,结果表明混合采样算法能够大大提升稀少类样本的识别率,进一步对数据进行特征选择之后,分类准确率显著提升。而基于Stacking两层架构的多分类器融合算法在分类性能和泛化能力方面比单一模型更优越。

【Abstract】 With the rise of big data,massive unstructured data presents high-dimensional and unbalanced characteristics,and existing algorithms cannot handle it effectively.Although some researches are aimed at improving high-dimensional features or data imbalances,these methods are only suitable for application in certain specific fields and have poor generalization.How to improve the performance and generalization of the model has become a key direction in this field.This thesis combines the characteristics of the data and improves it from three levels of hybrid sampling,feature selection and ensemble learning.(1)Aiming at the repetitive,abnormal and redundant interference samples that appear in the data set,an editing neighborhood noise cleaning algorithm is designed.Aiming at the characteristics of unbalanced data distribution,an improved hybridsampling based on boundary neighborhood partition and K-means clustering algorithm is proposed.By using an improved oversampling algorithm based on boundary neighborhood partition to oversample the minority samples that are the most difficult to classify on the boundary,synthesize new minority samples,and then use an improved under-sampling algorithm based on K-means clustering to under-sample the majority edge samples that are away from cluster center.(2)Further analyze the redundancy and correlation of features.By introducing the complementarity of features,a FCBF feature selection algorithm based on complementarity is designed.The MIC coefficient and FCBF algorithm are used to filter out irrelevant and redundant features.Then according to the level of feature complementarity,use the classification effect of the C4.5 classifier to evaluate the feature subset and select the optimal feature subset.(3)After hybrid sampling and feature selection,the preprocessed sample is obtained.So as to further obtain a classification model with high classification accuracy,strong adaptability,and good robustness,a two-layer model framework of Stacking is constructed,and support vector machines,Decision trees,random forests,and adaptive boosting are used as the classifiers of the base model layer.These learners have great differences and a single learner has good classification performance.The XGBoost algorithm,which runs faster and has a strong classification effect,is selected as the meta Model layer classifier.Finally,this thesis shows through experimental results that the hybrid sampling algorithm can greatly improve the recognition rate of minority samples.After further selecting the characteristics of the data,the classification accuracy rate is significantly improved.The multi-classifier fusion algorithm based on the Stacking two-layer architecture is superior to a single model in terms of classification performance and generalization ability.

  • 【网络出版投稿人】 湖南大学
  • 【网络出版年期】2022年 09期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络