节点文献

基于BDE-MICI的无监督特征选择的研究

Research on Unsupervised Feature Selection Based on BDE-MICI

【作者】 朱强;

【导师】 陈进源;

【作者基本信息】 兰州大学 , 应用统计(专业学位), 2019, 硕士

【摘要】 随着大数据时代的到来,很多高维度和没有标签的数据已经大量出现在如今的现实生活中,比如医疗、金融等领域产生的数据。人们在处理这些数据的时候,发现并不是所有的特征都是必要的。对于大数据集来说,我们会发现一些特征是冗余的特征或者与其它特征是高度相关的。所以,我们对数据进行预处理时,往往会去除这些冗余的特征和嘈杂的特征,这对后面的进一步学习是不可缺少的。基于原始集的样本有没有类别标签,特征选择不妨被划分成监督的和无监督的这两类方法。然而在现实中的数据多半是没有带有类别标签的。无监督特征选择算法的研究和应用成为了如今的一个热点研究问题,在对无标签数据的处理上体现了它无法代替的重要位置。本文对无监督特征选择问题进行了研究和分析。本文利用二进制微分进化算法和最大信息压缩指数的原理,提出了一种基于二进制微分进化与最大信息压缩指数的无监督特征选择算法。该算法利用最大信息压缩指数的性质所构造的适应函数作为候选子集的评价准则,该适应函数在特征选择中用于减少冗余性特征和不相关性特征。通过对现有微分进化算法的对比与分析,将实数编码方式改为0-1编码方式,使得二进制微分进化算法既具有微分进化的优化速度,又在特征选择上操作简单。我们在二进制微分进化中引入了自调节变异算子,避免了早熟现象,增加了搜索到全局最佳解的概率。通过在七个不同类型的数据集进行性能比较分析,得出的结果证明所改进的算法优于其他现有的四种无监督特征选择算法以及GA-MICI算法。

【Abstract】 With the advent of the age of big data,many high-dimensional and unlabeled data have appeared in today’s real life,such as data generated in the medical and financial fields.For big data sets,we will find that some features are redundant features or are highly related to other features.Therefore,when we preprocess the data,we often remove these redundant features and noisy features.This is indispensable for further learning.Based on whether the sample in the original set contains class label,feature selection can be divided into two methods: supervised and unsupervised.Most of the data in reality is without class label.The research and application of unsupervised feature selection algorithm has become a hot research issue today,and it is an important position that can’t be replaced in the processing of unlabeled data.This thesis makes an research and analysis on the unsupervised feature selection problem.In this thesis,based on the principle of binary differential evolution algorithm and maximum information compression index,an unsupervised feature selection algorithm based on binary differential evolution and maximum information compression index is proposed.The algorithm utilizes the adaptive function constructed by the properties of the maximum information compression index as the evaluation criterion of the candidate subset,which is used to reduce the redundancy feature and the irrelevance feature in feature selection.By comparing and analyzing the existing differential evolution algorithms,the real number coding method is changed to 0-1 coding mode,which makes the binary differential evolution algorithm have both the optimization speed of differential evolution and the simple operation of feature selection.We introduce a self-regulating mutation operator in binary differential evolution,avoiding the premature phenomenon and increasing the probability of searching for the global optimal solution.Through performance comparison analysis in seven different types of data sets,the results show that the improved algorithm is superior to other existing four unsupervised feature selection algorithms and GA-MICI algorithm.

  • 【网络出版投稿人】 兰州大学
  • 【网络出版年期】2019年 08期
  • 【分类号】C81
  • 【下载频次】49
节点文献中: