节点文献

基于模糊熵的多标记和弱标记特征选择算法的研究

Feature Selection Research for Multi-label and Weak-label Based on Fuzzy Entroy

【作者】 胡虎

【导师】 代建华;

【作者基本信息】 浙江大学 , 计算机技术, 2018, 硕士

【摘要】 在计算机和互联网飞速发展的今天,不仅数据量变得越来越大,数据的形式也变的越来复杂,对数据的智能处理也变的尤为重要。其中模式识别,数据挖掘,机器学习、深度学习等已经成为了处理和挖掘数据的主要方式。面对越来越复杂多样的数据,数据降维扮演着越来越重要的角色。多标记数据是传统单标记数据的一种延伸,数据形式变得复杂,同时在多标记中标记的缺失也是一个重要的问题,因此针对多标记和弱标记进行特征选择也变得重要起来。传统的粗糙集理论是一种有效的处理不确定性的工具,在特征选择中有着广泛的应用,但是其只能处理离散型数据。之后模糊粗糙集的出现解决这一问题,并扩展出了模糊信息论,其中有模糊熵和模糊互信息等。基于传统信息论的特征选择算法已经得到了很多的研究,但是基于模糊信息论的多标记特征选择算法却比较少,其可以直接处理数值型数据以及混合型数据。因此这里我们将利用模糊粗糙集来处理多标记数据中的特征选择问题。同时针对多标记数据的特点,提出了新的多标记特征选择算法。同时多标记中的标记缺失是一种常见的问题,针对这一情况,结合了粗糙集中不完备信息系统中缺失值的处理方式,多标记中标记缺失情况下的特征选择问题也得到了处理。针对上面提出的问题,本文提出了结合模糊粗糙集,多标记数据和不完备信息系统的多标记特征选择算法和弱标记特征选择算法。本文的主要成果如下:·基于模糊信息论和标记之间的相关性,提出了多标记场景下特征选择算法,并通过实验进行了效果分析。·基于特征相关性,提出了可以去除冗余特征的多标记特征选择算法,并通过实验进行了效果分析。·结合不完备信息系统中处理标记缺失的方式,提出了标记缺失下的多标记特征选择算法,并通过实验对比了不同处理方式的效果和缺失率对算法的影响。

【Abstract】 Nowadays,with the rapid development of computer science and Internet,not only the amount of data is increasing,but also the form of data is becoming more and more complex.So intelligent processing of data is becoming more and more important.Among them,pattern recognition,data mining,machine learning and deep learning have become the main way of processing and mining data.In the face of more and more complex and diverse data,data reduction plays a more and more important role.Multi label data is an extension of traditional single labeled data,and the form of data becomes more complex.Meanwhile,missing markers in multiple tags is also an important problem.Therefore,feature selection for multiple and weak label date is also important.Traditional rough set theory is an effective tool for dealing with uncertainty.It has been widely applied in feature selection,but it can deal with discrete data only.After the emergence of fuzzy rough sets,the problem is solved,and the fuzzy information theory is extended,including fuzzy entropy and fuzzy mutual information.The feature selection algorithm based on traditional information theory has been studied a lot.However,fuzzy information theory based multi label feature selection algorithm is relatively few,which can directly handle numerical data and hybrid data.Therefore,we will use fuzzy rough sets to deal with the problem of feature selection in multi label data.At the same time,a new multi label feature selection algorithm is proposed in view of the characteristics of multi label data.At the same time,multiple label with missing values is a common problem.In view of this situation,combined with the way of missing values in incomplete information system of rough set,the problem of feature selection under multi label with missing values is also dealt with.● Based on the fuzzy information theory and the correlation between labels,a multi label feature selection algorithm is proposed,and the effect analysis is carried out by the experiment.● Based on the feature correlation,a multi label feature selection algorithm which can remove redundant features is proposed,and the effect analysis is carried out through experiments.● Combined with the way of processing missing values in incomplete information system,a multi label feature selection algorithm under weak label is proposed,and the effect of different processing methods and the influence of missing rate on algorithms are compared through experiments.

  • 【网络出版投稿人】 浙江大学
  • 【网络出版年期】2018年 12期
节点文献中: