节点文献

基于模糊粗糙集的数据挖掘关键技术研究

Research on Key Technologies of Data Mining Based on Fuzzy Rough Set

【作者】 李代伟;

【导师】 李天瑞;

【作者基本信息】 西南交通大学 , 计算机科学与技术, 2023, 博士

【摘要】 大数据时代,数据纷繁复杂,具有很多不确定性。大数据智能主要在于发现和理解信息内容及信息间的关系。客观世界的不确定性,决定了人类主观认知过程的不确定性。这些大量的、不完全的、模糊的、有噪声的、随机的数据大多难以有效利用,其中隐藏的未知和有价值的信息不易被发现,数据挖掘是探究这些未知和有价值信息的有效技术。然而,数据的充足性、确定性和完备性是数据挖掘的有力保证。数据的缺失通常是不可避免且不利于数据的进一步分析。数据的不完备性和不确定性会导致数据分析的效率降低,增加数据分析的复杂度,导致无效的预测和分类等。粗糙集和模糊集理论是处理不完备和不确定数据的有效工具。本文以模糊粗糙集理论为研究工具,以缺失值填补和分类为主要技术,研究了不完备数据中的数据挖掘方法,主要工作如下:(1)针对处理缺失值数据时容易产生偏差和填补准确率偏低问题,从降噪和原始数据集的真实信息提取两方面研究缺失值填补算法。首先,为了最大限度地保留真实信息,无偏差地填补缺失值,利用最近邻的现有值预测缺失值,采用修正阈值改变决策属性集的模糊邻域以得到更精确的上、下近似值。其次,根据得到的上、下近似值和模糊决策计算出各邻域模糊隶属度所属的正确区间范围,然后从邻域获得的信息预测最终的隶属度,从而计算出需要填补的缺失值,提出了基于模糊粗糙最近邻拟合的缺失值填补算法Fitted FRNNI。考虑VQRS比implicator/t-模更能处理噪声数据,通过采用模糊量化的粗糙集计算对象的上、下边界值,进而取代Fitted FRNNI算法中通过模糊粗糙集计算的上、下边界值以提升算法效果,从而得到改进缺失值填补算法Fitted VQNNI。最后,通过分类精度、最近邻个数参数的影响以及决策集的模糊邻域半径和对象的模糊邻域半径等方面的算法评估和准确率分析验证了所提出算法的有效性。(2)针对如何通过准确表达训练集和测试集间的模糊相似关系提升缺失数据填补准确率问题,基于最近邻和聚类思想,结合模糊C均值(FCM)技术,进行缺失值填补算法研究。首先,利用FCM算法将完整的数据对象进行聚类,再在每个具有缺失值的对象对应的聚类簇中找到k个最邻近对象;采用不可区分矩阵、容差关系和模糊隶属关系,根据对应的相似对象和相关聚类识别潜在的最接近填补值;通过多次迭代确定模糊C均值聚类中的聚类参数和加权因子获取最佳预测精度,根据计算得到的隶属度、质心值和相应隶属度的总和获得缺失属性值,提出了联合模糊C均值和模糊量化最近邻的缺失值填补算法。其次,在前述算法基础上,进一步综合分析了每个簇对象从属特征的模糊隶属度,基于其最相关决策属性值,采用修正阈值改变决策属性集的模糊邻域获得更精确的上、下近似值,再针对每个对象对其相关簇执行模糊决策调整,进行缺失值的拟合填补,提出了联合模糊C均值和拟合VQNN填补的缺失值填补算法。通过RMSE和MAE评估分析、填补值与真实值的比较分析和分类精度结果分析验证了所提出算法的有效性。(3)针对基于模糊粗糙集的算法在分类决策过程中缺乏处理犹豫不确定性信息能力的问题,考虑犹豫模糊集具有传递犹豫信息能力,探索了犹豫模糊集和模糊粗糙集的融合,通过犹豫模糊元间的等价关系探究了高维犹豫模糊元的降维,从而扩展犹豫模糊集在决策分类中处理犹豫判断的能力。基于犹豫模糊元间的犹豫模糊相似性给出了犹豫模糊粗糙集上下近似的新定义,使用犹豫归一化汉明距离和犹豫归一化欧几里德距离融合产生的平均相似度作为决策要素,并根据获得的最大平均相似度进行分类,提出了犹豫模糊粗糙集最近邻分类算法HFRNN。通过分类准确度、执行时间、最近邻个数参数的影响等分析验证了算法的有效性。本文基于模糊集、粗糙集和犹豫模糊集理论,融合最近邻思想,结合模糊聚类策略,主要研究了缺失值填补技术和分类方法,提出了相应的缺失值填补算法和分类算法以提高基于模糊粗糙集的数据挖掘的效果。研究工作拓展了模糊粗糙集理论及应用的研究范畴,丰富了数据挖掘技术的研究手段,为大数据环境中的不确定性数据挖掘和智能决策提供了新的研究思路和方法。

【Abstract】 In the era of big data,data is complex and has many uncertainties.Big data intelligence mainly lies in discovering and understanding information content and the relationships between information.The uncertainty of the objective world determines the uncertainty of human subjective cognitive process.Most of these large-volume,incomplete,fuzzy,noisy and random data are difficult to use effectively,and the hidden unknown and valuable information is not easy to be found.Data mining is an effective technology for mining these unknown and valuable information.The adequacy,certainty and completeness of data are a powerful guarantee for data mining.The existence of missing data is unavoidable and not conducive to further analysis of data.The incompleteness and uncertainty of data can reduce the efficiency of data analysis,increase the complexity of data analysis,and lead to ineffective prediction and classification.The rough set theory and fuzzy set theory are effective mathematical tools to deal with incomplete and uncertain data.Based on the theories of fuzzy rough sets,this dissertation focuses on the missing data imputation and classification for data mining from incomplete data.The main research works and innovations are listed as follows:(1)Aiming at the problems of bias and low efficiency in imputation performance when processing data sets with missing values,the missing value imputation algorithms are studied by introducing the removal of noise information and real information extraction from original data sets.First,in order to retain the original information to the maximum extent and fill the missing values without bias,the existing values of the nearest neighbors are used to predict the missing values,and the correction thresholds are used to change the fuzzy neighborhood of the decision attribute set to obtain more accurate upper and lower approximations.Secondly,according to the upper and lower approximations and fuzzy decisions obtained,the correct range of the fuzzy membership of each neighborhood are calculated,and then the final membership value is predicted from the information obtained from the neighborhood,so as to calculate the missing values to be filled.A fitted fuzzy-rough imputation algorithms named Fitted Fuzzy-Rough Nearest Neighbor Imputation algorithm(Fitted FRNNI)is then proposed.Considering that VQRS is more capable of processing noisy data than t-norm,VQRS are used to calculate the upper and lower boundaries of the objects,and then replace the corresponding upper and lower boundary values in Fitted FRNNI algorithm respectively to improve the algorithm effect.Thus an improved missing value imputation algorithm named Fitted VQNNI algorithm is proposed.Finally,performance analysis has been conducted including classification accuracy analysis,the impact of the nearest neighbors parameter,and the weight coefficient of the fuzzy neighborhood radius of the decision set and the fuzzy neighborhood radius of the object to evaluate the proposed Fitted FRNNI and Fitted VQNNI algorithms.The experiments verify the effectiveness of the proposed algorithms.(2)Aiming at the problem of how to improve the accuracy of missing data filling by accurately expressing the fuzzy similarity between training sets and test sets,the missing value imputation algorithms are studied based on the nearest neighbor and clustering ideas,combined with Fuzzy C-Means(FCM)technology.Firstly,the FCM algorithm is used to cluster the complete data objects,and then find the k nearest neighbor objects in the corresponding cluster for each object with missing values.Specially,indistinguishable matrices,tolerance relations,and fuzzy membership relations are adopted to identify the potential closest filled values based on corresponding similar objects and related clusters.The best prediction accuracy is obtained by determining the clustering parameters and weighting factors in the fuzzy C-means clustering through multiple iterations.The missing attribute values are obtained according to the calculated membership degree and the sum of the centroid value and the corresponding membership degree.A missing value imputation algorithm named Jointly Fuzzy C-Means and VQNN(Vaguely Quantified Nearest Neighbor)Imputation(JFCM-VQNNI)is then proposed.Secondly,on the basis of JFCM-VQNNI algorithm,this dissertation synthetic analyzes the fuzzy membership of the dependent features for instances with each cluster,and considers the highly related decision attribute values.The correction thresholds are used to change the fuzzy neighborhood of the decision attribute set to obtain more accurate upper and lower approximations,and the fuzzy decision membership adjustment is performed on the related clusters for each object.A missing value imputation algorithm named Jointly Fuzzy C-Means and Fitted VQNN Imputation(JFCM-FVQNNI)is then proposed.The effectiveness of the proposed algorithms are verified by the analysis of RMSE(Root Mean Square Error),MAE(Mean Absolute Error),imputation values with actual values comparison,and classification accuracy results.(3)Aiming at the problem that the algorithm based on fuzzy rough set lacks the ability to deal with hesitation information in the process of classification and decision-making,considering that hesitant fuzzy sets has the ability to transmit hesitation and uncertainty information,and exploring the combination of hesitation fuzzy set and fuzzy rough set,a dimension reduction of high-dimensional hesitant fuzzy elements is explored through the equivalence relation between hesitant fuzzy elements,thus expanding the ability of hesitant fuzzy sets to deal with hesitant judgments in classification.Based on the hesitant fuzzy similarity between hesitant fuzzy elements,a new definition of upper and lower approximations of hesitant fuzzy rough sets is given.The average similarity generated from the fusion of Hesitation Normalized Hamming Distance and Hesitation Normalized Euclidean Distance as decision elements.The target instances are classified by employing the maximum average similarity obtained.So the capacity with hesitant fuzzy judgments is extended.An algorithm named Hesitant Fuzzy-Rough Nearest-Neighbor(HFRNN)is then proposed.The effectiveness of the proposed algorithms is validated through analysis of classification accuracy results,execution time,and the impact of the number of nearest neighbors.Based on the theories of fuzzy sets and rough sets and hesitation fuzzy sets,this dissertation integrates the ideas of nearest neighbor and fuzzy clustering strategy,mainly studies missing value imputation techniques and classification methods,and proposes corresponding missing value imputation algorithms and classification algorithms to improve the effect of data mining based on fuzzy rough sets.The research work extends the study fields of fuzzy rough set theory and applications,and further enriches the research methods of data mining,which provides new ideas and methods for uncertain data mining and intelligent decision-making in the big data environment.

  • 【分类号】TP18;TP311.13
节点文献中: