节点文献

基于粗糙集的属性约简算法及其应用研究

Knowledge Reduction Algorithms and Application Based on Rough Sets

【作者】 颜艳

【导师】 杨慧中;

【作者基本信息】 江南大学 , 检测技术与自动化装置, 2008, 硕士

【摘要】 Pawlak Z提出的粗糙集(Rough Sets,简称RS)理论是一种全新的刻划不完整性和不确定性的新的数学工具,它能有效地分析和处理不精确、不完整、不一致等数据,并从中发现隐含的知识,揭示潜在的规律。目前该理论已得到了国际众多学者的重视,RS理论已被广泛应用于数据挖掘、机器学习、数据库知识发现、决策支持系统、故障诊断等领域。本文着重对粗糙集理论的核心问题之一——决策表的属性约简问题进行了研究。属性约简是在保持知识库的分类或决策能力不变的前提下,删除其中不相关或不重要的知识。具体研究内容如下:对于完备的离散型的信息系统,从信息论的角度考虑粗糙集属性约简问题。基于互信息的概念定义了一种新的属性重要度,并以此属性重要度为启发式信息提出了一种基于改进的互信息增益率的启发式算法。利用条件熵计算属性间的相关性,并将属性相关性的定义融入到遗传算法的适值函数中,使得约简结果含有较少的属性,而且降低了它们的相关性。经典粗糙集理论不能处理不完备信息系统,在深入学习和研究了现有的几个关于不完备信息系统的粗糙集扩展模型的基础上,指出它们的不足之处。因为条件属性的重要性存在差异,通过引入差异度,对不完备信息系统中属性的重要性进行了定义,提出了一种基于权重联系度的属性约简算法,通过实例仿真说明该算法的优越性。制约粗糙集理论发展和应用的另一方面是,该理论无法直接用于连续数据。目前处理连续数据的方法大部分是基于数据离散化,但是这种方法在某种程度会造成信息的损失。引入样本之间的相似性和改进的属性广义区分度的概念,并定义属性的全局相似性程度,根据样本之间的全局相似关系直接对属性值为连续数据的决策系统进行属性约简,避免了数据离散化过程中信息的丢失。最后将其应用于汽轮机组故障诊断系统中,实验结果表明该方法的有效性。

【Abstract】 Rough Set (RS) theory, introduced by Pawlak Z, is a novel mathematical tool to deal with vagueness and uncertainty. It is a powerful mathematical tool for analyzing uncertain, fuzzy knowledge and can effectively deal with the imprecise, incomplete, or uncertain data. Now it has attracted much attention of researchers around the word. In recent years, it has been successfully applied to data mining, machine learning, knowledge discovery from database, decision support systems, fault diagnosis etc.This article emphatically studies on one of the important problem of Rough Set theory—the reduction of the decision table. Attribute reduction preserves the original meaning and reduces the irrelevant and unimportant knowledge. The details are studied as follows:In regard to a complete and discrete information system, consider attribute reduction in the view of information theory. A developed attribute importance measure method is defined based on the mutual information between selected attribute and decision attribute, and the measure is used as the heuristic information in the proposed algorithm. Conditional information entropy is used to compute relevance of attributes and it is used in fitness function of genetic algorithm to assure reduction has few attributes and relevance between attributes.Traditional Rough Set theory is generally incapable of handling incomplete information system. After studying the extensions of Rough Set model, point out their shortages. For essentiality of attribute existing difference, a developed attribute importance measure method is defined based on the difference degree of attributes. It’s proposed an attribute reduction algorithm based on connection degree of essentiality of attribute. An example shows that the proposed algorithm is an effective method.Another the Rough Set theory defect which blocks its development and application is that it can not be employed on continuous values directly. Previously discretization method is applied beforehand in order to transform the data into discrete values, but this may result in information loss. The notions of similarity between objects and improved general important degree of an attribute are introduced. The global similarity measure between objects is defined by them. A direct reduction method is applied to continuous attributes using tolerance relation by the global similarity relation. This method avoids losing the information in the data’s discretization progress. Finally, the method is applied to fault data, and the result shows that the method is effective.

  • 【网络出版投稿人】 江南大学
  • 【网络出版年期】2009年 03期
节点文献中: