节点文献
基于改进限制容差关系的信息系统知识约简
Knowledge Reduction of Information System Based on Improved Limited-Tolerance Relation
【作者】 刘娟;
【导师】 秦克云;
【作者基本信息】 西南交通大学 , 基础数学, 2005, 硕士
【摘要】 粗糙集理论是波兰科学家Z.Pawlak于1982年首先提出的一种数据分析理论,目前已发展成为一种处理不确定性信息的数学理论,并且成功地应用于机器学习、数据挖掘(data mining)、智能数据分析、控制算法获得等领域。 Pawlak最初提出的粗糙集理论是建立在等价关系基础之上的,然而相应的理论不适用于处理不完备的信息系统,现实中不完备信息系统的广泛存在极大地限制了粗集理论的应用领域。于是在后来的粗糙集理论研究中,研究者提出了各种扩充的粗糙集模型,如一般关系下的粗糙集模型、变精度粗糙集模型、模糊粗糙集模型、概率粗糙集模型等。针对不完备信息系统,为了刻划对象间的不可区分关系,Krysckiewcz提出了容差关系;Stefanowki等人提出了非对称相似关系和量化容差关系;王国胤在容差关系和非对称相似关系基础上提出了介于两者之间的限制容差关系。本文在分析以上关系的基础上,提出了改进限制容差关系。该关系的特点是:通过引入阈值先将原不完备信息系统进行划分,再利用联系度的概念确定改进限制容差类,基于产生的这些类得到上下近似。本文接着讨论了上下近似的代数性质,并把在完备信息系统基础上建立的一些粗糙集理论的重要概念引入到不完备信息系统中,对不完备信息系统进行了更深入地探讨。 属性约简是信息系统知识发现研究的核心内容之一,对完备信息系统的约简问题,目前学术界进行了大量的研究,其中包括基于正域的约简、基于信息熵的约简、基于包含度的约简等。本文基于改进限制容差关系,把正域约简、信息熵约简以及张文修等针对不一致决策表提出的分布约简、分配约简、最大分布约简和近似约简引入不完备信息系统,并讨论它们之间的关系,且证明了对于相容的不完备决策表,熵约简、分布约简、正域约简、最大分布约简、分配约简及近似约简都是等价的;文中通过定义属性的信息量,给出了分配约简的一种启发式算法:条件信息量约简算法,分析了该算法的时间复杂度。经实验检验,该算法是有效的。
【Abstract】 Rough set theory is a new theory of data analysis, it was first put forward by Poland scientist Z.Pawlak. At present it has developed to be a new math theory tool to deal with vagueness and uncertainty. It has been applied to many areas successfully including machine learning, pattern recognition, decision support, data mining, process control and predictive modeling.Pawlak rough set was established at the base of equivalent relation. But it is not suit to deal with incomplete information system. So, in the later research of rough set theory, the researchers address all kinds of models in order to expand its applying aspect. Examples are gerneralized binary-relation rough set model, variety precision rough set model, vague rough set model, probability rough set model etc. In order to describe the indistinct relationship of objects in incomplete information system, Krysckiewcz presented tolerance relation; Stefanowki etc. presented similarity relation and quantitative tolerance relation. Based on these relations, Wang G Y presents limited tolerance relation, followed by proving this model is superior to tolerance relation and similarity relation. In this paper, these relations are analyzed and improved limited-tolerance relation is presented. In this relation, thresholds are introduced and an incomplete system is divided into two parts. The aim of tolerant relation based on connection degree is to ascertain classes by which the lower approximation set and upper approximation set of a set in universe are confirmed. In the second part, this paper discusses algorithm property and presents some useful exploration about incomplete information system by introducing some important definition of complete information system.Knowledge reduction is one of the most important problems in the study of information system and knowledge discovery. There are many types of knowledge reductions in the area of rough sets. In this paper, some types of reductions of complete information are first presented to incomplete information system, followed by the relationship between these reduction methods. Also, for incomplete consistent decision tables, entropy reduction, distribution reduction, positive domain reduction, maximum distribution reduction, assignment reduction and approximate reduction are all proved to be equivalent. Based on condition information quantity, a heuristic algorithm for assignment reduction is presented, and the complexity of this algorithm is analyzed. Finally, the experimental result shows this algorithm can find this assignment reduction for incomplete information system.
【Key words】 rough set; incomplete information system; information quantity; knowledge reduction;
- 【网络出版投稿人】 西南交通大学 【网络出版年期】2005年 06期
- 【分类号】O159
- 【被引频次】2
- 【下载频次】156