节点文献

一种基于粗糙集的不完备信息处理方法研究

Research on an Approach of Incomplete Information Processing Based on the Rough Set Theory

【作者】 张在美

【导师】 李仁发;

【作者基本信息】 湖南大学 , 计算机系统结构, 2007, 硕士

【摘要】 在现实数据库知识发现过程中,由于数据采集能力有限或数据丢失等原因,使得所面临的数据库往往是不完备的信息系统,即可能存在部分对象的某些属性值未知的情况。空缺数据的处理非常关键,因为不完备的数据能够使知识挖掘过程陷入混乱,导致不可靠的输出,将严重影响挖掘的效果。粗糙集理论作为一种处理模糊、不确定知识的数学方法,其显著的优点是无需提供所需处理的数据集合之外的任何先验信息,近年来已在知识发现上取得了令人瞩目的研究成果。目前,基于粗糙集理论的不完备信息系统知识发现的理论框架已基本完整,但在具体知识获取的多样性及知识质量的提高方面还需要进一步努力。本文的主要工作就是以粗糙集理论为工具,对知识发现过程中信息不完备问题的处理方法进行研究,以提高知识发现的质量和效率。不完备信息系统的知识发现有两种实现途径:一是采用数据补齐算法对缺失值进行填充,在完备化的信息系统基础上进行知识获取;二是在不改变原不完备信息系统的基础上直接进行知识获取。本文从这两种途径入手,利用粗糙集的方法,提出了两个不完备信息处理的有效算法。首先,分析了目前数据补齐算法存在的缺陷及产生这些缺陷的原因。通过对拓展粗糙集理论模型作进一步的改进,并合理引入分治思想,提出了一种新的数据补齐算法。结合理论分析和实例阐述了算法的有效性,并通过在UCI机器学习数据库中选取的两个数据集上进行实验,验证了该算法不仅能够提高补齐率,而且能显著降低算法复杂性。其次,本文在不改变原不完备信息系统的基础上,分析了现有知识约简算法的局限性,扩展定义了不完备熵概念,与传统粗糙熵结合,对不完备信息系统中的属性重要性进行了定义,并以此作为启发式信息,提出了一种优化的不完备信息系统知识约简算法,与传统方法相比能够找出更优的最小约简。通过理论和实例分析说明了算法的有效性。

【Abstract】 In the process of Knowledge Discovery in Databases, people often face incomplete information system, that is, a substantial proportion of the data may be missing in real-world applications. It is very important to deal with incomplete data, because it may lead to confusion and irresponsible outputs in data minning. As a new mathematical tool for dealing with inexact, uncertainty or vague knowledge, the rough set theory has got great success in KDD in recent years, and the most prominent advantage is that, it needs only the data provided in the information systems, relying on no other model assumptions. At present, the theoretical frame of KDD in incomplete information system based on rough set theory is basiclly completed, but the variety and quality of knowledge extracted is still need to be improved.The main work of this paper is to give in-depth study on the processing method of incomplete data problem using rough set theory, to improve the quality and effiency of KDD. There are two methods of KDD in incomplete information systems: one is to complete the incomplete information system first, and then extract knowledge based on the completed system; the other is to extract knowledge directly from the incomplete information system with no change on original system. This paper starts with this two kinds method, provides two new algrithms under rough set theory to improve the KDD performance. Firstly, this paper analyzes the limitation of data filling algorithms in existence, extends the valued tolerance relation matrix in rough set theory, introduces divide-and-conquer idea, and then provides a new algorithm RSDIDA. The experimental result demonstrates that it improves the filling ratio and efficiency greatly. Secondly, this paper provides an optimized knowledge reduction algorithm for incomplete information system with no change on original system. It first analyzes the limitation of traditional rough entropy and correlative knowledge reduction algorithm, extends the incomplete entropy, which describes the uncertainty of knowledge more precisely, and then uses rough entropy and new incomplete entropy to define attribute significance. Accordingly, the new knowledge reduction algorithm is provided. The example analysis proves its validity.

  • 【网络出版投稿人】 湖南大学
  • 【网络出版年期】2007年 02期
  • 【分类号】TP311.5
  • 【被引频次】27
  • 【下载频次】587
节点文献中: