节点文献

S-粗集与数据筛选—过滤

S-rough Sets and Data Sieve-filtration

【作者】 蔡成闻

【导师】 史开泉;

【作者基本信息】 山东大学 , 模式识别与智能系统, 2008, 硕士

【摘要】 自二十世纪七十年代大规模集成电路、超大规模集成电路诞生以来,计算机已经成为现代工业、商业、农业等各个领域必不可少的一个工具,但随之而来的是数据的迅速膨胀,使得人类在一个极短的时间里进入了数据爆炸的时代。这些数据具有巨大性、随机性、不确定性等特征,并且数据的生成过程又往往存在着动态特征。实际上,在这些大型的、复杂的、信息丰富的数据中,只有一小部分是人们真正需要的,如何从其中提取出人们所需要的信息,已经成为目前一个重要的课题。粗集理论是波兰数学家Z.Pawlak在1982年首次提出的,这是一种处理不完整、不精确问题的新型数学工具,它通过等价关系和近似概念对数据进行约简以获取知识。粗集知识系统是一个基于规则的系统,它不需要精确的数学描述,而是对经验的总结,因此非常适合数据处理过程中直观、简单、易于理解、人性化、智能化的要求,为数据挖掘技术提供了理论基础和研究思路。传统的数据挖掘方法是建立在数据不会发生变化的假设下进行讨论的,可以说是一种静态的数据挖掘方法,实际上数据不可能是一成不变的,当数据发生变化时,静态的数据挖掘方法便失去了效用,因此传统的数据挖掘方法具有局限性。奇异粗集(Singular Rough Sets,简称S-粗集)是Z.Pawlak粗集的一种改进形式。它是山东大学史开泉教授于2002年提出的,是基于元素迁移的概念建立起来的一种动态粗集。S-粗集具有三种形式:单向S-粗集(One directionS-rough sets),单向S-粗集对偶(Dual of one direction S-rough sets),双向S-粗集(Two direction S-rough sets)。S-粗集的动态特征、遗传特征、粒度特征等特性,S-粗集的提出为我们研究动态数据挖掘开辟了一个全新的方向并提供了必要的理论保证。本文的主要工作如下:1.主要介绍了数据挖掘的发展研究现状以及数据挖掘的分类;阐述了粗集理论提出的背景、发展状况、研究的内容和方向;介绍了S-粗集提出的背景及研究现状;并将S-粗集的理论进行了简单的介绍。2.利用S-粗集的动态特征、遗传特征、粒度特征等特性,给出了S-粗集与数据筛选-过滤的研究,讨论了数据的粒度特征、单向筛选-过滤、双向筛选-过滤,给出了f-筛选-过滤度、(?)-筛选-过滤度和(?)-筛选-过滤度的概念,并提出了筛选-过滤定理和筛选-过滤准则。3.提出了基于S-粗集的动态聚类方法,利用第3章给出结果,提出了一种基于S-粗集的动态聚类算法。利用此算法改进了无线传感器网络的分簇算法,通过仿真,并与现有算法比较后,得到这样的结论:使得每个节点的能量得到均匀的使用,提高了节点的能效比,满足了无线传感器网络节能的要求。

【Abstract】 Since LSI (Large-scale integration) and SLSI (super-large-scale integration) have been produced from 1970s, computers have been an indispensable implement in modern industries, business and agriculture. However, with data expanding rapidly at the same time, the mankind enters a data explosive period very soon. These data have the characteristic of hugeness, random and uncertainty, whose generating process has dynamic characteristic. In fact, it is only a small part needed by people in the large-scale and complicated data which includes abundant information, then how to mine out the information needed becomes an important question for discussion. Rough sets theory was proposed firstly by a Polish mathematician Z.Pawlak in 1982, which is a new mathematic implement used to solve problems about incompletion and imprecision, and it gets the knowledge by data reduction using equivalent relation and approximation concept. Rough sets knowledge system is based on rule, which is a conclusion about experience and needn’t the exact mathematic description, so it meets the intuitionistic, simple, comprehensible, humanized and intelligentized demands in the data disposal process, providing the theory base and research direction for data mining technique.The traditional data mining method is discussed under the assumption that the data won’t change, which can be called a static data mining method. But actually data can’t keep unaltered, so when the data changes, the static data mining method won’t do its work, which is one limitation of the traditional data mining method. S-rough sets (Singular Rough Sets) which is proposed by Professor Shi Kaiquan with Shandong University in 2002, is an improvement based on Z.Pawlak rough sets, and it is a dynamic rough sets based on element transfer. S-rough sets has three forms: one direction S-rough sets (One direction Singular rough sets), dual of one direction S-rough sets (Dual of one direction Singular rough sets), and two direction S-rough sets (Two direction Singular rough sets). S-rough sets has dynamic, hereditary and granularity characteristic, which proposes a new research direction and provides the theory bases for dynamic data mining.The main work in this paper is as following:1. Introducing the research actuality of the data mining and its classes, and explaining the background, development, the research content and the direction of Rough sets theory, and outlining S-rough sets theory.2. By using dynamic, hereditary and granularity characteristic of S-rough sets, this paper carries the research of S-rough sets and data filter-filtration which includes discussing the granularity characteristic, one direction filter-filtration and two direction filter- filtration of data, and giving the concepts of f-filter-filtration degree, f-filter-filtration degree and (?)-filter-filtration degree, and proposing filter-filtration theorem and filter-filtration criterion.3. Proposing a kind of dynamic clustering algorithm based on S-rough sets with the results in chapter 3, which improves the clustering algorithm of WSN (wireless sensor net). Comparing it with the existing algorithms, we make such a conclusion: the algorithm makes the energy used equably at every node and the ratio of energy consumption to energy available reduced, so it meets the actual requirement for WSN.

  • 【网络出版投稿人】 山东大学
  • 【网络出版年期】2009年 01期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络