节点文献

基于二进制可辨矩阵的属性约简研究

Research of Attribute Reduction Based on Binary Discernable Matrix

【作者】 汪小燕

【导师】 王浩;

【作者基本信息】 合肥工业大学 , 计算机应用技术, 2006, 硕士

【摘要】 数据挖掘(DM)是从数据中提取人们感兴趣的、潜在的、可用的知识,并表示成用户可理解的形式。分类是数据挖掘的一个重要分支,粗糙集方法是数据挖掘中的重要分类技术之一。粗糙集理论是一种处理模糊和不精确知识的数学工具,它具有很强的知识获取能力。粗糙集理论在数据挖掘中的应用是一个较新的研究领域。由于粗糙集理论提供了严格的处理数据分类问题的数学方法,不需要任何数据的附加信息,能够搜索数据的最小集合,可以使用定性与定量的数据,并从数据中产生决策规则集合等优点而得到广泛的应用。 对于分类来说,并非所有的条件属性都是必要的,有些是多余的,去除这些属性不会影响原来的分类效果,反而会提高系统潜在知识的清晰度。决策表的属性约简就是约简决策表中的条件属性,约简后的决策表具有约简前决策表的功能,但是约简后的决策表具有更少的条件属性。 本文主要对粗糙集理论中的二进制可辨矩阵进行研究,研究了基于二进制可辨矩阵的知识粒度的有关理论和计算公式。利用获得的公式可计算知识的分辨度和粒度,以及属性的重要度。并利用得出的有关理论进行决策表的属性约简和值约简,提出了两种约简算法:一种是基于二进制可辨矩阵的属性及属性值约简算法,该算法只要扫描一次二进制可辨矩阵,就可求得核属性和去除核属性后,所增加的不能被正确分类的对象,从而得出核值。同时将吸收律应用于各析取式,可求得条件属性的约简集,从而得到具有约简属性的核值表。该算法使得属性约简和属性值约简得以一致计算,大大缩短了约简时间。 另一种是基于二进制可辨矩阵的重要度的属性及属性值约简算法(BDMSR):该算法利用二进制可辨矩阵的属性重要度作为属性选择标准,以在获取核属性的基础上,通过逐个增加属性构成决策表的最小约简。该算法也使得属性约简和属性值约简得以一致计算。 此外,我们设计了基于BDMSR算法和基于二进制可辨矩阵的属性约简算法(BDMR)的原型系统,在此统一的平台上,我们通过对UCI提供的多个标准测试数据集进行测试,对两种算法的性能进行比较。实验证明,BDMSR算法确实优于BDMR算法。

【Abstract】 Data Mining is the process of mining the interesting, potentially useful, and understandable knowledge in data. Classification is an important sub-branch of Data Mining and the method of Rough Set is one of the important techniques of Classification. Rough Set is a new mathematical tool to deal with fuzzy and uncertain knowledge. It has strong knowledge obtaining ability. It is the new research domain that the theories of Rough Set are applied in Data Mining. Because the theories of Rough Set provide the strict mathematics method of the problem of dealing with datas Classification, without the additional information of any datas, they can search for the smallest assembly of the datas , may use datas of determining the nature and fixing quantity and also can generate the rule assembly of policy decision from the datas,etc. Rough Set has got the extensive application.Not all the condition attributes are necessary for classification, some are unnecessary. Doing away with these attributes won’t affect original classification effect, on the contrary, it will improve the articulation of the latent knowledge in system. Attribute reduction is reducing condition attributes in a decision table .The decision table after being reduced has the function of the one before being reduced,but the former has the less condition attributes .The dissertation is mainly researching on Binary Discernable Matrix in the theories of Rough Set and related theories and calculation formulas based on the knowledge granulation. We can calculate the discernment degree, granulation of knowledge and attribute significance making use of formulas obtained. We can also go on attribute reduction and value reduction of the decision table utilizing related formulas obtained and present two reduction algorithms:one is attribute and its value reduction algorithm based on Binary Discernable Matrix.Only scanning Binary Discernable Matrix once, the algorithm can get core attributes and objects that can’t be classified correctly, so we can obtain core values. Reduction assemblies of condition attributes can be acquired if we apply the role of absorptive to every disjunct normal form, so we can get a core value table having attributes which are reduced. The algorithm makes that attribute reduction and its value reduction are calculated identically, so it shortens time of reduction greatly.The other is the attribute and its value reduction algorithm based on attribute significance of Binary Discernable Matrix(BDMSR):on the basic of getting core attributes,the smallest reduction assembly of a decision table is formed by increasing a attribute one by one utilizing attribute significance of the Binary Discernable Matrix as a criterion of attribute selection. The algorithm makes that attribute reduction and attribute value reduction are calculated identically.too.In addition, we designed a prototype system based on BDMSR algorithm model and attribute reduction algorithm model based on Binary Discernable Matrix(BDMR). On this uniform flatform, we compared the BDMSR algorithm and BDMR algorithm by using the standard UCI data sets. From the experiment, we can see the BDMSR algorithm is superior to BDMR algorithm indeed.

  • 【分类号】TP311.13
  • 【被引频次】3
  • 【下载频次】170
节点文献中: 

本文链接的文献网络图示:

本文的引文网络