节点文献

区间数据库的构建及其在知识发现中的应用

Construction of Interval Databases and Its Applications in Knowledge Discovery

【作者】 尹云飞

【导师】 张师超; 严小卫;

【作者基本信息】 广西师范大学 , 计算机软件与理论, 2005, 硕士

【摘要】 关联规则挖掘是数据挖掘中的一个重要研究课题。它是搜索强相关的项集合的一个过程。挖掘一个超市数据库, 可以找到不同商品之间的销售联系(它反映了顾客的消费行为),例如:面包与牛奶、咖啡与方糖、牙膏与牙刷等通常被同时销售。这些是常识性知识。有趣的是, 关联规则挖掘能找到,像“啤酒与尿不湿”被同时销售, 这种非常识性知识。这导致关联规则挖掘被深入研究和广泛的应用。例如, 它被进一步用于解决库存控制(stock control) 、商品促销(Sales promotion) 、消费者行为分析(Customer behavior analysis)等问题。随着超市和日用品工业的发展,捆绑销售(Binding sale)方式——捆绑商品(Binding commodities)销售已成为方便顾客并提升利润的一种重要手段。这正是关联规则挖掘的用武之地。本论文深入细致地研究了这个问题,并提出了挖掘区间值规则:A→[B, C]的思想和方法。捆绑商品借助区间值(Interval values)来表示有很多优点。首先区间值包含了比单个具体数据更多的信息。因为单个数据提供的只是单个数据本身,而区间值提供的是一个分布,即, 可以取区间内的任意一个数。其次区间值比平均数有更强的表达能力,也就是说区间值的信息熵(Interval entropy)要大于平均数的信息熵(Mean entropy)。再者,区间值数据库挖掘可以发现哪些商品适合于捆绑、哪些商品不适合于捆绑。这有重要的实际应用价值。论文在对区间值聚类算法研究的基础上, 提出将传统关系数据库的两个字段看成一个新字段,并用其中一个来表示新字段的“左端点域”(区间值左端点)用另一个来表示新字段的“右端点域”( 区间值右端点),由此形成了区间值数据库。论文深入研究了强关联规则( 亲属关联规则) 的挖掘算法,给出了强关联规则的区间函数公式; 在对这些区间函数值研究的基础上,构建了一种完备区间格系统,并利用完备区间格满足的一个性质:A∧C=B∧C且A∨C=B∨C ?A=B 来对商品进行捆绑。区间值关联规则挖掘的实质是对捆绑商品的挖掘,也就是研究哪些商品应该被捆绑。本论文的主要工作分为如下四个部分: (1) 提出传统数据挖掘中存在许多模式遗漏问题,并从物理学、数学、生物学等角度论述研究这些遗漏模式的重要意义。(2) 针对这些遗漏模式构建一种新型的数据库结构来存放和处理它们,这种新型的数据库就称为区间值数据库。(3) 提出了区间值关联规则的概念,并深入研究了区间值规则的真正内涵。(4) 区间值规则挖掘算法的研究。最后对本论文的主要工作做了总结,指出今后的改进方向。

【Abstract】 Association rule mining is one of important topics in data mining. It is a procedure of identifying strong interactions among itemsets in databases. For example, mining a transaction database in a supermarket can discovery associations (customer behaviors) within different commodities, such as bread and milk, coffee and sugar, toothpaste and toothbrush. While they are commonsense, association rule mining can find many other interesting interactions, such as ‘beer and diaper’. This leads to a deeper research on development and wide applications of association rule mining. For example, it can be helpful in solving stock control, sales promotion, and customer behavior analysis in supermarkets. With the development of the supermarket and the commodity industries, binding sale, namely binding commodity, is rapidly popularized and becomes an important meanings of gaining profits. Association rule mining assists in binding sales. After our deeper understanding and researching, a kind of novel association rules of the form A→[B, C], referred to as interval rules, is proposed and studied in this thesis. There are many advantages for using interval values to represent binding commodities. Firstly, an interval contains more information than a single value. Because a single value offers the single value itself, while an interval offers the distribution of the values, i.e., any numbers in the interval can be taken. To follow up, an interval is more expressive than a mean, that is to say entropy of an interval is larger than the entropy of a mean. Thirdly, in an interval database it can be discovered which commodities are fit to be bound. This is much more important in reality. Based on the research about the interval clustering algorithms, this paper proposes to take the two fields in traditional database as a new field, and uses one of them to express the ‘left field’of the new field (the left boundary of the interval), and uses the other to express the ‘right field’of the new field (the right boundary of the interval), So we obtain the interval database. This paper conducts a deeper research on the algorithms of strong association rules (affinity rules), and offers the function formula for strong association rules. On the basis of the research about these function values, a complete interval allocation lattice system is constructed, and a property satisfied by the complete interval allocation lattice system, i.e. A∧C=B∧C and A∨C=B∨C ?A=B, is used to bind commodities. The essence of interval association rule mining is how to find out the binding commodities, i.e., to decide which commodities should be bound. The main contributions of this paper are divided into the following four parts: (1) Limitations of the traditional data mining, i.e. pattern missing, are studied in views of physics, math, and biology. (2) For the pattern missing problem, a kind of new data structure, i.e. interval database, is designed. (3) The concept of interval association rules mining is proposed, and a deeper research is conducted from the real meanings of the interval. (4) An algorithm for mining interval rules is designed, simulated and tested. For future work, some suggestions for improvements are given as well.

  • 【分类号】TP311.13
  • 【下载频次】119
节点文献中: 

本文链接的文献网络图示:

本文的引文网络