节点文献

条件独立性在关联规则挖掘中的研究和应用

Research and Application of Conditional Independences in Association Rule Mining

【作者】 陈斌

【导师】 倪天倪;

【作者基本信息】 河海大学 , 计算机应用技术, 2004, 硕士

【摘要】 随着信息技术的应用普及,数据爆炸和知识贫乏之间的矛盾越来越大,使数据挖掘的深入研究和广泛应用势在必行。在数据挖掘的各分支中,关联规则挖掘的研究最为深入和广泛。对关联规则挖掘的研究又主要集中在频繁集的生成优化和事物集的扫描次数两个方面,并且主要基于支持度---可信度框架,由于这种框架的自身缺陷,使挖掘的关联规则中用户感兴趣的却不多,因此如何使用户对挖掘的关联规则更感兴趣又成为一项新研究任务,不少学者不采用支持度---可信度框架,尝试采用新方法来进行关联规则挖掘,以提高用户的满意度和兴趣度。本文正是在这种背景下,研究基于条件独立性的关联规则挖掘的算法框架;研究如何在传统关联规则挖掘的基础上,利用条件独立性进行后处理,提高关联规则的有趣性。 本文的主要内容如下: 1.探讨了传统关联规则挖掘的主要思想和技术,分析了各种频繁集裁剪技术和兴趣度度量。 2.给出了一种利用马尔可夫覆盖进行关联规则挖掘的算法框架,并研究了算法中的各个组成部分。 3.提出了多项集的马尔可夫覆盖的生成方法,证明了其正确性,然后探讨了多变量马尔可夫覆盖的贝叶斯网络表示形式。 4.面向教育评估系统中的具体应用,本文提出了对原有系统中采用的Apriori算法挖掘的关联规则进行后处理的方法:采用条件独立性和传统支持度---可信度框架相结合的方法进行关联规则的过滤,并从中发现存在的条件独立性限制。

【Abstract】 Because of broadly using of information technologies, the conflict between explosion of data and poorness of knowledge has been more and more acute, and the necessity of data mining becomes more and more urgent too. Among all branches of data mining, the research of association rule mining was the deepest one and the application of association rule mining was the most widely used one. Research of association rule mining almost concentrated on the optimality of frequent sets generation and the reduction of scanning times of transaction sets. Most of them were based on support-confidence framework. For the instinct shortcoming of the frame, few association rules were interesting. So the exploration of how to improve the interest of association rules became a novel and popular task in the research of association rule mining. Several experts employed new measures to improve the contentness and interest of association rules without adopting support-confidence framework in association rule mining. Under such backgrounds, this paper makes a study of the algorithmic framework of association rule mining based on conditional independences; explores how to perform post-processing and how to improve the interest of association rules with constraints of conditional independences after the processing of traditional association rule mining.Firstly, this paper explores main ideas and popular algorithms of traditional association rule mining, and analyzes different pruning technologies of frequent sets and different measures of interest. Secondly, this paper puts forward the algorithmic framework of association rule mining using Markov Blanket, and also explores each components of algorithmic framwork. Thirdly, this paper gives out a method of discovering Markov Blanket of multi-sets rather singleton, and proves the correctness of the method, afterwards this paper expresses it using Bayes Network too. Finally, aimed at the application of educational evaluation, this paper performs post-processing at those rules mined by Apriori algorithm, and filters those rules with the constraints of conditional independences, and reads out conditional independences from rules.

  • 【网络出版投稿人】 河海大学
  • 【网络出版年期】2004年 03期
  • 【分类号】TP311.13
  • 【被引频次】3
  • 【下载频次】167
节点文献中: 

本文链接的文献网络图示:

本文的引文网络