节点文献

FP-growth算法的实现方法研究

Research on Implementation of the FP-growth Algorithm

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 王新宇杜孝平谢昆青

【Author】 Wang Xinyu 1 Du Xiaoping 2 Xie Kunqing 11 (School of Information Science and Technology,Peking University,Beijing100871) 2 (Software School of Beijing University of Aeronautics and Astronautics,Beijing100083)

【机构】 北京大学信息科学技术学院北京航空航天大学软件学院北京大学信息科学技术学院 北京100871北京100083北京100871

【摘要】 事务数据库中频繁模式的挖掘研究作为关联规则等许多数据挖掘问题的核心工作,已经研究了许多年。早期算法大都是Apriori型算法,即首先产生候选集,然后在候选集的基础上找出频繁模式,候选集的产生往往是耗时的,特别是挖掘富模式或长模式时。JianweiHan等人提出了一种新颖的数据结构FP-tree及基于其上的FP-growth算法,用于有效的富模式与长模式挖掘。由于不同的实现方法可能会导致不同的挖掘效率,该文在讨论FP-growth算法的基础上,采用了几种不同的方法来实现它,并用几个数据库对它们的性能进行了比较。

【Abstract】 Mining frequent patterns in transaction databases,as an essential role in many data mining tasks such as the association rule mining,has been widely studied for many years.Most of the previous studies adopt an Apriori-like candidate set generation-and-test approach.However,candidate set generation is costly if there exist prolific patterns or long patterns.Jianwei Han et al propose a frequent pattern tree structure and a FP-growth algorithm based on this structure that can mine the frequent patterns by pattern fragment growth.Due to different methods will result in different performance,in this paper several methods to implement the FP-growth algorithm are discussed.The performance is studied,analyzed and compared on several canonical datasets.

【关键词】 频繁模式关联规则数据挖掘算法
【Key words】 Frequent PatternAssociation RuleData MiningAlgorithm
【基金】 国家973重点基础研究发展规划项目(编号:G1999032705);留学回国人员科研启动基金资助
  • 【文献出处】 计算机工程与应用 ,Computer Engineering and Applications , 编辑部邮箱 ,2004年09期
  • 【分类号】TP311.13
  • 【被引频次】74
  • 【下载频次】1029
节点文献中: 

本文链接的文献网络图示:

本文的引文网络