节点文献

一种挖掘压缩序列模式的有效算法

An efficient algorithm for mining compressed sequential patterns

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 童咏昕张媛媛袁玫马世龙于丹赵莉

【Author】 Tong Yongxin~1,Zhang Yuanyuan~2,Yuan Mei~3,Ma Shilong~1,Yu Dan~1,and Zhao Li~1 1(State Key Lab.of Software Development Environment,Beihang University,Beijing,100191) 2(China Academy of Telecommunication Technology,Beijing,100191) 3(College of Information,Beijing Union University,Beijing,100084)

【机构】 北京航空航天大学软件开发环境国家重点实验室电信科学技术研究院北京联合大学信息学院

【摘要】 从序列数据库中挖掘频繁序列模式是数据挖掘领域的一个中心研究主题,而且该领域已经提出和研究了各种有效的序列模式挖掘算法。由于在挖掘过程中会产生大量的频繁序列模式,最近许多研究者已经不再聚焦于序列模式挖掘算法的效率,而更关注于如何让用户更容易地理解序列模式的结果集。本文受到压缩频繁项集思想的启发,提出了一种CFSP(CompressingFrequent Sequential Patterns)算法,其可挖掘出少量的有代表性的序列模式来表达全部频繁序列模式的信息,并且清除了大量的冗余序列模式。CFSP是一种two-steps的算法:在第一步,其获得了全部闭序列模式作为有代表性序列模式的候选集,与此同时还得到大多数的有代表性模式;在第二步,该算法只花费了少量的时间去发现剩余的有代表性序列模式.一个采用真实数据集与模拟数据集的实验研究也证明了CFSP算法具有高效性.

【Abstract】 Mining frequent sequential patterns from sequence databases has been a central research topic in data mining and various efficient mining sequential patterns algorithms have been proposed and studied.Recently,many researchers have not focused on the efficiency of sequential patterns mining algorithm,but have paid attention to how to make users understand the result set of sequential patterns easily,due to the huge number of frequent sequential patterns generated by the mining process.In this paper,the problem of compressing frequent sequential patterns was studied.Inspired by the ideas of compressing frequent itemsets,an algorithm,CFSP (Compressing Frequent Sequential Patterns),was developed to mine a few representative sequential patterns to express all information of all frequent sequential patterns and eliminate a large number of redundant sequential patterns.The CFSP adopted a two-steps approach:in the first step,all closed sequential patterns as the candidate set of representative sequential patterns was obtained,and at the same time the most of representative sequential patterns were gotten;in the second step,the time on finding the remaining representative sequential patterns is only a little.An empirical study with both real and synthetic data sets proves that the CFSP has good performance.

【基金】 国家“九七三”重点基础研究发展规划基金项目(2005CB321902);北京市教委科技计划基金项目(KM200911417003)~~
  • 【会议录名称】 第26届中国数据库学术会议论文集(A辑)
  • 【会议名称】第26届中国数据库学术会议
  • 【会议时间】2009-10-15
  • 【会议地点】中国江西南昌
  • 【分类号】TP311.13
  • 【主办单位】中国计算机学会数据库专业委员会
节点文献中: 

本文链接的文献网络图示:

本文的引文网络