节点文献

基于图结构的候选序列生成算法

Graph-Based Candidate Frequent Patterns Generating Algorithm

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 郭平刘潭仁

【Author】 QUO Ping LIU Tan-Ren (College of Computer Science,Chongqing University,Chongqing,400044)

【机构】 重庆大学计算机学院重庆大学计算机学院 重庆 400044重庆 400044

【摘要】 先生成候选序列再判断候选序列是否为频繁序列,最后获得频繁序列是序列数据挖掘中基于候选序列挖掘算法的一般结构,如Apriori类算法,GSP算法,SPADE算法等。因此,研究候选序列生成算法具有普遍意义。本文首先研究了序列数据集(序列数据库)与图结构间的关系,证明了一个序列是频繁序列的必要条件是该序列对应于一个完全子图。以此为基础提出了基于图结构的候选序列生成算法,文中给出了算法正确性证明。在T25I10D10K和T25I20D100K数据集上的挖掘实验表明在本文提出的候选序列生成算法上进行挖掘比用Apriori算法进行挖掘的效率更高。

【Abstract】 In candidate-sequence-based mining algorithms, the common procedure is first Generating the candidate frequent patterns, then identifying the frequent patterns based on the candidate frequent patterns, last getting the frequent patterns, such as the apriori-like algorithm, GSP algorithm, SPADE algorithm and so on. Thus there is universal meaning to research the candidate frequent patterns generating algorithm- In this article, firstly we investigate the relationship between sequence data set (sequence database)and graph structure and prove that the requirement of a sequence to be a frequent sequence is that there is a corresponding complete subgraph to the sequence in the graph. Then a graph-based candidate frequent patterns generating algorithm is proposed and the correctness of the algorithm is proved in the article. Lastly the algorithm is applied to T25I10D10K and T25I20D100K data set and comparing with the apriori algorithm it higher the efficiency of sequence data mining.

【基金】 国家十五攻关项目(编号:2002BA107B)
  • 【文献出处】 计算机科学 ,Computer Science , 编辑部邮箱 ,2004年01期
  • 【分类号】TP311.13
  • 【被引频次】12
  • 【下载频次】135
节点文献中: 

本文链接的文献网络图示:

本文的引文网络