节点文献

挖掘数值型数据流中的最大频繁模式

Mining Maximal Frequent Patterns over Numerical Data Streams

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 高晶李建中张兆功

【Author】 GAO Jing~1,LI Jian-Zhong~(1,2),and ZHANG Zhao-Gong~(1,2) 1(School of Computer Science and Technology,Harbin Institute of Technoloy,Harbin 150001) 2(School of Computer Science and Technology,Heilongjiang University,Harbin 150080)

【机构】 哈尔滨工业大学计算机科学与技术学院黑龙江大学计算机科学技术学院

【摘要】 近年来,很多应用中产生了大量的、源源不断的数据,即数据流.如何对数据流进行分析和挖掘,从而得到有用的知识已成为热点问题.挖掘频繁模式是常用的挖掘方法,已有的数据流频繁模式挖掘算法集中在挖掘传统交易数据,这些方法不能用于属性是数值型的数据流上.针对这一问题,提出了一种挖掘数值型数据流中最大频繁模式的有效方法.采用基于距离的方法将数据离散化,并重新定义了最大频繁模式的概念,并在此基础上设计了一种新的算法.该算法用聚类的方法产生频繁项,通过增量更新及时快速地输出最大频繁模式.实验结果证明了该算法的有效性.

【Abstract】 In recent years,many applications generate huge-volume,continuously-arriving data,i.e.data streams.It has become a hot issue to analyze and mine data streams to obtain valuable knowledge.Frequent patterns mining is a central method,but now the algorithms of mining frequent patterns over data streams focus on mining from the traditional transaction data,which can not fit in data streams with numerical attributes.To solve the problem,an efficient method is presented to mine maximal frequent patterns over numerical data streams.The data are discretized using a distance-based method and the concept of maximal frequent patterns are redefined.Based on the definition,a new algorithm is designed,which finds frequent items by clustering and outputs the maximal frequent patterns in a timely and quick manner by incremental update.Experimental results show its effectiveness and efficiency.

【基金】 国家“九七三”重点基础研究发展计划基金项目(G1999032704);国家自然科学基金项目(60273082);哈尔滨市科学研究基金项目
  • 【会议录名称】 第二十一届中国数据库学术会议论文集(研究报告篇)
  • 【会议名称】第二十一届中国数据库学术会议
  • 【会议时间】2004-10-14
  • 【会议地点】中国福建厦门
  • 【分类号】TP311.13
  • 【主办单位】中国计算机学会数据库专业委员会
节点文献中: 

本文链接的文献网络图示:

本文的引文网络