节点文献
一种改进的LIPI数据挖掘算法的仿真分析
Simulation Analysis of an Improved LIPI Data Mining Algorithm
【摘要】 在传统LIPI数据挖掘算法中,需要反复扫描投影数据库寻找局部频繁项并重复构造大量重复投影,造成数据挖掘耗时,效率低下的不足。为了提高算法的计算速度,提出改进的LIPI数据挖掘算法。算法借助连接2-序列位置信息表(LIPI)找到序列模式的下一项,完成K-1序列位置信息与2-序列位置信息的连接,实现序列模式放缩式增长,得出K-序列与K-序列相应的位置信息数据,避免对投影数据库反复扫描;引入了BIDE算法的前后向剪枝策略,检查相同末项序列位置信息表进行前向剪枝,消除大量重复投影的构建,提高挖掘算法的效率。实验结果表明,改进后的算法能快速的寻找到局部频繁项,有效提高了数据挖掘的效率。
【Abstract】 In the traditional LIPI data mining algorithm,it needs to scan projection databaserepeatedly,to look for local frequent items,and to repeatedly construct a large number of repeat projection,which results in time- consuming and inefficiency of data mining. To solve this problem,this paper presented an improved LIPI data mining algorithm. Firstly,the next item of sequence mode is found by means of connection 2- the serial position information table( LIPI) in the algorithm,to complete the connection of K- 1 sequence position and 2 sequence position information,achieve the scaling type growth of sequential pattern,and get theK – sequence and its corresponding position information data,which can avoid to scan projection databaserepeatedly. Then,BIDE algorithm is introducedinto the forward- backward pruning strategy,and the same last item of the sequence information table is checked for the forward pruning to eliminatea large number of constructions of repeat projection,and improve the efficiency of mining algorithm. The experimental results show that the improved algorithm can find local frequent items quicklyand improve the efficiency of data mining effectively.
【Key words】 Scaling type growth; Sequential pattern mining(SPM); Location information; Projection database; Frequent prefix;
- 【文献出处】 计算机仿真 ,Computer Simulation , 编辑部邮箱 ,2014年08期
- 【分类号】TP311.13
- 【被引频次】2
- 【下载频次】96