节点文献
基于迭代事务集与交集剪枝的最大频繁项集挖掘算法
An Algorithm for Mining Maximal Frequent Itemsets Based on Datasets Iteration and Intersection Pruning
【摘要】 挖掘最大频繁项目集是多种数据挖掘应用中的关键问题,如果采用Apriori类的候选项目集生成-检验方法,则候选项目集生成的代价通常很高。为寻求避免生成大量候选项集或生成频繁模式树的挖掘算法,提出一种从事务项集交集求最大频繁项集的迭代算法DIIP(Datasets Iteration and Intersection Pruning Algorithm),通过不断缩减事务集数据量和尽可能早地对项目集进行修剪实现最大频繁项集的挖掘,该算法有别于已有的最大频繁项集经典算法,实验表明该算法有效可行。
【Abstract】 Mining frequent itemsets is a key issue in data mining applications,if an algorithm uses Apriori-like candidate itemsets generating-testing approach,the generation process is usually costly.To seek for a method that can avoid the generating of vast volume of candidate itemsets or frequent pattern trees,an iterative algorithm is proposed to discover maximal frequent itemsets from itemsets intersections(DIIP),it condenses the volume of datasets continuously and performs itemsets pruning as early as possible.It differs from classical MFI discovering algorithms.Experiments show that this algorithm is valid and mildly efficient.
【Key words】 data mining; maximal frequent itemsets; candidate itemsets; intersection pruning; iteration;
- 【文献出处】 南开大学学报(自然科学版) ,Acta Scientiarum Naturalium Universitatis Nankaiensis , 编辑部邮箱 ,2009年04期
- 【分类号】TP311.13
- 【被引频次】6
- 【下载频次】77