节点文献

基于数据仓库的关联规则挖掘算法的研究与应用

Research and Application of Algorithem of Mining Association Rule Based on Data Warehouse

【作者】 何胜文

【导师】 周绍梅;

【作者基本信息】 南昌大学 , 计算机软件与理论, 2007, 硕士

【摘要】 在过去的数十年中,无论是商业企业、科研机构或者政府部门,MIS系统(Management Information System,管理信息系统)都被广泛地应用在信息管理上。以事务处理为主的MIS系统在方便数据管理的同时,也积累了海量的、十分繁杂的数据。由于爆炸性增长的数据量与相对贫乏的知识之间的矛盾,使得数据挖掘成为目前国际上信息决策和人工智能领域的最前沿研究方向之一,其中发现大量数据中项集的相关联系的关联规则(Association Rule)挖掘是它的一个重要方向。在本文中,对经典的关联规则算法进行了深入的分析和研究,并对原算法存在的不足之处,提出了一些改进方法,取得了一定的效果,研究内容主要包括;(1)改变了经典算法的单向搜索方法,运用了自顶向下和自底向上相结合的搜索策略。无论项目集数目多少或是最小支持度大小,都能够较快的找到频繁集,实验证明能提高算法的效率。(2)利用自底向上生成的非频繁项目集合来指导自顶向下的降维操作,可减少自顶向下的侯选频繁集的数量。(3)利用矩阵的结构保存事务数据库,减少计算机的I/O操作,利用事务数据库压缩存储的性质对数据进行裁减,提高了对数据遍历的效率。(4)研究了数值属性关联规则的挖掘算法,利用聚类算法来划分区间,然后将划分后的区间映射为布尔属性,最后发现用户感兴趣的关联规则。最后,将理论知识与实践相结合,以学生成绩分析数据仓库为数据源,进行关联规则的挖掘,同时以OLAP等工具对数据进行多维展现,实现数据分析的可视化。

【Abstract】 In the past several ten years, both commercial enterprises and scientific research institutions or government departments, the MIS system has been widely used in information management. The MIS system which is based on affair processing accumulates a massive and complicated data. Meanwhile it is also convenient for data management. The contradiction between the explosive growth of data and comparative lack of knowledge makes data mining become one of the foremost frontier directions of research in regions of international information-making and artificial intelligence. Among the research, the exploration indicates that association rule is an important direction of data mining.This thesis has deeply analysis and research on the classical association rule of algorithm and provides some improvements to the demerits of original algorithm. Fortunately it achieved some certain effect, the research includes:(1) Changing the classical algorithm for one-way search methods, and use a top-down and bottom-up two-way search strategy. Both items that are set for the number or size of the smallest support, can quickly find Frequent Item sets, experiments improved that it can greatly increases the efficiency of algorithms.(2) Using the strategy of bottom-up to generate infrequent item sets for dimension reduction of top-down operation can greatly reduced the number of frequent top-down candidates set.(3) Using the structure of the matrix to preserve database can reduce computer I/O operation, using the database storage and compress function to reduce the data, will improve the efficiency of data traversing.(4) Research the data mining of quantitative association rule, using clustering algorithm, to classify the space and then mapped those spaces to Boolean attribution. After that, the research discovers association rule which is interested by the user.Eventually, combine the theory with practice, use the Data Warehouse of student’s score as the source of the data, mining the association rules, meanwhile use OLAP(Online Analytical Processing) and other facilities to show that in many dimensions, then realizing the visualization of the data analysis.

  • 【网络出版投稿人】 南昌大学
  • 【网络出版年期】2008年 08期
  • 【分类号】TP311.13
  • 【被引频次】3
  • 【下载频次】263
节点文献中: 

本文链接的文献网络图示:

本文的引文网络