节点文献

基于粗糙集的Web日志挖掘研究

The Research of Web Log Mining Based on Rough Set

【作者】 杨志勇

【导师】 张永;

【作者基本信息】 兰州理工大学 , 计算机应用技术, 2006, 硕士

【摘要】 随着Internet的迅猛发展,使得World Wide Web已经深入到社会生活的方方面面。Web已经发展成为拥有数十亿页面,蕴涵着具有巨大潜在价值的分布式信息空间。人们迫切需要从这些海量的数据中查找出对自己有用的信息,对数据挖掘研究提出了新的挑战。Web数据挖掘是一种将传统数据挖掘与Web结合起来的技术,它将随着Internet的发展越来越受到各方面的关注。 本文工作主要包括以下几个方面: 首先,Web日志挖掘数据预处理。Web数据复杂多样,首先需要确定研究对象,Web日志挖掘的对象不是网络上的原始数据而是从用户和网络交互过程中抽取出来的第二手数据,它包括所请求的URL、发出请求的IP地址和时间戳等,这些数据提供了有关用户访问的丰富信息。本文在这部分的研究重点是如何提取有关用户访问的特征(如用户的访问行为、频度、内容等),以及建立基于用户访问行为的数据模型。 其次,基于粗糙集理论的Web日志挖掘。以前的方法对Web同志数据库中潜在信息的挖掘采用先将数据组织成传统的数掘挖掘方法能够处理的数据模型,然后用数据挖掘算法(如关联规则算法等)进行处理。这种方法虽然暂时解决了Web挖掘的需求但是对于Web数据库来说不能满足其动态增长的需要。在粗糙集理论中,知识被看成是一种分类能力,即在域上构造分区的能力。本文在基于粗糙集理论的思想上对预处理后的数据进行离散化,并建立了一种新的数据模型,最后改进约简算法并约简提取出稳定的分类规则。同时考虑到不一致规则的存在,还研究了缺省情况下如何获得决策规则。 最后,对本论文的内容进行了总结,并对下一步日志挖掘研究进行了展望。

【Abstract】 Accompanying with the quick development of the Internet, World Wide Web has already been related to every aspects of social life. Web has developed to be a distributed information space which owns billions of websites and contains knowledge of great and potential value. That people want look for the useful information they need in the rich database, provides new challenge for the research on data mining. Web data mining is a kind of technique combining traditional data mining and Web data, which will gain more attention from all the aspects along with the development of the Internet.The paper mainly includes several aspects:Firstly, the data pretreatment of Web log mining. The Web data is complex and various. First we must determine the research object. The object of Web log mining is not the original data on the Web but the secondhand data abstracted from the interactive process of users and the Internet, which includes the appealed URL, the appealing IP and the time stab, etc. All these logs offer rich information about users’ visits. The research focal point of this part in the paper is how to get the characteristics of the visits (such as behaviors, frequency, content, etc. of users’ visits) and to establish the data model based on the behaviors of users’ visits.Secondly, the research of Rough set. The former way of mining the potential information in the Web Log database is to transform the data into a data model which can be manipulated by the traditional data mining and then manipulate them by data mining technique (such as the algorithm of association rule). Although this way meets the needs of Web mining temporarily, it can not satisfy its dynamic increasing demand. In Rough set, knowledge is considered as an ability of classification, which is the ability of constructing partition in the domain. According to the thoughts of Rough set, the paper researches the dispersiveness of pretreated data, sets up a new data model and improves reduction algorithm and abstracts the static rule of classification ultimately. At the same time, it takes into account the existence of the incoherence rule, and it researches on how to achieve the decision rule on the absent occasion.Finally, the paper makes a conclusion and opens a new prospect to the next step of log mining.

  • 【分类号】TP393.09;TP311.13
  • 【下载频次】180
节点文献中: 

本文链接的文献网络图示:

本文的引文网络