节点文献

Web使用挖掘若干关键问题研究

Study on Some Key Issues of Web Usage Mining

【作者】 阮备军

【导师】 朱扬勇;

【作者基本信息】 复旦大学 , 计算机软件与理论, 2004, 博士

【摘要】 Web使用挖掘(Web Usage Mining)是应用数据挖掘技术从Web数据中发现使用模式的过程。Web提供了一种不受时空限制的人机交互界面,为大规模记录,收集,分析和抽取用户行为信息提供了巨大的技术发展空间。在此背景下,Web使用挖掘研究得到了学术界和工业界的广泛关注,由此衍生的技术大量应用在科学研究,软件设计以及商业智能等领域。 本文总结了目前Web使用挖掘研究的现状,对其中存在的一些问题作了深入的研究和探讨。这些问题分别涉及频繁序列模式挖掘,Web用户行为特征相似性/差别的量化方法,以及支持Web站点设计优化的数据挖掘技术。 本文的主要贡献如下: (1)提出了一个称为TD-WAP-Mine的频繁序列模式挖掘算法。和已有的算法相比,它采用了新的频繁模式搜索策略,大幅度减少了在构造中间数据方面的工作量。大量的实验结果表明此算法在运行速度方面好于原有的算法,特别适合用在需要挖掘大量频繁模式的场合。 (2)提出了一种使用Web结构数据所蕴涵的语义信息量化使用行为特征差别的方法。与已有的研究相比,特征项之间的关系表示结构从有向根树扩展到了有向无环图。基于核心概念“最大相似宽度”,此方法为量化使用行为特征在语义上的差别定义了一组距离函数。在关系表示结构是有向根树的条件下,这些距离函数均满足三角不等式特性,在提高搜索效率方面具有优势,弥补了以往研究存在的缺陷。实验初步表明此类距离函数在最近邻查询效果和计算速度方面可与已有研究媲美。 (3)提出了一种新的支持站点设计优化的Web使用挖掘方案。此方案基于历史搜寻路径统计用户寻找目标花费的平均时间,用以量化Web页面的搜寻费用。在此基础上提出了一种高效的数据挖掘方法,寻找一组能够有效压缩搜寻路径(降低搜寻费用)的超链接。实验表明挖掘的结果能够提供许多有用的信息,帮助管理者及时发现站点设计中存在的问题。

【Abstract】 Web Usage Mining (WUM) is the process of applying data mining techniques to the discovery of usage patterns from Web data. As a kind of human-computer interface that can be used anywhere and anytime, the Web offers lots of opportunities for developing techniques to record, collect, analyze and extract the usage information on a large-scale level. In this context, WUM; attracts enormous interest from both the academic and industrial communities. The WUM techniques have broad applications in science study, softeware design and business intelligence.In this thesis, an up-to-date survey of the WUM research is given and the results of the study on some key issues of WUM are presented. The investigated issues are related with frequent sequential-pattern mining, methods for measuring the differences between two user behaviors, and the data mining techniques for optimizing web-site design.The main contributions are as follows.(1) A frequent sequential-pattern algorithm, called TD-WAP-Mine, is proposed. It differs from the previous algorithms in applying a new frequent-pattern searching strategy, which greatly reduces the workload of building intermediate data. The experimental results for various datasets show this algorithm performs better than the previous ones, especially when the datasets contain prolific frequent patterns.(2) A new method is proposed for measuring the difference between two user behaviors based on the semantic information contained in the Web structure data. The presentation structure of the relationship between feature items is formalized as a special kind of directed acyclic graph. Based on a core concept, called maximum similarity width, serveral distance functions are defined for quantifying the difference between two user behaviors in terms of the set of feature items. When the presentation structure of relationship is a rooted directed tree, these distance functions satisfy the property of triangle inequality. The property of triangle inequality is very useful for making searching more efficient, but the investigation on this property is lacking in the previous research. The results of preliminary experiments show these new functions can perform similarity with the previous ones in computing speed and nearest neighbor searching.(3) A n ew d ata m ining m ethod i s p roposed for optimizing Web s ite d esign. Itmeasures the searching cost of Web pages by computing the average searching time based on the details of the information foraging paths. Moreover, a kind of efficient data mining method is proposed to discovery a group of hyperlinks that are useful for reducing the searching time. The experimental results show the mining results can provide useful information for identifying problems in Web site design.

  • 【网络出版投稿人】 复旦大学
  • 【网络出版年期】2005年 01期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络