节点文献

搜索引擎查询日志的聚类

Clustering of Search Engine Query Log

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 张玉连李彦威王权原福永

【Author】 ZHANG Yu-lian, LI Yan-wei, WANG Quan, YUAN Fu-yong (College of Information Science and Engineering, Yanshan University, Qinhuangdao 066004)

【机构】 燕山大学信息科学与工程学院

【摘要】 随着搜索引擎技术和网络数据挖掘技术的发展,怎样从搜索引擎查询日志中找到有用的信息成为研究热点。该文在讨论Beeferman提出的算法及Chan对其改进的算法的优缺点后,提出一个基于用户网页兴趣度的改进算法。该算法能进一步减小噪声数据的影响,并通过模拟实验对3种不同的算法进行了对比。

【Abstract】 In recent years, with the search engine technology and the network data mining technology development, how to find the useful information from the search engine query log becomes an important research direction. This paper discusses the excellences and the disadvantages of the clustering algorithm proposed by Beeferman and the improved algorithm which is proposed by Chan. A new improved algorithm based on the user profile of the Webpage is proposed that can weaken the influence of the noises data. And the simulation experiment proves that the new algorithm is better than the Beeferman algorithm and the Chan algorithm.

  • 【文献出处】 计算机工程 ,Computer Engineering , 编辑部邮箱 ,2009年01期
  • 【分类号】TP391.3
  • 【被引频次】13
  • 【下载频次】335
节点文献中: