节点文献

聚焦爬虫技术研究综述

Survey on the research of focused crawling technique

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 周立柱林玲

【Author】 ZHOU Li-zhu,LIN Ling(Department of Computer Science and Technology,Tsinghua University,Beijing 10084,China)

【机构】 清华大学计算机科学与技术系清华大学计算机科学与技术系 北京100084北京100084

【摘要】 因特网的迅速发展对万维网信息的查找与发现提出了巨大的挑战。对于大多用户提出的与主题或领域相关的查询需求,传统的通用搜索引擎往往不能提供令人满意的结果网页。为了克服通用搜索引擎的以上不足,提出了面向主题的聚焦爬虫的研究。至今,聚焦爬虫已成为有关万维网的研究热点之一。文中对这一热点研究进行综述,给出聚焦爬虫(Focused Crawler)的基本概念,概述其工作原理;并根据研究的发展现状,对聚焦爬虫的关键技术(抓取目标描述,网页分析算法和网页搜索策略等)作系统介绍和深入分析。在此基础上,提出聚焦爬虫今后的一些研究方向,包括面向数据分析和挖掘的爬虫技术研究,主题的描述与定义,相关资源的发现,W eb数据清洗,以及搜索空间的扩展等。

【Abstract】 The survey of focused crawling starts with the motivation for this new research and an introduction on basic concepts of focused crawling.The key issues in focused crawling are reviewed,such as webpage analyzing algorithms and the searching strategy on the Web.How to crawl relevant data and information according to different requirements is discussed in detail and three representative architectures of focused crawler systems are analyzed.Some future works for focused crawling research are indicated,including crawling for data analysis and data mining,topic description,finding relevant Web pages,Web data cleaning,and the extension of search space.

【基金】 国家自然科学基金资助项目(60173008)
  • 【文献出处】 计算机应用 ,Computer Applications , 编辑部邮箱 ,2005年09期
  • 【分类号】TP393.02
  • 【被引频次】552
  • 【下载频次】6421
节点文献中: 

本文链接的文献网络图示:

本文的引文网络