节点文献

一种高性能分布式Web Crawler的设计与实现

Design and Implementation of a Distributed High-Performance Web Crawler

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 张岭叶允明宋晖于水马范援

【Author】 ZHANG Ling,YE Yun-ming,SONG Hui,YU Shui,MA Fan-yuan(Dept. of Computer Science and Eng., Shanghai Jiaotong Univ., Shanghai 200030, China)

【机构】 上海交通大学计算机科学与工程系上海交通大学计算机科学与工程系 上海200030上海200030上海200030

【摘要】 介绍了一种大规模、高性能、分布式的Web信息搜集器的设计及其Java实现.提出了Crawler设计中数据结构、系统功能模块和相关算法新的设计思想;对设计与实现过程中需要解决的关键问题分布式协调机制、基于内存的URL存储管理等进行了讨论,并提供了现阶段的设计、实现方法和分布式无损链接分析算法.

【Abstract】 Web crawler is the core component of WWW search engine and information retrieval systems. This paper discussed the architecture of a distributed Web crawler and the design ideas about the Web crawler data structure, system modules and related algorithms. The key problems encountered in the design and implementations were also commented, and the solutions to those problems were presented.

【关键词】 Web信息搜集器分布式系统搜索引擎
【Key words】 Web crawlerdistributed systemsearch engineJava
【基金】 上海市科委重点基础研究项目(02DJ14045)
  • 【文献出处】 上海交通大学学报 ,Journal of Shanghai Jiaotong University , 编辑部邮箱 ,2004年01期
  • 【分类号】TP393.09
  • 【被引频次】29
  • 【下载频次】367
节点文献中: 

本文链接的文献网络图示:

本文的引文网络