节点文献

主题Web信息采集的研究与设计

Research and Design of Focused Web Crawler

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 李盛韬吴丽辉于满泉潘文锋余智华王斌程学旗

【Author】 Li Shengtao Wu Lihui Yu Manquan Pan Wenfeng Yu Zhihua Wang Bin Cheng XueqiSoftware Division, Institute of Computing Technology, Chinese Academy of Sciences, Beijing 100080

【机构】 中国科学院计算技术研究所软件研究室

【摘要】 主题Web信息采集是信息检索领域内一个将采集技术与过滤方法结合的新兴方向,也是信息处理技术中的一个研究热点。本文分析了主题Web信息采集的基本问题,提出了难点以及相关的解决方案,并在此基础上设计了“天达”主题Web信息采集系统。

【Abstract】 Focused web crawling is a new crawling direction in the field of information retrieval which is combined with filtering methods. And it also is a research hotspot in the information processing technologies. This paper argues the principles, difficulties and measures of the focused web crawler, and then detailedly analyses the design of our SkyReach focused web crawler.

【基金】 国家973项目(G1998030413);中科院计算所领域前沿青年基金(20016280-8)资助
  • 【会议录名称】 语言计算与基于内容的文本处理——全国第七届计算语言学联合学术会议论文集
  • 【会议名称】全国第七届计算语言学联合学术会议
  • 【会议时间】2003-08
  • 【会议地点】中国哈尔滨
  • 【分类号】TP393.092
  • 【主办单位】哈尔滨工业大学计算机科学与技术学院、清华大学智能技术与系统国家重点实验室
节点文献中: 

本文链接的文献网络图示:

本文的引文网络