节点文献
主题Web信息采集的研究与设计
Research and Design of Focused Web Crawler
【作者】 李盛韬; 吴丽辉; 于满泉; 潘文锋; 余智华; 王斌; 程学旗;
【Author】 Li Shengtao Wu Lihui Yu Manquan Pan Wenfeng Yu Zhihua Wang Bin Cheng XueqiSoftware Division, Institute of Computing Technology, Chinese Academy of Sciences, Beijing 100080
【机构】 中国科学院计算技术研究所软件研究室;
【摘要】 主题Web信息采集是信息检索领域内一个将采集技术与过滤方法结合的新兴方向,也是信息处理技术中的一个研究热点。本文分析了主题Web信息采集的基本问题,提出了难点以及相关的解决方案,并在此基础上设计了“天达”主题Web信息采集系统。
【Abstract】 Focused web crawling is a new crawling direction in the field of information retrieval which is combined with filtering methods. And it also is a research hotspot in the information processing technologies. This paper argues the principles, difficulties and measures of the focused web crawler, and then detailedly analyses the design of our SkyReach focused web crawler.
【Key words】 Web Crawler; Information Retrieval; Information Processing; Focused Crawler;
- 【会议录名称】 语言计算与基于内容的文本处理——全国第七届计算语言学联合学术会议论文集
- 【会议名称】全国第七届计算语言学联合学术会议
- 【会议时间】2003-08
- 【会议地点】中国哈尔滨
- 【分类号】TP393.092
- 【主办单位】哈尔滨工业大学计算机科学与技术学院、清华大学智能技术与系统国家重点实验室