节点文献

Deep Web爬虫研究与设计

On the research and design of deep web crawler

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 郑冬冬赵朋朋崔志明

【Author】 ZHENG Dongdong, ZHAO Pengpeng, CUI ZhimingDepartment of Computer Science and Technology, Soochou University, Suzhou 215006, China

【机构】 苏州大学计算机科学与技术系苏州大学计算机科学与技术系 苏州215006苏州215006苏州215006

【摘要】 随着W eb的发展,越来越多的数据可以通过表单提交来获取,这些表单提交所产生信息是由D eep W eb后台数据库动态产生的。在这种情况下,信息集成就更加需要W eb爬虫来自动获取这些页面以进一步地处理数据。为了帮助用户完成这样的任务,提出一种用于搜集D eep W eb页面的爬虫的设计方法。此方法使用一个预定义的领域本体知识库来识别这些页面的内容,同时利用一些来自W eb站点的导航模式来识别自动填写表单时所需进行的路径导航。通过对来自不同领域的D eep W eb站点的大量实验,验证了此方法是非常有效的。

【Abstract】 As the web grows, more and more data has become available under dynamic forms of publication, such as legacy databases accessed by an HTML form. In this way, integration of this data relies more and more on the Web Crawler that can automatically fetch pages for further processing. As a result, there is an increasing need for tools that can help users generate such agents. The method is described for automatically generating agents to collect Deep Web pages. This method uses a pre-defined ontology repository for identifying the contents of these pages and takes the advantage of some patterns that can be found among web sites to identify the navigation paths to follow. The results of a number of experiments carried out with sites from different domains demonstrate the accuracy of the method.

【基金】 Deep Web关键技术研究
  • 【文献出处】 清华大学学报(自然科学版) ,Journal of Tsinghua University(Science and Technology) , 编辑部邮箱 ,2005年S1期
  • 【分类号】TP393.09;
  • 【被引频次】116
  • 【下载频次】1588
节点文献中: 

本文链接的文献网络图示:

本文的引文网络