节点文献
基于标记树对象抽取技术的Hidden Web获取研究
Research on the Hidden Web Retrieval with Tag-Tree-based Object Extraction Technique
【摘要】 目前标准的搜索引擎能够检索的仅仅是WorldWideWeb提供的小部分称为可索引的Web信息。大量的HiddenWeb信息(估计容量是可索引Web的500倍)对这些搜索引擎是不可见的。这些信息隐藏在Web页面的搜索表单后面,保存在大型的动态数据库中。该文提出了一套检索HiddenWeb信息的方法,给出了系统的框架结构,并详细讨论了实现的关键技术。系统采用新的基于标记树的对象抽取(Tag-Tree-basedObjectExtraction)方法自动地从Web页面中抽取HiddenWeb信息,然后在此基础上给出了结构化的HiddenWeb信息查询算法。文章最后对实验结果进行了讨论。
【Abstract】 Current traditional search engines retrieve only a small portion of World Wide Web,which called the publicly indexable Web.In particular,they ignore the tremendous amount information″hidden″behind search forms ,in large searchable electronic databases.The size of hidden Web is about 500times larger than the publicly indexable Web.This paper addresses this problem of designing a system for extracting and retrieving hidden Web information.It presents a generic operational model of the hidden Web information retrieval and describes the key techniques.It also introduces a new Tag-Tree-based Object Extraction Technique for automatically extracting hidden Web information from web pages.Based on this technique,the retrieval algorithm for structured query of hidden Web information is implemented.At last the test results are reported.
【Key words】 Hidden Web; Information Retrieval; Object Extraction; Structured Query; Tag Tree;
- 【文献出处】 计算机工程与应用 ,Computer Engineering and Applications , 编辑部邮箱 ,2002年23期
- 【分类号】TP391.4
- 【被引频次】50
- 【下载频次】188