节点文献
一种全自动生成网页信息抽取Wrapper的方法
Fully Automatic Wrapper Generation for Web Information Extraction
【摘要】 Web网页信息抽取是近年来广泛关注的话题。如何最快最准地从大量Web网页中获取主要数据成为该领域的一个研究重点。文章中提出了一种全自动化生成网页信息抽取Wrapper的方法。该方法充分利用网页设计模版的结构化、层次化特点,运用网页链接分类算法和网页结构分离算法,抽取出网页中各个信息单元,并输出相应Wrapper。利用Wrapper能够对同类网页自动地进行信息抽取。实验结果表明,该方法同时实现了对网页中严格的结构化信息和松散的结构化信息的自动化抽取,抽取结果达到非常高的准确率。
【Abstract】 Web information extraction has been a hot topic in recent years.The challenge is how to extract important information from a large number of web pages as quickly and accurately as it can.In this paper a novel method is proposed for fully automatic wrapper generation for Web information extraction.This method makes use of structure of Web templates abundantly.It uses Web Page Link_Sort algorithm and Web Page Structure_Seperator algorithm to extract information from Web pages and output a wrapper accordingly.Experimental results showed that this method performs well in both rigidly and loosely structured records in Web pages.
【Key words】 computer application; Chinese information processing; Web information extraction; Web structure seperator; wrapper;
- 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2008年01期
- 【分类号】TP391.1
- 【被引频次】57
- 【下载频次】1068