节点文献

一种全自动生成网页信息抽取Wrapper的方法

Fully Automatic Wrapper Generation for Web Information Extraction

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 梅雪程学旗郭岩张刚丁国栋

【Author】 MEI Xue~1,2,CHENG Xue-qi~1,GUO Yan~1,ZHANG Gang~1,DING Guo-dong~1(1.Institute of Computing Technology,Chinese Academy of Sciences,Beijing 100080,China;2.Graduate University of Chinese Academy of Sciences,Beijing 100049,China)

【机构】 中国科学院计算技术研究所中国科学院计算技术研究所 北京100080中国科学院研究生院北京100049北京100080

【摘要】 Web网页信息抽取是近年来广泛关注的话题。如何最快最准地从大量Web网页中获取主要数据成为该领域的一个研究重点。文章中提出了一种全自动化生成网页信息抽取Wrapper的方法。该方法充分利用网页设计模版的结构化、层次化特点,运用网页链接分类算法和网页结构分离算法,抽取出网页中各个信息单元,并输出相应Wrapper。利用Wrapper能够对同类网页自动地进行信息抽取。实验结果表明,该方法同时实现了对网页中严格的结构化信息和松散的结构化信息的自动化抽取,抽取结果达到非常高的准确率。

【Abstract】 Web information extraction has been a hot topic in recent years.The challenge is how to extract important information from a large number of web pages as quickly and accurately as it can.In this paper a novel method is proposed for fully automatic wrapper generation for Web information extraction.This method makes use of structure of Web templates abundantly.It uses Web Page Link_Sort algorithm and Web Page Structure_Seperator algorithm to extract information from Web pages and output a wrapper accordingly.Experimental results showed that this method performs well in both rigidly and loosely structured records in Web pages.

【基金】 国家高技术研究发展计划(863)资助项目(2005AA142110)
  • 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2008年01期
  • 【分类号】TP391.1
  • 【被引频次】57
  • 【下载频次】1068
节点文献中: 

本文链接的文献网络图示:

本文的引文网络