节点文献

Web信息抽取技术综述

Survey of Web information extraction technologies

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 陈钊张冬梅

【Author】 CHEN Zhao,ZHANG Dong-mei ( School of Information Science & Technology,Beijing Forestry University,Beijing 100083,China)

【机构】 北京林业大学信息学院

【摘要】 快速高效地获取网页主题信息的需求使得Web信息抽取技术成为信息技术领域的研究热点。现有的Web信息抽取技术大致可以归纳为基于统计理论的、基于视觉特征的、基于DOM树结构的和基于模板的几类。由于网页文本本身具有树结构并且具有一定的相似性,基于DOM树结构和基于模板的抽取技术发展很快而且已经得到了广泛的应用。分别论述了上述几类技术在近几年来的研究进展,从自动化程度、适用范围和复杂性三个角度分析对比了几类技术的优缺点。

【Abstract】 Web information extraction technology has been made the focus of the field of information technology by the needs of obtaining the topic contents of Web pages more efficiently. Existing technologies of this field could be classified into the following four categories,statistics based technology,vision based technology,DOM tree based technology and template based technology. The DOM tree based technology and template based technology had gained a rapid development and a wide employment because of the special structure and similarity owned by Web pages. This paper made a detailed survey and analysis of the above four technologies as well as the comparison of their advantages and disadvantages from points of automation,application filed and complexity.

【基金】 中央高校基本科研业务费专项资金资助项目(BLYX200928)
  • 【文献出处】 计算机应用研究 ,Application Research of Computers , 编辑部邮箱 ,2010年12期
  • 【分类号】TP393.09
  • 【被引频次】116
  • 【下载频次】1585
节点文献中: 

本文链接的文献网络图示:

本文的引文网络