节点文献

网页文本信息自动提取技术综述

Survey on text information extraction from Web page

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 张俊英胡侠卜佳俊

【Author】 ZHANG Jun-ying1,HU Xia2,BU Jia-jun1(1.College of Computer Science & Technology,Zhejiang University,Hangzhou 310027,China;2.Hangzhou Science & Technology Information Institute,Hangzhou 310000,China)

【机构】 浙江大学计算机学院杭州市科技信息研究院

【摘要】 对W eb网页文本信息自动提取技术提供了一个较为全面的综述。通过分析在这个领域常用到的三种信息提取模型和四类机器学习算法的发展,较为全面地阐述了当前主流的网页文本信息自动提取技术,对比了各种方法的应用范围,最后对于该领域当前的热点问题和发展趋势进行了展望。

【Abstract】 This paper supplied a comprehensive survey of the text information extraction from Web page.By presenting and analyzing the development of three kinds of extraction modules and four types of the learning algorithms used in this area,it comprehensively surveyed the relative technologies of the text information extraction from Web page,and analyzed the application scenarios of different technologies.Finally,discussed the difficulties and the trend of the development of this area.

【关键词】 信息提取机器学习网页
【Key words】 information extractioncomputer learningWeb page
【基金】 浙江省科技资助项目(2007C23086)
  • 【文献出处】 计算机应用研究 ,Application Research of Computers , 编辑部邮箱 ,2009年08期
  • 【分类号】TP391.41
  • 【被引频次】34
  • 【下载频次】823
节点文献中: 

本文链接的文献网络图示:

本文的引文网络