节点文献

基于视觉特征和领域本体的Web信息抽取

Visual Features and Domain Ontology-Based Web Information Extraction

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 张鑫陈梅王翰虎王嫣然

【Author】 ZHANG Xin,CHEN Mei,WANG Han-hu,WANG Yan-ran(College of Computer Science and Information,Guizhou University,Guiyang 550025,China)

【机构】 贵州大学计算机科学与信息学院

【摘要】 为了解决网页信息的自动抽取,该文提出了一种基于视觉特征和领域本体的Web信息抽取算法。该算法以基于领域本体的信息抽取为基础,根据网页的视觉特征来准确划定信息抽取区域,然后结合DOM树技术和抽取路径的启发式学习,获得Web页面中信息项的抽取路径。通过信息项的抽取路径自动生成信息项的领域本体,通过信息项的领域本体解析出信息项的抽取规则。使用本算法来进行Web信息的抽取,具有查全率与查准率高、时间复杂度低、用户负担较轻和自动化程度高的特点。

【Abstract】 Put forward a Web information extraction algorithm based on visual features and domain ontology in order to solve the problem of Web information automatic extraction.This algorithm is on base of domain ontology-based Web page information extraction,according to the visual characteristics of the sample Web page to accurately delineated the area of information extraction,and get the Web page information item extraction path by combining DOM tree technology and extraction path heuristic learning.Through the domain ontology which is automatically generated by the extraction path,get the extraction rules of the information items.Using this algorithm for Web information extraction has many advantages,such as higher recall and precision rate,lower time complexity,lighter user burden and higher degree of automation.

【基金】 贵州省2008年省级信息化专项基金项目(0830);贵州省科技计划工业攻关基金项目(黔科合GY字[2008]3035)
  • 【文献出处】 计算机技术与发展 ,Computer Technology and Development , 编辑部邮箱 ,2011年02期
  • 【分类号】TP393.09
  • 【被引频次】14
  • 【下载频次】229
节点文献中: 

本文链接的文献网络图示:

本文的引文网络