节点文献
基于视觉特征和领域本体的Web信息抽取
Visual Features and Domain Ontology-Based Web Information Extraction
【摘要】 为了解决网页信息的自动抽取,该文提出了一种基于视觉特征和领域本体的Web信息抽取算法。该算法以基于领域本体的信息抽取为基础,根据网页的视觉特征来准确划定信息抽取区域,然后结合DOM树技术和抽取路径的启发式学习,获得Web页面中信息项的抽取路径。通过信息项的抽取路径自动生成信息项的领域本体,通过信息项的领域本体解析出信息项的抽取规则。使用本算法来进行Web信息的抽取,具有查全率与查准率高、时间复杂度低、用户负担较轻和自动化程度高的特点。
【Abstract】 Put forward a Web information extraction algorithm based on visual features and domain ontology in order to solve the problem of Web information automatic extraction.This algorithm is on base of domain ontology-based Web page information extraction,according to the visual characteristics of the sample Web page to accurately delineated the area of information extraction,and get the Web page information item extraction path by combining DOM tree technology and extraction path heuristic learning.Through the domain ontology which is automatically generated by the extraction path,get the extraction rules of the information items.Using this algorithm for Web information extraction has many advantages,such as higher recall and precision rate,lower time complexity,lighter user burden and higher degree of automation.
【Key words】 visual features; domain ontology; Web information extraction; path learning; discovery learning;
- 【文献出处】 计算机技术与发展 ,Computer Technology and Development , 编辑部邮箱 ,2011年02期
- 【分类号】TP393.09
- 【被引频次】14
- 【下载频次】229