节点文献

基于DOM树的DeepWeb接口属性自动提取算法

The Algorithm for Automatic Extraction of Deep Web Interface Attributes based on DOM Tree

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 朱杨段青玲

【Author】 Yang,Zhu Qingling Duan, 1(College of Information and Electrical Engineering,China Agricultural University,Beijing 100083,China)

【机构】 中国农业大学信息与电气工程学院计算机系

【摘要】 Deep Web接口集成是为了向用户提供一个统一的查询接口来获取Deep Web信息。要完成Deep Web接口集成,首先需对各Deep Web接口的属性进行自动提取,它们是后续集成工作的基础,如何将属性与其对应的语义文本进行准确的匹配是其中的难点。本文提出了一种基于表单DOM树的Deep Web接口属性自动提取算法,以控件节点作为起始节点,然后通过自右向左遍历的方式逐层寻找与控件相对应的语义文本,从而确定每个属性的语义信息,最后将提取的接口属性集采用XML格式保存,实验结果表明此算法具有较高的提取准确率。

【Abstract】 Deep Web interface integration is in order to provide a uniform query interface for users to access Deep Web information.Automatic extraction of interface attributes is needed first to complete the integration,which is basis for the follow-up integration work.The difficulty is finding the matching semantic text for each attribute.This paper presents an algorithm for automatic extraction of deep web interface attributes based on DOM tree,which traverses the nodes from right to left to search the matching text for each attribute starting from the control to determine the semantic information of each attribute,and conserve the attributes with XML The experiment results show good performance of the algorithm.

【关键词】 深网查询接口表单属性提取
【Key words】 deep webquery interfaceformattribute extraction
【基金】 国家科技支撑计划课题(2006BAJ09B05)
  • 【会议录名称】 中国畜牧兽医学会信息技术分会2012年学术研讨会论文集
  • 【会议名称】中国畜牧兽医学会信息技术分会2012年学术研讨会
  • 【会议时间】2012-08-13
  • 【会议地点】中国广西北海
  • 【分类号】TP391.3
  • 【主办单位】中国畜牧兽医学会信息技术分会
节点文献中: