节点文献
基于预定义模式的Web信息抽取
Schema-guided Data Extraction from the Web
【机构】 中国人民大学数据与知识研究所;
【摘要】 <正>1.引言随着Internet的快速发展,Web已经成为一种主要的信息来源。目前Web上的数据主要以HTML文档形式存在。最近出现的大量关于Web的研究主要有Web上的信息集成,智能信息代理,数据源间的互操作,客户应用的快速构建等。它们都迫切需要访问HTML
【Abstract】 This article presents a schema-guided data extraction approach- With the predefined semantic schema by user,user defines mappings between the HTML syntactic structure and the semantic schema.The extraction rule can be induced from mappings automatically.This semi-automatic approach can help user generate wrapper for specific HTML source quickly and conveniently.We focus on the rule induction algorithm in this paper,and also describe the two models adopted by us.Finally,we give the system architecture with module illustrations.
【Key words】 Schema-guided Data Extraction;
Data Integration;
Wrapper;
HTML;
- 【会议录名称】 第十八届全国数据库学术会议论文集(研究报告篇)
- 【会议名称】第十八届全国数据库学术会议
- 【会议时间】2001-08-26
- 【会议地点】中国河北秦皇岛
- 【分类号】TP393.09
- 【主办单位】中国计算机学会数据库专业委员会