节点文献
基于实例的Web信息抽取
Web Infomation Extraction based on samples
【作者】 张绍华;
【导师】 李天柱;
【作者基本信息】 河北大学 , 计算机及应用, 2001, 硕士
【摘要】 随着Internet的迅猛发展,Web已经成为全球传播与共享科研、教育、商业和社会信息等最重要和最具潜力的巨大信息源。由于Web信息的动态性,不规则性,信息量巨大,给信息搜索和查询带来了很大困难,Web搜索和查询是目前WWW和DB界研究的热点。本文提出了一种基于实例的快速从HTML页面中抽取信息的方法,该方法将抽取信息按对象关系模型进行重组存放在数据库中,以支持查询及各种应用。将信息抽取过程划分为两个阶段:学习阶段和抽取阶段,同时在抽取阶段中分为两个部分:抽取部分和集成部分。通过用户的少量参与,选定样本实例,预先定义模式,生成具有特点和高效的各种抽取规则(左右边界规则、文本特征、前导标识和关联规则),并保存在知识库中,然后根据知识库自动进行信息的抽取。基于这种抽取方法的原型系统可直接应用于Web查询和搜索,也可用于其它应用(例如数据仓库和数据挖掘等)的数据准备,抽取效果良好。
【Abstract】 As Internet rapidly developing, Web has already become the most important and potential information resources for global broadcasting, sharing science and studying, education, commercial and social information. Web has characteristics of dynamics, irregular and tremendous, so it is very difficult for user extracting information from Web. The paper presents a samples-based method of rapid realization of information extraction from HTML pages. The method utilizes object-relation model restructuring extracted information to support web querying and other applications, and divides the information extraction process into two phases: learning phase and extracting phase, and divides the extracting phase into two parts: extraction part and integration part. By some user抯 attending and a few selected samples and predefined schema, system form some extracting rules including left and right -border rule, text characteristics, predecessor text and dependency rule and store them into knowledge base, then system automatic extract information by rules of knowledge base. According to the prototype抯 experiment, the method is directly applied to web querying and also to data preparation of other applications, such as data warehouse and data mining, and the effect of extracting is good.
【Key words】 semi-structured; schema; information extraction; wrapper; samples-based learning;
- 【网络出版投稿人】 河北大学 【网络出版年期】2002年 01期
- 【分类号】TP393.03
- 【被引频次】5
- 【下载频次】264