节点文献

DynamicView中信息抽取系统的设计与实现

An Information Extraction System for DynamicView

【作者】 何娟

【导师】 高志强;

【作者基本信息】 东南大学 , 计算机软件与理论, 2006, 硕士

【摘要】 万维网(WWW)技术的不断发展促进了Web信息检索(Web Information Retrieval,WIR)和Web信息抽取技术(Web Information Extraction,WIE)的迅猛发展,如何从Web中抽取相关信息引起了人们的广泛关注。Web信息检索可用于从Web上的海量页面中找到相关信息所在的页面地址。与Web信息检索不同,Web信息抽取可以从一个具体的Web页面中抽取出相关信息,并以结构化的形式描述。现有的Web信息抽取算法可以分为以下两类:一是基于页面半结构化特征的信息抽取,例如html页面结构文法推断(Grammar Inference)和页面分段(Page Segmentation);二是基于自然语言文本特征的抽取,例如模版-槽填充(Template Filling)。与自然语言文本(Free Text)信息抽取相比,Web上某个具体领域中已标记的页面数量较少,因此如何在减少手工工作量的基础上保证较高的信息抽取系统的精度和召回率是有待解决的重要问题之一。本文在分析现有信息抽取算法的基础上,从DynamicView项目中信息抽取面临的问题出发,以准确探测研究员主页中的研究兴趣为目的,设计了基于列表页面导航特性和结构模版规则参数学习的研究员主页发现算法和基于页面分段技术的研究兴趣信息抽取算法。前者用于获取研究员的姓名及其主页地址,它将Web信息检索技术和Web信息抽取技术结合,能够高精度地获取具有相同特征的页面集合的问题。后者通过基于分隔符的页面分段算法过滤无关数据,并根据本体表示的领域知识从相关段落中抽取研究兴趣。本文将这两种方法运用到DynamicView系统中,实验结果证明这种方法是高效的、可靠的。

【Abstract】 With the growing of World Wide Web (WWW), WIR and WIE techonology has been developed rapidly. More and more researchers are paying attation to how to extraction information from the Web.WIR can be used for locating a specific page that contains the relevant information on the Web. Unlike WIR, WIE can extract revelant information from a specific page directly and transform the revelant information into structural format.Generally speaking, WIE methods can be divided into two categories: one is based on structure of a page, such as Page Structure Grammar Inference and Page Segmentation; the other is based on language feature of a page, such as template filling. Unlike free text information extraction, the number of annotated web page for a specific domain on the Web is small. Hence, how to extract information with high accuracy without increasing the tedious manual work is a critical problem to be solved.Based on the analysis of existing WIR and WIE algorithm and the target of DynamicView project, this thesis proposes a WIR algorithm based on structure template to get the faculty’s homepage from the Web and a page segmentation based WIE algorithm to extract the facultys’research interest from their homepages. The WIR algorithm applies WIE technology into WIR algorithm. In this way, the web pages with the same attributes can be found easily. The page segmentation algorithm DeSeA (Delimiter based Segmentation Algorithm) for WIE can be used to filter irrelevant information out in a web page. After this, research interestes can be extracted easily from the relevant segments using the domain knowledge. Experiments show that these two algorithms fit commendably with DynamicView.

  • 【网络出版投稿人】 东南大学
  • 【网络出版年期】2007年 04期
  • 【分类号】TP311.52
  • 【下载频次】146
节点文献中: 

本文链接的文献网络图示:

本文的引文网络