节点文献

基于信息抽取的知识生成系统

A Knowledge Generation System Based on Information Extraction

【作者】 白曦

【导师】 孙吉贵;

【作者基本信息】 吉林大学 , 计算机软件与理论, 2008, 硕士

【摘要】 随着互联网技术的迅猛发展以及越来越多的网页被发布,海量的信息以电子文档的形式出现在我们面前。为了及时应对信息大爆炸所带来的严重挑战,人们迫切需要借助一些自动化的工具从海量的信息源中去粗取精,去伪存真,迅速找到自己需要的有价值信息。信息抽取技术正是在这种背景下产生出来的。该技术原来的目标是从自然语言文档中找到特定的信息,是自然语言处理领域的一个重要分支。包装器是一种广泛应用的信息抽取技术,利用它人们可以把网页转化为结构化的数据。但是不同网页的结构不尽相同,而且同一网页在更新时网页结构也有可能发生改变。因此,找到一个相对通用的方法来自动或者半自动地产生包装器并对网上信息进行准确抽取是当今信息处理领域中的一个热点问题。由于传统的抽取技术不是基于语义的,提取出来的信息无法被计算机所理解,因而达不到数据处理智能化的目的。语义网是一种能够让计算机理解的新型的Web内容形式,在它的辅助之下,计算机会根据关键名称定义的超链接和逻辑推理规则发现语义数据的含义。在上述背景下,本文对基于本体的信息抽取和知识生成技术进行了深入地探讨和研究,基于模式发现和领域本体,利用后缀数组将包装器学习出来,领域本体自动将抽取出来的原始数据进行映射,并形成知识存贮在RDF文件当中,从而实现了从网页中半结构化的内容中抽取知识。本文同时设计并实现了一个平台独立的基于领域本体的手工标注工具。其主要功能是通过预匹配、本体呈现、实例名推荐等来指导用户对网页进行语义标注,最终生成知识并存放在RDF文件当中。

【Abstract】 With the emerging of the Web technology, a great amount of information appears on the internet. Facing to the information, people can not get helpful information effectively and efficiently. Therefore, giving specific semantics to the information becomes more and more important nowadays. Information Extraction (IE) structurally processes the information contained in the electronic documents and transforms them into the tabular formations. The inputs are the original documents and the output are the information points with stable formations. The original goal of IE is to find out the specific information from documents containing natural language content. IE is a very important sub-domain of the natural language processing domain. The IE systems should not only process the structured text containing the tabular information but also process the free text. IE is especially helpful for extracting specific facts from a good amount of documents and offers people an avenue for obtaining the information on the Web. The traditional IE techniques are based on the statistics and rules. The main idea is to do the statistics on the key words and create some auxiliary rules to do the matching within the extraction process. However, since traditional IE techniques are not based on the semantics, the extracted information can not be understood by computers or processed intelligently. The goal of developing the Semantic Web is to establish a new brand environment in which all the data can be readable by computers. So the computers can serve the people intelligently. The core of this technique is to make the computer understand the semantics of the data when the data are extracted. It is a challenge work that non-semantic information is transformed into the knowledge carrying the semantics. Among many current information processing techniques, the wrappers are widely used to transform the Web documents into the structured data. The Web documents from different Web sites have diverse structures and the structure of a single document will be changed with the updating of the Web site. Therefore, looking for an automatic or semiautomatic approach for generating wrappers and extracting information hidden in the Web documents has become one of the hottest problems in the information processing community.Based on the above analysis, this paper deeply discusses and researches on the IE and knowledge generation techniques based on domain ontologies. The research mainly focuses on the following areas: pattern-discovery-based IE and template-based IE, the design and the implementation of the manually annotation tool, and ontology learning based on the two-stage clustering.1. Based on the pattern discovery and domain ontologies, a tool for extracting knowledge from semi-structured content is implemented. Within the extracting process, wrappers are learnt through the suffix arrays. Then domain ontologies automatically map the raw data to the knowledge stored in the RDF documents. After merged, newly generated knowledge are finally added to the knowledge base which will be further used to be queried by users. An evaluation method is proposed to fulfill the comparison between our method and the traditional methods.2. A manually annotation tool is designed and implemented. The main functions of this tool are guiding users to annotate the Web pages by pre-matching, displaying ontologies hierarchies, recommending individual names. Finally, new knowledge will be generated and stored in the OWL documents. This tool contains the following modules: the inputting documents module, the manually annotation module, the tips module, the ontologies displaying module and the saving module.3. A two-stage clustering method for learning domain ontologies is proposed. It is based on the SOM natural network and the hierarchical clustering. Chinese lexical analysis and XPath are used within the extraction process. This paper also compares the performance for our two-stage clustering method with that of the traditional ontology learning method.Within the process of fulfilling the above tasks, the author finds the following problems. The performance of the knowledge extraction system depends on the correctness and the completeness of the related domain ontologies. In order to model all the content and obtain the knowledge, we need the supervision of the domain experts to further evolve the existing ontologies. When users annotate the Web page, they may find the disadvantages of the current involved ontologies. How to give users some reasonable guidance and help them safely modify the structure of ontologies becomes a challenge for developing the annotation tools. Within the process of ontology learning, the error information and the conflicts should be detected and corrected. Since the involved reasoning techniques are not mature enough nowadays, the two-stage method has not implemented the detection and the correction. In this paper, the knowledge extraction system based on the domain ontologies and the pattern discovery is a prototype system and since some related techniques need to be improved, its functionalities are restrained. However, the fruits of this paper will definitely benefit the related researches in the future.

【关键词】 信息抽取知识标注语义网
  • 【网络出版投稿人】 吉林大学
  • 【网络出版年期】2008年 10期
  • 【分类号】TP391.1
  • 【被引频次】1
  • 【下载频次】261
节点文献中: 

本文链接的文献网络图示:

本文的引文网络