节点文献
基于语义本体的智能搜索引擎研究
【作者】 陆峰;
【作者基本信息】 复旦大学 , 软件工程, 2009, 硕士
【摘要】 互联网改变了人与人之间的交流方式以及商业的运作模式。它的核心是一场革命,这场革命目前正逐渐把世界向知识型经济和社会转变。然而,要为网络用户提供更好服务还存在着一个主要的障碍,即目前机器是不能理解网页的具体含义的。一些基于关键词的搜索引擎,如AltaVista,雅虎和谷歌,是当今网络的主要搜索工具。显然如果没有搜索引擎,网络不会获得如此巨大的成功。但是它们的使用仍然存在一些严重的问题如:撤消频率高,精度低,搜索结果对词汇高度敏感等。为克服以上存在的部分问题,本文提出了一种基于语义技术的智能搜索引擎。首先,本文分析了现有本体的构建方法,结合它们的优点,提出了新闻领域的本体知识库构建方法;在此基础上,本文提出了一种基于领域本体和网页视觉分块技术相结合的领域信息收集器,该收集器能够在语义层次上对网页内容进行理解,同时结合网页视觉分块技术,使其能在不同模块中对相似度进行分别计算,提高主题识别率。其次,本文提出了一种基于领域本体的智能搜索引擎模型,它实现了搜索引擎在概念层次上对信息进行检索,部分克服了传统搜索引擎基于关键词检索的局限性;实现了对于多信息源的融合检索机制,克服了原来单一的文本检索。最后介绍了基于领域本体的智能搜索引擎相似度计算方法,提出了基于传统向量模型和领域本体相结合的相似度计算方法,以提高用户的检索体验。
【Abstract】 The Internet and World Wide Web has changed the way people communicate with each other and the way business is conducted. It lies at the heart of a revolution that is currently transforming the world toward a knowledge economy and knowledge society. However, the main obstacle to provide better support to Web users is that, at present, the meaning of Web content is not machine-accessible. Keyword-based search engines, such as AltaVista, Yahoo, and Google, which are the main search tools of the existing network. It is clear that the Web would not have been the huge success it was, were it not for search engines. However, there are serious problems associated with their use: high recall, low precision; high recall, low precision; results are highly sensitive to vocabulary. This paper proposes a structure of intelligent search engine base on semantic techniques.Firstly, the paper analyzed the constructions of existing ontology. Through combining these advantages, a method of building the news-based knowledge is presented in this foundation, the paper introduced a kind of information collector on basis of combining the domain ontology with web page visual block technology. This collector could learn the page content on semantic hierarchy and calculate the similarity respectively in different modules depending on the page visual block technology in order to increasing the identify ratio of page information.Secondly, the paper proposed a kind of intelligent search engine model based on the domain ontology. It realized the information retrieval on concept hierarchy, partly overcoming the limitations of traditional keywords-based retrieving. It also realized the retrieval of multi-information fusion, overcoming the original simplex method of text searching.Finally the paper illustrates a computational method of intelligent search engine similarity based on the domain ontology, and another method which is based on combining the traditional vector models with the domain ontology, for the purpose of improving the users’ feeling of searching.
【Key words】 Search Engine; Domain Ontology; Intelligent Reptiles; Ontology Similarity;
- 【网络出版投稿人】 复旦大学 【网络出版年期】2011年 S1期
- 【分类号】TP391.3
- 【被引频次】3
- 【下载频次】493