节点文献
个性化智能元搜索引擎模型研究
Research on Personalized Intelligent Meta Search Engine
【作者】 李娅;
【导师】 余建桥;
【作者基本信息】 西南大学 , 农业机械化工程, 2006, 硕士
【摘要】 Internet自诞生以来不断成长,尤其是最近几年更是得到长足发展,功能不断扩展,信息容量呈爆炸性趋势增长,然而在信息极大丰富的同时,用户也面临着信息过载和资源迷向的问题,Internet网络环境下的信息检索于是成了一个新的研究热点。根据专家评测,目前主要搜索引擎返回的相关结果比率不足45%,用户要想获得一个比较全面、准确的结果,就必须反复调用多个搜索引擎。元搜索引擎的出现,在一定程度上解决了这些问题。 元搜索引擎技术是一种集成搜索引擎技术,它主要通过成员搜索引擎选择、文本选择、结果融合三个主要步骤来完成信息检索任务,如果系统策略设计得当,成员搜索引擎选择方法合适,那么相对于独立的传统搜索引擎来说,元搜索引擎一般可以达到更高的搜索覆盖率和更好的查询效果。但是元搜索引擎也会面临与传统搜索引擎一样的问题,就是不能对用户进行个性化分析和提供相应的有针对性的服务,而且如果系统的集成策略设计地过于简单和机械化,则元搜索引擎多数情况下并不会取得更好的信息检索效果。 本文试图通过设计一个个性化智能元搜索引擎模型来改善传统元搜索引擎所面临的不足。个性化是指模型可以针对不同的用户建立不同的用户兴趣模型,采用兴趣模型将查询定位到用户兴趣领域中并扩展用户查询,能更清晰、准确的表达用户查询;通过用户兴趣模型来过滤和筛选搜索结果,使结果的返回更有针对性。智能是指成员搜索引擎的选择,可以根据成员搜索引擎以往性能表现动态的决定每次的调度策略,选出那些可能对某个特定的领域有良好检索效果的子引擎来参与最终的搜索任务。本文取得了如下研究成果: 1.基于Ontology技术的用户兴趣模型构建 用户兴趣模型的构建对元搜索引擎的性能表现起着至关重要的作用,本论文研究了现有用户兴趣模型的构建方法,元搜索引擎中采用的兴趣模型大多使用传统的词频法来衡量某个用户的兴趣,用二元组(兴趣词条,兴趣权重)或三元组(兴趣词条,兴趣权重,词条新鲜度)表示,主要通过从用户访问记录中抽取部分主题词作为用户感兴趣的词条,同时计算其出现的概率表达用户对该词条的感兴趣程度,即:兴趣权重。 但单使用词条作为用户感兴趣的模型可能会出现用户的兴趣领域相当分散,使用该分散的兴趣模型指导用户查询的针对性不强;同时用该分散的用户兴趣模型过滤出的结果可能仍然存在不少不相关结果。为使用户模型能比较集中的反映用户对某领域的兴趣,本文提出用领域Ontology来表示用户兴趣,建立的模型包括用户感兴趣的领域以及反映对该领域感兴趣程度的主题词。建立好基于领域Ontology的用户兴趣模型后,用户的查询请求可与主题词相匹配,映射到最相关的领域主题中,使得用户的兴趣范围更明确。 2.成员搜索引擎的调度策略 本论文首先研究了现有的几种基于定性、基于定量、基于学习法的成员引擎(也称成员数据库)调度策略,基于定性、定量的调度策略需要成员搜索引擎的数据库描述信息,但很
【Abstract】 The capacity of information has been increasing massively, since the Internet was invented. However, people urgently need an effective retrieval tool to help them find the right information quickly in the infinite data domain. According to the experts’ investigation, the average precision of numerous famous Search Engine systems is below 0.45. Users have to seek help for the other Search engines in order to get the more comprehensive, veracious retrieved information. The arising of the Meta Search engine technique has solved this problem in a sense.The Meta Search engine is an integration Search engine technique, and it is constructed by several single Search Engines. When a meta-engine receives a query from a user, it invokes the underlying search engines to retrieve useful information for the user. The Meta Search engine itself involves three problems: the database selection problem (sub-engines selection), the document selection problem and the result merging problem. If the system policies are designed properly, the meta-engine has high possibility to achieve high coverage, precision and recall. But Meta Search Engine is also confronted with how to analyze personalizing characteristics of information requirements and to provide service with pertinence. If the system integration policy is too simple and there is no mechanism to solve the individualized service, the Meta Search engine would not achieve better effect compare with single search engine.A personalized and intelligent meta search engine is designed in this dissertation in order to improve the insufficiencies faced by traditional meta search engines. Personalization means to set up a user interest model pertinently and to allocate users’ queries to their interest domain for the sake of extend query in it. Thus the users’ queries can be expressed in a more accurate and clear way. The user interest model can also be applied in result filtering. "Intelligent" means dynamic sub-engines selecting decision on the basis of their performance on particular subjects demanded by the users’ queries to some best engines. The research results spread out as follows: 1. Construction of user interest model based on Ontology technologyConstruction of user interest model plays an important role in the performance of Meta search engine. Construction of user interest available is studied in this dissertation. In traditional approaches, word frequency is wildly used to measure user interest and 2-tuple (interest items, interest weight) or 3-tuple (interest items, interest weight, freshness) have been used to expressed user interest model. Interest items are extracted from user’s visiting records and interest weight is the arisen probability in the user’s visiting records.However, the user interest model constructed only by words may result in the interest domain decentralization. This model can not guide user query pertinently and quite a lot unrelated result may
【Key words】 meta search engine; Ontology technology; user interest; database selection;
- 【网络出版投稿人】 西南大学 【网络出版年期】2006年 11期
- 【分类号】TP391.3
- 【被引频次】9
- 【下载频次】325