节点文献

基于时间域的信息检索系统的设计与实现

The Design and Implement of Information Retrieval System Based on Time Range

【作者】 孙琦

【导师】 牛军钰;

【作者基本信息】 复旦大学 , 计算机应用技术, 2008, 硕士

【摘要】 上世纪90年代,互联网的兴起加速了信息与知识的传播。近年来,随着计算机的普及以及硬件性能的加速提升,以文本方式呈现的信息数据正急速膨胀着。大规模信息检索系统的出现为人们查找所需信息提供了很好的帮助,因此,信息检索的相关技术也一直是研究的焦点。这其中包括:索引的结构与构建算法,索引的压缩与维护,检索模型,查询反馈与扩展,top-k的高性能查询处理算法等。这些技术为信息检索系统的发展提供了坚实的基础。但随着时间的推移,信息一直在不断地积累着,人们对历史数据信息逐渐产生兴趣,这种需求随着数据的积累会逐渐显著,尤其是近年来web2.0的发展,各类社区以及用户blog中的信息不断更新,加速了人们对该领域的研究。目前,已经有一些研究者注意到这一问题,并试图提出一些解决方案。本文综述了信息检索系统的基本原理,详细介绍了文本检索系统的各主要构件的实现细节。提出了动态文本环境中高性能的支持任意时间段检索的索引结构以及查询算法,实现了以高校社区站点为对象检索系统。本文主要工作包括:●本文提出了一种支持高性能时间段查询的索引组织方式;●本文在新的需求环境下,改进了时间段索引中压缩算法;●本文详细分析了各检索模型的主要特征,使用一种简化的模型NRA-Okapi,有效地支持了高性能top-k算法;●本文对以上方法在TREC 2006 Genomics Ad-hoc语料进行了评测●针对社区文本不断演化的特征,本文设计并实现了一个面向高校社区的检索系统。

【Abstract】 In the 1990s, the rise of the Internet has accelerated the spread of information and knowledge. In recent years, with the popular of computer and high boosted performance of hardware, the text information presents a rapid expansion. Large-scale information retrieval system has made a great help for people who want to find some information they need. These technologies relevant with information retrieval have been the focus of study. Such as, the index structure and its construction algorithms, index compression and maintenance methods, document scoring model, query feedback and expansion, top-k, high-performance query processing algorithms, and so on. These developed technologies provide a solid foundation for the development of information retrieval system.But with the lapse of time, information has kept been accumulated continually as historical data, in which people are gradually interested. This demand for mining historical data is growing significantly, especially with the development of web 2.0 in recent years. The situation that internet users constantly updated information in various communities and their own blogs greatly mounted up the quantity of the data which accelerated researching in the area. At present, researchers have been aware of the issue and proposed some solutions.This paper makes a survey of the basic principles of information retrieval system dives into details of how to implement the major components in the IR system. A new index structure and the associated query algorithm are proposed to efficiently support the retrieval in arbitrary time frame in the frequently updated text environment. The paper also demonstrates a retrieval system targeted to college community. This paper makes the following contributions:Propose a high-performance index organization for retrieval in arbitrary time frame;Improve the index compression algorithm in the new index structure;Analysis the features of retrieval model, employing a simplified model of NRA-Okapi to effectively support the top-k query in text retrieval system;Evaluate the above methods in the corpus from TREC 2006 Genomics Ad-hoc; Design and implement a retrieval system targeted in college community;

  • 【网络出版投稿人】 复旦大学
  • 【网络出版年期】2009年 03期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络