节点文献

信息检索相关技术研究

Research on Information Retrieval Technology

【作者】 王树梅

【导师】 吴慧中;

【作者基本信息】 南京理工大学 , 计算机应用, 2007, 博士

【摘要】 随着互联网的快速发展,网上的信息呈指数级增长。因此,如何处理网上的海量信息成为非常重要的研究课题。文本分类和信息检索的研究可以帮助人们有效的从网上找到自己感兴趣的信息,帮助用户在日益增多的信息中发现对自己有用的知识。本文从以下三个方面对信息检索的相关问题进行了研究:首先对文本分类相关技术进行讨论。主要包括:1)引入义类的概念,设计了一个图结构的同义词词典,并给出了该词典的生成算法。应用该词典可以按语义对向量维数进行压缩;词典作为文本分类系统的启发式知识,可以提高系统的模拟推理能力、增加系统对开放语料的处理能力。2)提出一种仿人文本分类算法,该算法一方面基于文章的标题可以突出内容的观点,在处理特征向量时增加标题的权重;另一方面,设计了一维加权因子ω向量,用以模仿人工分类专家的略读和跳读,对大量出现在正例集而较少出现在反例集中的特征项,在计算文档聚类中心时增加它们的权重。实验表明:该算法可以较好的提高文本分类系统的性能。其次,是对网页检索相关问题的研究。主要研究内容:1)针对搜索引擎检索的对象是Web页面这一特点,通过分析HTML标签的修饰功能,结合传统的tf-idf加权公式,对网页进行加权索引。实验证明对于精确匹配,在查全率较低时系统的查准率有较大的提高。2)利用词间相关性进行查询结果重排。根据Web页面篇幅较小的特点,提出“网页主题关键词集合”的概念。利用词间相关性计算用户查询词集合与网页主题关键词集合之间的距离,对检索结果重新排序。将与用户查询需求相关性较大的网页排在前面。3)查询扩展是提高信息检索效果的一个有效方法,而扩展词的选择是查询扩展的一个难点。通过词共现分析,提出了一种新的词间相关性计算方法,应用于查询扩展,所选扩展词和查询整体关联,较好地反映了查询主题。实验表明,基于这种词间相关性进行查询扩展,对于信息检索性能有一定提高。最后,对基于内容的多媒体信息检索进行研究。分别对MPEG-7标准的部分描述子进行多媒体检索实验研究。在此基础上,1)提出了一种利用MPEG-7标准中的主颜色描述子抽取镜头视频关键帧的方法,并进行了相应的实验;2)利用主颜色描述子与同构型纹理描述子所适应的检索范围不同,结合两者对关键帧进行了检索实验;3)将以上研究结果应用于“CG(Computer Graphics)制作环境项目管理系统”。

【Abstract】 With the spread and the rapid development of Internet, online information increases greatly.So, how to organize and process the large amount of this information becomes a challenge.The research of text classification and information retrieval helps people efficiently findtheir interested information online, which means helps people find what they truly needfrom increasingly information. Three aspects, which related to information retrievaltechnology, will be discussed in this paper.In first part, technique about text classification will be discussed. We will (1) proposesemantic category, and construct a dictionary of graphic structure, along with an algorithmfor this graphic structure. As a enlightening knowledge of text classification, the dictionaryimproves the ability of simulating illation and processing opening corpus of the system; (2)propose an algorithm, which imitates human’s behavior, On one hand the algorithm isbased on the point that the information of an document can be tell by its title, so whenfeature vector is processed the algorithm enhances its weight; on the other hand, a weightparameterωvector is designed to simulate human’s skimming and skipping behaviorfor calculating method of a document cluster center, and a weight of the feature that thereare more positive examples than negative ones is enhanced. The experiment shows: Thealgorithm greatly improves the performance of a text classification system.Questions about Web pages will be discussed in the second part, including: (1) Giving akey technique to weight the index in information retrieval. As for search engines aredesigned to find the Web pages, which the user need. In order to weight the index, weexplore the feature of the Web pages that written in HTML. The experiment demonstratesthat the precision is improved compared with the traditional method (tf-idf) when the recallis low.(2) Bringing forward a new concept "Topic Keywords Set" (TKS). As forinformation retrieval online, the objectives searched are Web pages, the feature of thesepages is that they often small, presenting just one subject. TKS along with the explorationof the words’ relationship, by calculating distance between the user’s query and TKS,re-sort the result list. (3) Query expansion is an efficient way in improving informationretrieval quality. And in query expansion the selection of expansion words is a crucial anddifficult step. By analyzing the words co-occurrence, we proposes a new method to evaluate words’ relevance. With this method, selected expansion words are relevant withthe whole query, capable of representing the theme of query, and effient in improving theperformance, which proved by experiment.At last, a research on multimedia information retrieval, which based on content, will bediscussed. The discussion will be on basis of some different descriptors under the MPEG-7standard. According to above, we will: (1) propose a method, using dominant colordescriptor in MPEG-7, to extract the key frames from the scenes, along with an experiment;(2) give an experiment in key frame retrieval, taking advantage of the different searchingarea of dominant color descriptor and homogeneous texture descriptor; (3) apply the twoabove achievements into the material base of "CG(Computer Graphics) producingproject management system".

节点文献中: 

本文链接的文献网络图示:

本文的引文网络