节点文献
基于案例推理的科技文献推荐系统研究
Study on Recommendation System of Technological Literatures Using Case-Based Reasoning(CRB)
【作者】 席俊红;
【导师】 姜丽红;
【作者基本信息】 华东师范大学 , 软件工程, 2005, 硕士
【摘要】 在过去的几年里,因特网(Internet)迅速发展,人们能够更容易、更直接地获取各种形式的信息。由于Internet是一个具有开放性、动态性和异构性的全球网络,资源分布很分散,且没有统一的管理和结构,这就导致了信息获取的困难。如何快速、准确地从浩瀚的信息资源中找到所需资源已经成为困扰网络用户的一大困难。现在Internet上有许多搜索引擎,如Yahoo,Web Crawter等等,可以帮助人们搜索Internet上的各种信息。它们能够根据用户提供的关键字采取一定的匹配策略,在索引库中进行查找,然后将结果地址集返回给用户。由于语言的模糊性,词语的多义性,一个主题通常可以用多个关键字来表示,用户常常难以选择合适的关键字把他的兴趣准确地表达出来。搜索引擎返回的地址集经常包含很多用户不需要的无关信息,使得用户常常花费很长的时间却没有找到对自己有用的信息。个性化推荐技术就是针对这个问题提出的,它针对不同的用户,提供不同的服务,以满足不同的需求,个性化推荐系统通过收集和分析用户信息来学习用户的兴趣和行为,从而实现主动推荐的目的。目前存在着许多个性化推荐模型,大致上分为两种:基于内容的个性化推荐系统和基于协同过滤的个性化推荐系统。后者的优点是能为用户发现新的感兴趣的信息,缺点是存在两个很难解决的问题:一个是稀疏性,亦即在系统使用初期,由于系统资源还未获得足够多的评价,系统很难利用这些评价来发现相似的用户;另一个是可扩展性,亦即随着系统用户和资源的增多,系统的性能会越来越低。本文对上述的协同过滤技术中存在的稀疏性问题采用案例推理进行一定程度的改善;并在计算用户间的相似度时,采用归一化的欧几里德距离来代替余弦表示,使得效率有所提高,从而对可扩展性问题有所改进;同时对用户偏好的更新提出一种简单可行的策略。
【Abstract】 In the past few years, the Internet develops very fast, and it is easier and more direct to get various information. Because Internet is an open, dynamic and non-homogeneous network, lacking unified architecture and management, it is virtually impossible to separate the documents that interest a given user from myriad of other documents. Now, how to find information valuable is a time consuming process. There are many search engines, such as Yahoo and WebCrawter, which can help people to search the web based on keywords offered by users and certain match strategies. They can find relevant URLs in the index database and send back to users. Because of the fuzzy character in languages, the same keywords often appear in different contexts, and it is difficult for users to select suitable keywords to express what they want accurately. The result of search engines often contains so much non-relevant information that people have to spend much time and effort to navigate but may not find any personalized information.Personalized Recommendation technique, which is developed to attack this problem. According to different users’ various tastes, it holds the promise to serve their customers to find resources on WWW voluntarily, by collecting and analyzing the users’ interests and behaviors. For the moment, there exist many Personalized Recommendation models, being classified in two classes, in substance: Content-based system and Collaborative Filtering system. The latter has the vantage of finding new resources which may interest users, but do have two scabrous problems: one is the sparsity problem which the system can not find similar users due to system resources’ lacking of enough ratios, at the initial stage of system; The other is the expansibility that the system’s performance will become more worse because of its users and resources become more and more.By using Case-Based Reasoning(CBR), this article presents an approach to address the sparsity problem ; And uses Euclid distance to replace cosine representation as the computational method of users’ similarity in order to attack the efficiency problem; At the same time, we also address a simple, workable tactic in the light of updating user’s profile.
【Key words】 Personalized Recommendation technique; Collaborative Filtering; Case-Based Reasoning (CBR);
- 【网络出版投稿人】 华东师范大学 【网络出版年期】2005年 05期
- 【分类号】TP319
- 【被引频次】6
- 【下载频次】353