节点文献
基于社会化媒体的自适应信息推荐机制研究
Research of Adaptive Information Recommendation Mechanism Based on Social Media
【作者】 王佳;
【导师】 李庆;
【作者基本信息】 西南财经大学 , 计算机应用技术, 2011, 硕士
【摘要】 由于互联网的优越特性,在其上发布信息极为便捷,这就使得互联网上的信息数量以近乎爆炸的速度增长。如此多的信息即使浏览一遍都无法做到,用户希望能找到感兴趣的部分更是不可能的。传统的搜索方法只能呈现给所有用户一样的排序结果,无法针对不同用户的兴趣偏好提供相应的服务。信息的爆炸使得信息的利用率反而降低,这种现象被称之为“信息过载”。推荐系统是为解决互联网上的信息过载问题而提出的一种智能代理系统,能从互联网的大量信息、中向用户自动推荐出符合其兴趣偏好或需求的资源。在当前Web 2.0的环境下,社会化媒体的出现使得用户不仅是网络内容的浏览者,也是网络内容的制造者。它的发展进一步加剧了网络时代的信息爆炸。传统的推荐系统通过让用户回答问题或者主动定制的方式来获取用户的兴趣,进而实现推荐。然而,用户的兴趣不是一成不变的,它会随着时间的推移而变化。针对该点,本文提出了一种自适应信息推荐机制,来及时跟踪用户兴趣变化,推荐用户感兴趣的资源。社会化媒体形式多样,如论坛、博客、内容社区、社交网络等。在这些形式下,用户可以发布或者转帖一篇文章,其他用户可以对其阅读或评论,这些评论本身又会被其他用户阅读或评论。从用户评论中,可以观察出用户当前感兴趣的话题。传统的基于内容的推荐方法一般根据原文的内容信息来推荐相关文章。然而,我们知道,随着用户讨论的继续,讨论的主题也会发生变化,即用户兴趣也会发生变化。这时,如果仅仅依据原文本体进行推荐,则返回的文章往往不是用户当前最感兴趣的,从而会降低用户的满意度。因此,本文考虑了结合用户评论和原文本体来构建主题模型,利用该模型来选择相关文章。根据观察发现,每条评论对推荐结果的影响应该是不一样的,如有些评论对原文内容有深刻的见解,而有些评论完全是无意义的讨论。所以,当利用用户评论信息来跟踪主题演变时,区分开每条评论的影响非常重要。这里,我们从用户评论中抽取出评论间语义关系、结构关系以及用户权威来区别每条评论对推荐的影响。分析事件报道在网络上的传播,可以发现其存在如下四个特点:转载重合、报道重合、包含重合和追踪重合。这些特点使得基于内容的推荐系统存在一个严重问题—重复推荐,即推荐文章的内容与原文含有相同的信息,这样会增加用户的阅读负担。于是,本文提出了一种方法来解释推荐文章与原文本体之间的逻辑关系(包括一般化、特殊化和重复),以此降低重复内容的推荐,推荐出符合用户需求的文章。本文第一部分介绍了课题的研究背景、研究目的和意义,对文中涉及到的一些基本概念作了简单介绍。介绍了推荐系统的定义;四种主要方法,即基于内容的推荐、协同过滤推荐、混合型推荐和基于数据挖掘技术的推荐;针对四种方法,分别以一个系统实例解释其工作模式;对推荐系统的评测标准进行了汇总。还介绍了社会化媒体的概念以及与传统媒体相比,其具有的一些特点。最后,总结了本文的主要工作和贡献如下:(1)本研究是在国内外率先结合用户评论来协助信息推荐服务的研究,为基于社会化媒体的信息推荐研究提供一条崭新的研究思路,将信息推荐的研究从Web 1.0的传统静态媒体延伸到了Web 2.0的社会化媒体模式。(2)为了充分利用社会化媒体的用户交互体验特征,我们独创性地设计了一套基于图论的用户评论信息挖掘机制,可以准确地捕捉用户对事件的关注焦点,并将其与原文本体内容相结合,使得推荐的结果既反映了作者的观点,也反映了读者的观点。(3)为了减轻用户的认知负担,我们创新性地提出了一套基于信息熵理论来判断文本逻辑关系的机制。通过该机制,我们可以获得推荐文章与原文章的逻辑关系。此外,该研究成果可以广泛地应用到文本分析的内容逻辑判断中。例如,搜索引擎的结果呈现,基于内容的广告设置等。本文第二部分介绍了该课题的研究基础与背景。首先,针对本文的实验对象,即新闻和博客,对已有的相关研究工作进行了总结。新闻推荐从现有的商业新闻推荐系统和学术研究两个方面进行了介绍。接着,针对文中存在的主题漂移问题,对主题检测与跟踪技术的研究发展进行了汇总。最后,对本文将涉及到的相关理论知识作了简要介绍,如语言模型,PageRank算法、信息熵、T检验等。本文第三部分是核心部分,介绍了自适应信息推荐机制的设计。首先,展示了总体系统框架图,并对其运作流程进行简单介绍。然后,针对框架中的各个模块进行详细阐述。通过用户间关系建模计算用户权威,这里的关系包括了引用关系与回复关系。在整个社区中,根据一个用户对另一个用户的信息进行引用或者回复来构建图模型,然后利用PageRank算法计算每个用户的权威。接着,计算评论权重。这里,我们同样利用了图模型,不同的是,现在的模型是建立在用户评论之间的关系上,这里的关系包括了语义、引用和回复关系。语义关系指的是两条评论之间的内容相似性,引用或回复关系指的是一条评论对另一条评论的信息引用或者回复。模型构建好后,也利用PageRank算法得出评论的权重。一条评论质量的好坏,由其作者的权威和评论本身共同决定,因此,我们将用户权威和评论权重结合起来,计算出每条评论的最终权重。其次,将这些权重信息和原文本体、用户评论一起输入到合成器中,构建主题模型。利用该主题模型从数据库中检索出相关文章。最后,根据信息熵理论来解释相关文章与原文本体之间的逻辑关系,返回符合用户兴趣的文章。本文第四部分是实验设计与分析。介绍了系统开发环境、实验数据的获取以及详细信息。实验数据包括两部分:一个是新闻数据集,一个是博客数据集。由于我们获取的是整个网页数据,所以需要对网页进行解析,抽取出所需部分。还介绍了评测标准的选取,为了评测目的,我们除了选用一些常用的指标,还引入了一个新的评测指标—新颖度,来度量返回文章的主题多样性。接着,设计了一系列实验:1)将本文提出的方法与两种常用方法进行比较,结果表明,在新闻和博客数据集上,我们的方法都明显优于其它两种;2)分析了用户权威和评论对推荐效果的影响,实验结果表明结合用户权威和评论信息有利于提高推荐效果;3)分析了评论间关系对推荐效果的影响,实验结果显示,针对不同的文本形式,有不同的推荐效果。对于新闻数据,结合用户评论间的内容关系会导致推荐效果的降低;然而,对于博客数据,结合用户评论间的内容关系有助于推荐效果的提高;4)对推荐关系解释进行了评估。本文的最后一部分是对本文研究工作的总结和未来研究工作的展望。总结了本文研究的基于社会化媒体的自适应信息推荐系统的整体设计;针对本文的研究工作,指出了其存在的一些不足之处,并给出了以后的发展方向。
【Abstract】 In the Internet, it is very convenient to release some information, which makes the quantity of information grow explosively. So much information cannot be skim through, and of course it is more impossible that a user is hoping to find interested things. Traditionally search engines only present all users the same sorted results, and can not provide the corresponding services according to different users’ interest preferences. Information explosion leads to reverse utilization, which is called "information overload". To solve the problem of "information overload" on the Internet, recommender system is proposed as a kind of intelligent agent system, which can automatically recommend the resources to users from the Internet that meet their interest preferences or demands.In the current Web 2.0, the emergence of social media has made that each user can not only browse the Web, but also create and disseminate information, which leads to information overload more serious. The traditional recommender systems acquire users’ interests by letting users answer questions or active customization and recommend relevant items. However, users’ interests will change as time passes by. For the point, this paper puts forward an adaptive recommendation mechanism, to make timely follow-up of user interest’ change, and recommend users interested resources. Social media has various forms, such as BBS, Blog, content community, social network and so on. In these forms, a user can post or repost an article, and other users can read or comment on it, these comments themselves can be read and commented by other users. From the users’ comments, we can observe the current topic of common interest of users. The traditional content-based recommenders usually recommend related articles based on the original content. However, as we know, with users’ discussion going on, the topic will also be changing, namely the users’ interest will also be changing. At this moment, if recommendation is only based on the original, returned articles will be of no interest to users, which will further reduce user satisfaction. Therefore, in this paper, we consider combining user comments with the original to build a topic profile, and then utilize the topic profile to select some related articles. On the basis of our observation, the impact of each comment on recommendation varies according to its quality. In particular, some comments reflect insightful opinions on the original, which provides balanced views from both readers and authors, while some are meaningless discussions. Differentiating the contribution of each comment is important to utilize them properly in guiding the topic evolution for recommendation in social media. Here, we extract structural, semantic, and authority information carried by the comments to differentiate the importance of every comment. Analyzing coverage transmission on the Internet, we can find that it has four features as follows:reposting coincidence, reporting coincidence, containing superposition and tracking superposition. These characteristics make the content-based recommendation system has a serious problem-repeat recommendation, namely the recommended contains the same information as the original, which will increase users’reading burden. Hence, a method is proposed to explain the logical relationship between recommend articles and the original (including generalization, specialization and repetition), in order to reduce repetitive content and recommend the articles that meet user demands.The first part introduces the research background, research purpose, significance, and some basic concepts involved in this paper. About recommender system, it first introduces the definition and four common methods, including content-based recommendation, collaborative filtering recommendation, mixed recommendation and data mining based recommendation, then for these methods respectively exemplifies a system to explain its working mode, and summarizes the evaluation standards of recommender system. Besides, it also introduces the concept of social media and its features, compared with traditional media. At last the contributions are listed as follows:(1) This study is to take the lead in utilizing user comments to assist the information recommendation service at home and abroad. It provides a new idea for adaptive information recommendation in social media and extends the study of information recommendation from traditional static media in Web 1.0 to social media in Web 2.0.(2) In order to make the best use of engaging interaction among users in social media, we design a mechanism, which mines information from user comments by using the graph theory, to accurately capture users’ concern about some event. Then, it is combined with the original content, which can balance views of both authors and readers.(3) In order to reduce user cognitive burden, we put forward a creative approach to generate hints to indicate the logical relationship between articles based on the information entropy. Thus, we can judge the relationships between recommended articles and the original posting. In addition, the research can be widely applied to text analysis, for example, presentation of search engine results, based-content advertising.setting, etc.The second part introduces the research foundation and background. First of all, for our experimental objects, namely news and Blog, their existing works are summarized. News recommendation is presented from two aspects of existing business recommenders and academic research. Secondly, for a fundamental challenge in this paper, that is topic divergence, the development of topic detection and tracking (TDT) is in the summary. Finally, some involved theories are briefly introduced, such as language model (LM), PageRank algorithm, information entropy, T-test, and so on.The third part is the core, which introduces the design of our adaptive recommendation mechanism. First, it presents a framework for recommendation and briefly introduces the process. Then, each module expounds in the framework. The authority for each user is calculated through modeling the relations among users, which include quotation and reply. In the whole community, it constructs graph models when a user replies to another user’s posting or a user quotes another user’s posting separately, and employs a variant of the PageRank algorithm. Next, comment’s weight is calculated. Here, we still use the graph model, however, differently, the model is built on the relationships among users’ comments, including semantic, quotation and reply. The content relation means the semantic similarity between comments, and the quotation or reply means that a comment quotes another one, or replies to another one. After building the models, another variant of the PageRank algorithm is used to calculate the weight of every comment. The quality of a comment can be determined by its authority and itself, therefore, it combines user authority with comment weight to get the final weight for each comment. Then, this information along with the entire discussion thread is fed into a synthesizer to construct a topic profile, which balances the perspectives of both authors and readers. With the topic profile constructed, the retriever returns an ordered list of articles with decreasing relevance to the topic. Finally, we utilize the information entropy to explain the logical relationships between returned articles and the original, and recommend articles with users’ interest.The fourth part is about experimental design and analysis. It presents the system development environment, experimental data acquisition and its details. Experimental data includes two synthetic data sets:one is news, the other is Blog. Because we obtain the entire data in a web page, the page is first parsed to extract our required parts. It also introduces the selection of evaluation standards. Here, besides some common indexes, we introduce a metric of innovativeness to measure the topic diversity of returned articles. Then, we design a series of experiments to test our proposal:1) we compare our work to two baseline works. The results show that our approach performs significantly better than the baseline methods for both news and Blog data sets; 2) we study the effect of user authority and its integration to comment weighting and the result shows that with the assistance of user authority and comments, the recommendation precisions are improved for news and Blog; 3) we investigate the effect of the semantic and structural relations among comments, i.e. semantic similarity, reply, and quotation. The result indicates that semantic contents of user comments can play a fairly different role in a different form of social media. For the case of news, incorporating content information adversely affects recommendation precision. On the other hand, when we test the Blog data set, the trend is the opposite, i.e. content similarity does contribute to retrieval performance positively; 4) we evaluate the performance gain obtained from interpreting recommendation.The last part summarizes the research work and looks into the future research. It makes a summary of the overall design and implementation of adaptive recommendation system based on social media. Furthmore, it also points out some deficiencies and gives some future directions.
【Key words】 Recommender System; Social Media; Adaptive Recommendation; News Recommendation; Blog Recommendation; Content-based Filtering;