节点文献

基于历史上下文挖掘的“科技论文在线”用户行为研究

Study on User Behavior of "Sciencepaper Online"Based on the Mining of Historical Context

【作者】 张攀

【导师】 高曙;

【作者基本信息】 武汉理工大学 , 计算机应用技术, 2013, 硕士

【摘要】 “中国科技论文在线”是由教育部科技发展中心主办,以“阐述学术观点、保护知识产权、思想交流创新、论文快捷共享”为宗旨,为科研人员提供一个方便、快捷的交流的学术平台,以此平台为基础实现新成果的及时推广,科研创新思想的及时交流。作为一个信息获取类的网站,在它快捷、方便地带来大量信息的同时,也带来了许多难题:如何能使用户快速、准确地获得所需要的科研信息;如何理解已有的用户历史数据并用于预测用户未来的行为等。对于“科技论文在线”用户行为的研究可以有效地解决这些问题。在分析历史上下文信息与web信息各自的优缺点后,将历史上下文信息与web日志进行融合,融合后数据来源更为广泛,能较全面的体现用户访问页面时的环境状况,较准确的反映用户当时的情绪、心理状态,行为特征。在此两类数据基础上进行挖掘分析,可以较准确地得出用户的访问模式和访问特点。本文主要研究了历史上下文信息挖掘过程中的数据获取、融合及预处理的各阶段的算法并进行了部分改进和创新,然后利用改进的聚类分析算法DICA分析预处理得到的会话集,并根据聚类分析结果得出推荐集来实现网站站点结构改善和向用户提供推荐服务。本论文的工作主要集中在四个方面:(1)数据预处理:首先在较为全面的分析了历史上下文信息以及web日志的数据特点后,将多种历史上下文信息和服务器端的web日志进行去噪融合。然后通过会话划分算法将融合后的信息整理为会话集,在此基础上,利用用户访问轨迹重现算法模拟用户当时的访问轨迹,并以此再次细化会话集。最后利用历史上下文信息中的终端环境上下文信息,修正用户每个页面的浏览时间。(2)页面兴趣度计算:对于得到的会话集,采用基于多特征的页面兴趣度计算方法为每个页面赋权重值。针对以往权重计算算法中,不能体现用户浏览页面顺序的问题,本文提出了将会话中页面的序号作为一个特征加入页面权重的计算,有效地区分了多个用户采用不同的顺序访问某些特定页面的情况。(3)聚类分析用户行为:在对会话集中的页面赋值权重后,本文提出改进的k-means算法DICA。算法的自动获取最优聚类个数和初始聚类中心的特点有效的避免了k-means算法中需要依据经验设定初始聚类个数和随机设定初始聚类中心的缺陷。(4)生成推荐集:对带权重的会话集进行DICA算法聚类分析后得到基于群体用户的推荐集和基于个体用户的推荐集,并将这两个推荐集融合,以此来改善网站站点结构和向用户提供推荐服务。本文的研究工作得到教育部项目“基于上下文感知的“中国科技论文在线”用户行为研究”(项目编号:20121140004)的资助。

【Abstract】 The "Sciencepaper online" is sponsored by the Ministry of Education, Science and Technology Development Center. Its purpose is "expounding the academic view、 protecting the intellectual property right、exchanging ideas and innovation、sharing paper fast". On the one hand, It provides a convenient and fast academic platform for the researchers, promoting the new results timely, also making research and innovation ideas exchange timely based on this platform. As a web sit about information acquisition, on the other hand, it not only brought a lot of information quickly and easily, but also brought many problems such as:How to make users to obtain the research information quickly and accurately; How to understand the existing historical data of the user and predict the future behavior of the user? Research on the user behavior of "Sciencepaper online" can solve these problems effectively.After analyzing the advantages and disadvantages of historical context information and web information, we will integrate historical context information and log. Because of its widely data source, when user accesses the page, it can reflect both its state fully and user’s emotion and mental state accurately. Base on this data, mining analysis can obtain mode and features of the users when they accesse more accurately.This thesis mainly studies data acquisition、integration and the algorithm of each stage of pre-processing in the mining process of historical context information, and do some improvement and innovation. Then uses the improved clustering analysis algorithm to analyze the obtained session set by pre-processing, and uses the obtained recommendation sets by the results of cluster analysis to realize the improvement of the web site’s structure and provide recommended services for users.In this thesis, the research focuses on four aspects:(1) Data pre-processing:On the basis of a more comprehensive analysis of the historical context information and the data characteristics of web log, a variety of historical context information and web log on the server-side will do denoising and fusion, then use session partitioning algorithm to sort the information after fusion into session set. On this basis, we use users access the trajectory reconstruction algorithm to simulate user’s access track and refine the session set again. We use terminal context information of the historical context information to correct browsing time of each page.(2) Calculating page interest:For the obtained session set, using page interest computing method based on multi-feature to set weight for each page. As the former weight calculation algorithm will not reflect the browsing order, in this page, the serial number of the session page as a feature is added to calculate the page weight, it can distinguish multiple users who have accessed specific pages in different sequential effectively.(3) Cluster analyzing user behavior:After setting the weight of the page in session set, this paper improves the K-means algorithm. The feature of algorithm is that can automatically obtain the optimal number of clusters and the initial clustering centers, it can effectively avoid the defect in the K-means algorithm about setting the initial number of clusters based on the experience and initial cluster centers randomly.(4) Generate a recommended set:We can get recommended set based on groups of users and the individual user by weight session set DICA algorithm cluster analysis. In order to improve the web site structure and to offer recommended service for users. We will integration of the two recommended sets.

【关键词】 上下文web日志数据挖掘用户行为
【Key words】 contextweb logdata mininguser behavior
  • 【分类号】TP311.13
  • 【被引频次】2
  • 【下载频次】72
节点文献中: 

本文链接的文献网络图示:

本文的引文网络