节点文献

基于Spark的大数据珊瑚视频推荐系统的设计与实现

Design and Implementation of Big Data Coral Video Recommendation System Based on Spark

【作者】 黄瑞;

【导师】 黄振峰;

【作者基本信息】 广西大学 , 机械工程(专业学位), 2020, 硕士

【摘要】 随着现代科技的迅速发展,人们已经步入了信息化时代。但是,信息时代带来便捷的同时,也带来了信息量庞大而难以筛选的问题。在这个背景下推荐系统应运而生,在电子商务、电影、音乐推荐等各个领域均发挥着巨大的作用,可以将潜在的信息从大规模的数据中挖掘出来并精准的推送给用户。我校某学院实验室拍摄了大量的珊瑚视频数据,却难以精确、迅速查找某一类的实验视频以及视频的具体信息。故本文针对这个难题,通过离线推荐算法、实时推荐算法及相关统计算法来构建一个基于Spark的大数据珊瑚视频推荐系统,按照用户的需求来为之个性化推荐多个珊瑚视频数据并提供搜索引擎、视频分类等功能。首先,对大数据处理技术进行了研究,完成隐语义模型推荐算法的研究和改进,提出了一种过滤dislike因素干扰的隐语义模型推荐算法。该算法根据Slope-One算法计算出用户对物品的厌恶程度评分项,然后使用隐语义模型的ALS矩阵分解算法对其进行交叉过滤,避免了隐语义模型中包含的用户dislike评分项对其推荐效果的干扰。然后基于Spark平台,在Movielens数据集进行了相关的实验验证。结果表明,所提出的算法在准确度、精确率方面均优于ALS算法、基于物品的协同过滤方法以及传统的Slope-One算法。然后根据用户需求对以下几个功能进行实现:包括用户的注册登录模块、用户偏好选择模块、搜索与分类模块、离线推荐模块、实时推荐模块、离线统计模块(热门推荐、评分最多推荐、最新珊瑚视频推荐)、详情展示模块。用户注册与登录模块提供了新用户注册与登录功能;用户偏好选择模块新用户的主观意向解决了冷启动问题;搜索与分类模块为用户提供依据类别检索及模糊查询的功能;离线统计服务模块实现了热门推荐、评分最多、最新珊瑚视频的推荐;离线推荐模块综合所有历史数据精准地进行珊瑚视频的推荐,并保持一定周期的固定性;实时推荐模块是根据用户做出的操作来进行实时性的响应,调度相关的服务,展开珊瑚视频的推荐服务。最后,基于Spark相关大数据技术完成珊瑚视频推荐系统的搭建。前端架构基于Angular js框架为系统提供可视化服务。用户的数据会通过前端发送到系统的综合业务后台,而业务后台会利用了Spring开发框架,Tomcat做web容器,用来响应前端并处理数据。分别使用Mongodb、Redis做核心业务数据库和缓存数据库,部署Elasticsearch作为搜索服务器满足系统的搜索服务。使用Spark SQL实现离线统计服务。基于本文提出的改进的隐语义模型推荐算法预测用户需求,并利用Azkaban调度离线模块,将离线部分实现的结果写入业务数据库中。实时推荐部分通过业务后台搜集日志,利用Flume-ng来采集日志,然后将日志传递给Kafka用以消息缓冲,最后传递到Spark Streaming中进行实时推荐。

【Abstract】 In the era of big data,intelligent recommendation system has brought great convenience to our life.However,while the information age brings convenience,it also brings the problem of information overload.Recommendation system based on big data can effectively solve the problem of information overload,in electronic commerce,movies,music and recommend all areas are playing an important role,it can be viewed according to the user’s information to provide users with the corresponding function,the products and services,allows users to more efficiently get a desired information data from huge amounts of data.A large number of coral video data have been shot in the laboratory of Ocean College of our university,but it is difficult to accurately and quickly find a certain kind of experimental video and the specific information of the video.Therefore,in view of this problem,this paper builds a big data coral video recommendation system based on Spark through offline recommendation algorithm,real-time recommendation algorithm and relevant statistical algorithm,and provides personalized recommendation of multiple coral video data and search engine,video classification and other functions according to users’needs.Firstly,the big data processing technology is studied to complete the research and improvement of the recommendation algorithm of the latent factor model,and a recommendation algorithm of the latent factor model is proposed to filter the dislike factor interference.This algorithm calculates the user’s dislike degree score items according to Slope-One algorithm,and then uses ALS matrix decomposition algorithm of the latent factor model to cross-filter it,avoiding the interference of user dislike score items contained in the latent factor model to its recommendation effect.And based on Spark platform,relevant experimental verification was carried out in Movielens data set.The results show that the proposed algorithm is superior to ALS algorithm,item-based collaborative filtering algorithm and traditional Slope-One algorithm in accuracy and precision.Then,according to user needs,the following functions are implemented:the user registration and login module,user preference selection module,search and classification module,offline recommendation module,real-time recommendation module,offline statistics module(popular recommendation,highest rating recommendation,latest coral video recommendation)and detail display module.The user registration and login module provides the function of new user registration and login.The subjective intention of new users solves the problem of cold start.The search and classification module provides users with functions of category retrieval and fuzzy query.Offline statistics service module realizes the recommendation of popular recommendation,the most rated,and the latest coral videos.The offline recommendation module integrates all historical data to accurately recommend coral videos,and maintains a fixed periodicality.The real-time recommendation module makes real-time response according to the operation made by users,schedules related services,and develops the recommendation service of Coral video.Finally,the establishment of coral video recommendation system was completed based on Spark related big data technology.The front-end architecture provides visual services to the system based on the Angular JS framework.The user’s data will be sent to the integrated business background of the system through the front end,and the business background will use the Spring development framework and Tomcat as the Web container to respond to the front end and process data.Mongodb and Redis are used to build core business database and cache database respectively,and Elasticsearch is deployed as the search server to meet the system’s search service.Offline statistics services are implemented using Spark SQL.Based on the improved recommendation algorithm of the cryptic meaning model proposed in this paper,user requirements were predicted,and the offline implementation results were written into the business database by using the Offline scheduling module Azkaban.The real-time recommendation part collects logs through the business background,uses Flume-NG to collect logs,and then passes the logs to Kafka for message buffer,and finally passes them to Spark Streaming for real-time recommendation.

  • 【网络出版投稿人】 广西大学
  • 【网络出版年期】2025年 09期
  • 【分类号】TP391.3
节点文献中: