节点文献

电影评分的影响因素实证分析

An Empirical Analysis of the Influencing Factors of Film Rating

【作者】 陈琦;

【导师】 高文武;

【作者基本信息】 安徽大学 , 应用统计(专业学位), 2023, 硕士

【摘要】 近些年随着网络的发展,影评网站上的评论数不胜数,用户和制片商难以逐条浏览评论来了解大众对电影的看法和态度,基于此本文通过对电影评论进行情感分析和属性挖掘,进而分析出电影评分与电影评论情感以及电影评论属性之间的关系,即可帮助观影者购买符合自己需求的电影票。本文主要从三个方面进行研究:(1)为了实现影评文本的情感分类,本文首先使用词嵌入模型提取文本的特征向量,并构建不同的分类模型,从中选出了效果最好的Text CNN模型后,利用改进的损失函数Poly Loss替换Text CNN模型的Cross Entropy Loss,进一步提升了模型的效果。为了更好的挖掘上下文语义,使用Albert模型提取特征向量,在该模型的输入层中输入了位置信息,并采用transformer中的自注意力机制以此获得动态的词向量,将所得结果作为Text CNN模型的输入,最终的实验结果表明本文构建的Albert-PL-Text CNN模型性能最佳,与其他的模型相比,在各个评估指标上的结果都有了一定程度的提高。(2)为了挖掘出评论文本的属性,本文使用词性分析和依存句法分析结合的方法抽取评论文本中的属性,并将其作为侯选属性,然后利用K均值聚类法对侯选属性进行聚类,最后使用少量人工人为挑选出核心的属性,并将侯选属性与核心属性进行相似度匹配,得到最终的属性分布情况。(3)对电影评分的影响因素进行分析。将电影评论的各个属性以及电影评论的整体情感倾向作为预测电影评分的特征,利用随机森林算法进行电影评分的预测,并得到特征重要度评估结果,根据特征重要度可知每种类型电影的评分受哪些因素影响最大,日后可根据电影评论中的属性以及电影评论的整体情绪倾向预判电影评分,从而给未观看电影的观众的购票决策提供一定的帮助。

【Abstract】 With the development of the Internet in recent years,there are countless comments on the film review website.It is difficult for users and producers to browse comments one by one to understand the public’s views and attitudes of movies.Based on this article,this article conducts emotional analysis and attribute mining of movie reviews,and then then Analyzing the relationship between the movie score and the movie reviews and the attributes of the movie review can help the viewer to buy a movie ticket that meets his needs.This article mainly studies from three aspects:(1)In order to realize the emotional classification of film review text,this article first uses the word embedding model to extract text feature vector and builds different classification models.The Cross Entropy Loss of the Text CNN model was replaced by the improved loss function Poly Loss,which further enhanced the effect of the model.In order to better excavate the context,use the Albert model to extract feature vectors,enter the location information in the input layer of the model,and use the self attention mechanism in the Transformer to obtain a dynamic word vector,and use the result of the results as the Text CNN model of the Text CNN model.Input,the final experimental results show that the Albert-PL-Text CNN model constructed in this article is the best performance.Compared with other models,the results on each evaluation indicator have improved to a certain extent.(2)In order to dig out the attributes of the comment text,this article uses a combination of part-of-speech analysis and dependency parsing method to extract the attributes in the comment text,and uses them as candidate attributes,and then uses the K-means clustering method to cluster the candidate attributes,and finally uses a small number of artificial selection of the core attributes,and matches the similarity between the candidate attributes and the core attributes to obtain the final attribute distribution.(3)Analyze the influencing factors of film scores.Take the various attributes of the movie and the overall emotional tendency of the movie review as the characteristics of the forecast movie score,use the random forest algorithm to predict the movie score,and get the results of the characteristics of the characteristics of characteristics.What factors are affected most,in the future,they can predict movie scores based on the attributes in the movie reviews and the overall emotional tendency of movie reviews,so as to provide some help to the ticket purchase decisions of the audience without watching the movie.

  • 【网络出版投稿人】 安徽大学
  • 【网络出版年期】2025年 03期
  • 【分类号】TP391.1;J905
节点文献中: 

本文链接的文献网络图示:

本文的引文网络