节点文献

网络媒体的用户兴趣识别方法研究

A Research of User Interest Recognition Method on Internet

【作者】 王旭

【导师】 于富财;

【作者基本信息】 电子科技大学 , 工程硕士(专业学位), 2020, 硕士

【摘要】 随着社交网络媒体的兴起,越来越多的人通过社交媒体分享新闻见识和个人生活等内容,社交网络活动已经成为人们生活的一个重要组成部分。通过社交网络用户言论识别出社交网络用户的兴趣,可以针对性的开展商业广告投放等活动,也可对用户关心的社会态势进行一定程度的评估和预测。社交网络用语具有文本简短、内容分散、俚语多和语法不规范等特点使得传统基于长文本的兴趣提取方法不再适用,为此本文提出了结合三种语义的神经网络模型对推特用户进行兴趣识别的方法。为对用户兴趣进行细粒度分析,提出了融合注意力机制的神经网络模型完成推文元素抽取与搭配,最终实现推特用户兴趣识别与分析。本文主要工作如下:(1)兴趣推文筛选与预处理。为保证兴趣识别与分析实验的有效性,获取纯净兴趣推文数据集,本文利用Twitter API获取兴趣领域官方认证用户的推文。针对推文语法结构特点,进行了推文预处理。为保证推文中特征丰富度,根据推文长度对推文进行了筛选。最后利用兴趣关键词集与推文进行相似度计算的方法来进行推文打分,按分数排序筛选纯净兴趣推文数据集。(2)兴趣识别方法。为减少推文中噪声词对推文分类的干扰,本文利用半显性语义LLDA模型获取单词隶属兴趣的概率向量,为有效拓展推文中词汇的语义特征,本文提出链接外部知识图谱的显性语义模型来重编单词向量。为充分利用这些特征提升推文分类效果,将隐性语义模型向量、半显性语义模型向量和显性语义模型向量以多通道方式输入到文本卷积神经网络。通过实验对比证明了多语义模型获取的向量对推文分类效果的提升。最终统计用户推文分类结果作为用户兴趣识别结果。(3)兴趣分析方法。本文利用结合注意力机制的用户兴趣细粒度元素抽取方法,对用户兴趣进行分析,提取出用户推文中的兴趣实体与情感词汇。首先将本文提出的融合了文本语法特征的注意力机制嵌入到循环神经网络模型中,利用该模型得到推文中各个词的序列标注结果,提取出兴趣实体与情感词汇。随后根据推文中词汇间距对兴趣实体与情感词汇进行匹配。最后通过实验证明了该方法对于兴趣元素抽取效果的提升以及方法的鲁棒性。

【Abstract】 With the rise of social network media,more and more people share news and personal life through social media.Social network activities have become an important part of people’s life.Through the social network user speech to identify the interests of social network users,we can carry out targeted activities such as commercial advertising,and also can carry out a certain degree of evaluation and prediction of the social situation users care about.The traditional interest extraction method based on long text is no longer applicable due to the characteristics of short text,scattered content,many slang and nonstandard grammar of social network terms.Therefore,this thesis proposes a neural network model combining three kinds of semantics to identify the interest of Twitter users.In order to analyze the user’s interest in a fine-grained way,a neural network model based on attention mechanism is proposed to extract and match the tweet elements,and finally to recognize and analyze the user’s interest.The main work of this thesis is as follows:(1)Screening and preprocessing of tweets of interest.In order to ensure the effectiveness of interest recognition and analysis experiments and obtain pure interest tweet data set,this thesis uses twitter API to obtain the tweets of officially certified users in the interest field.According to the characteristics of the structure of tweet,the preprocessing of tweet is carried out.In order to ensure the feature richness of tweets,the tweets were selected according to the length of tweets.Finally,we use the similarity calculation method of interest keyword set and tweet to score tweets,and filter the pure interest tweet data set according to the score order.(2)Interest recognition method.In order to reduce the interference of noise words in tweets to the classification of tweets,this thesis uses the semi dominant semantic LLDA model to obtain the probability vector of the interest of the word membership.In order to effectively expand the semantic characteristics of the words in tweets,this thesis proposes the explicit semantic model linking the external knowledge map to recompile the word vector.In order to make full use of these features to improve the classification effect,the implicit semantic model vector,semi explicit semantic model vector and explicit semantic model vector are input into the text convolution neural network in a multi-channel way.The experimental results show that the vector obtained by multi semantic model can improve the classification effect of tweets.Finally,users’ tweet classification results are counted as user interest recognition results.(3)Interest analysis method.In this thesis,the user interest fine-grained element extraction method combined with attention mechanism is used to analyze the user interest and extract the interest entity and emotion vocabulary in the user tweet.First,we embed the attention mechanism which integrates the text grammatical features into the cyclic neural network model,and use the model to get the sequence annotation results of each word in the tweet,and extract the interest entities and emotion words.Then,according to the distance between words in the tweet,the interest entities and emotion words are matched.Finally,experiments show that the method improves the extraction effect of interest elements and the robustness of the method.

  • 【分类号】TP391.1;TP18
  • 【被引频次】1
  • 【下载频次】101
  • 攻读期成果
节点文献中: 

本文链接的文献网络图示:

本文的引文网络