节点文献

基于情感信息度量的中文情感文摘研究

The Research on Measuring the Sentiment Information in Chinese Opinion Summarization

【作者】 潘敏;

【导师】 陈水利;

【作者基本信息】 福州大学 , 计算机应用技术, 2014, 硕士

【摘要】 情感文摘是文本倾向性分析的重要组成部分,也是难点之一。情感文摘致力于分析和归纳具有明显倾向性的情感信息。目前,国内外已经开展了很多针对情感文摘的研究工作,并取得了一定的成果。但是这些工作没有充分考虑情感信息元素对情感信息的作用:其一,缺乏考虑极性强度、评价对象、评价短语三者的关联对情感信息情感强弱的影响,而这可能降低情感文摘的精度;其二,对于多样性问题,目前的方法没有考虑情感信息元素的作用,而这可能影响情感文摘的准确度。因此,本文针对这两点展开研究,进行了以下三方面的工作:(1)针对情感文摘中情感信息的情感强弱问题,提出了基于极性强度度量的中文情感文摘生成方法。算法利用逐点互信息原理,充分考虑评价对象、评价短语、极性强度三者之间的关联,度量主流情感信息。根据一个好的情感文摘应该包含主流情感信息的假设来获取情感文摘。实验结果显示,基于极性强度度量情感信息的方法相比没有考虑极性强度的方法,ROUGE-2提升2.21%,ROUGE-SU4提升2.01%,ROUGE-SU9提升2.45%。(2)针对情感文摘中多样性的问题,提出基于情感元素度量的中文情感文摘生成方法。算法首先利用谱聚类算法对语料数据进行分类,建立一个双层模型表示句子和句子、句子和类之间的关系。在该模型的基础上,假设句子的情感信息量,同关联的句子和所属的类有关。根据假设,利用PageRank原理度量主流情感信息。引入情感元素,采用欧氏距离的原理计算句子与句子之间的相似程度和差异程度,以实现情感信息的多样性要求。实验结果显示,基于情感元素度量情感信息的方法相比没有考虑情感元素的方法,ROUGE-2提高3.49%, ROUGE-SU4提高2.97%, ROUGE-SU9提高2.68%。(3)针对情感文摘结果的冗余度问题,提出基于最大边缘相关算法的中文情感文摘生成方法。综合考虑情感信息的情感强弱和多样性问题,运用最大边缘相关算法进行文摘句子的选择,使得选择的文摘句和已选文摘句之间冗余最小。实验结果证明,采用最大边缘相关算法去除冗余信息后的情感文摘,ROUGE-2提升1.32%, ROUGE-SU4提升1.34%,ROUGE-SU9提升1.38%。综上所述,本文在情感强弱、多样性和冗余度三方面,分别利用基于极性强度的方法、基于情感元素的方法以及最大边缘相关算法进行处理,有效提高了中文情感文摘的精度。

【Abstract】 Opinion Summarization is not only an important part of Opinion Mining, but also a difficult task. It endeavors to analyze and summarize the sentiment information of obvious tendentiousness. So far it has attracted lots of attention among domestic and overseas researchers. They have made some achievements in this domain. however, their work ignore the effect of sentiment information elements on sentiment information:Firstly, the lack of consideration on the relationship among reviewers, target and opinion expression may reduce the accuracy of opinion summarization; Secondly, previous work ignores the effect of sentiment information elements on sentences similarity, which may affect the diversity of opinion summarization. So in this paper we will focus on the affection of the two points. The details of the research are proposed as follows:(1)In this paper we propose a novel method to deal with the problem of the emotional strength of the sentiment information, which is based on polarity strength to measure. The algorithm uses the principle of PMI (Pointwise Mutual Information), fully considering the relationship among reviewers, target and opinion expression, to measure the mainstream sentiment information. According to the hypothesis that a good abstract should include mainstream sentiment information, get the opinion summarization. Experimental results show that the method based on polarity intensity measure sentiment information than the method which does not take the polarity intensity into account, ROUGE-2 improved by 2.21%, ROUGE-SU4 improved by 2.01%, ROUGE-SU9 improved by 2.45%.(2)In this paper we propose a novel method to solve the problem of diversity in the opinion summarization, which is based on the elements of sentiment information. The algorithm firstly uses spectral clustering to classify the data, lastly establishes a double layer model of sentences and sentences, sentences and classes.The model assumes that a sentence of sentiment information may be related to link sentences and the corresponding class. According to the hypothesis, use the principle of PageRank calculate to get mainstream sentiment information. We calculate the similarity and difference between sentence and sentence by introducing the sentiment elements, and the principle of Euclidean distance, thus, we can get the diverse opinion summarization of products. Experimental results show that the method based on sentiment elements to measure sentiment information is better than the method which does not take sentiment elements into account, ROUGE-2 improved by 3.49%, ROUGE-SU4 improved by 2.97%, ROUGE-SU9 improved by 2.68%.(3)Considering the problems of redundancy in the opinion summarization, we propose an algorithm based on Maximal Marginal Relevance. The algorithm considers both the emotional strength and diversity of the sentiment information to measure the sentiment information of sentence. Then, use the algorithm of Maximal Marginal Relevance to select sentences in order to make the minimum redundancy between sentence and the selected sentences. The experimental results show that using algorithm of Maximal Marginal Relevance to remove redundant information, ROUGE-2 improved by 1.32%, ROUGE-SU4 improved by 1.34%, ROUGE-SU9 improved by 1.38%.To sum up, On the issue of emotional intensity, diversity and redundancy, the paper propose the method based on the intensity of polarity, the method based on emotional elements and Maximal Marginal Relevance algorithm for processing respectively. The methods effectively improve the precision of Chinese opinion summarization.

  • 【网络出版投稿人】 福州大学
  • 【网络出版年期】2016年 12期
节点文献中: