节点文献
面向词对复杂依赖关系建模的结构化情感分析研究
Research on Structured Sentiment Analysis with Word Pair Complex Dependency Modeling
【作者】 石文轩;
【导师】 姬东鸿;
【作者基本信息】 武汉大学 , 网络空间安全, 2023, 硕士
【摘要】 近年来,随着社交媒体的快速发展,人们的情感和观点表达也越来越多样化和复杂化,为了更准确和及时的掌握公众对政治事件、公共话题等方面的看法和态度,细粒度的情感分析研究受到越来越多研究者的关注。细粒度情感分析任务旨在更准确的理解文本中的情感和观点表达,一般涉及对四类情感元素的分析,包括观点持有者、观点描述词、评价对象以及情感极性。然而,针对这四类元素,情感分析领域发展出了越来越多的细分子任务,每一个子任务通常只针对一部分情感元素进行研究,这种发展趋势妨碍了从一个更整体的视角来研究情感分析领域存在的问题和挑战。针对这种任务细分的趋势,结构化情感分析任务被提出,其不仅要求同时抽取文本中包含的四类情感元素,还需要分析它们之间的语义关系,输出情感四元组。然而,传统的基于序列标注的方法较难识别结构化情感分析任务中存在的具有长文本片段和重叠嵌套的情感元素。受启发于词对关系建模方法在信息抽取领域的亮眼表现,部分研究者探索了该方法在结构化情感分析任务中的应用,基于特殊的词对关系标签设计方案,可以将情感四元组抽取任务转化为词对关系分类和解码任务。然而,由于存在重叠嵌套和长文本片段的元素,以及跨从句的元素间关系,结构化情感分析数据中的词对具有更复杂的依赖关系,给标签和模型的设计带来了很大的挑战,之前的相关方法都存在词对复杂依赖关系建模的不充分问题,以及因词对标签设计不平衡影响元素关系预测的问题。针对上述挑战,本文将以词对复杂依赖关系建模为目标,进一步探索更适合结构化情感分析任务的词对标签设计方案以及相应的模型架构,具体地,本文提出的主要解决方案如下:(1)针对现有相关工作词对标签设计的不平衡问题,本文提出了一个解耦标签设计方案,通过设计两套词对标签集解耦情感四元组预测和词对复杂依赖关系建模两个任务,其中,一套比较稀疏但设计较为平衡的标签集用于在模型预测层解决重叠嵌套情感四元组预测任务,并使模型的元素片段预测和元素关系预测两方面能力更加平衡;另一套比较丰富的标签集充分利用了数据集中已标注的任务相关知识,用于模型隐层作为词对复杂依赖关系建模任务的学习目标。(2)针对现有模型设计中存在的词对复杂依赖关系建模的不充分问题,本文在模型隐层构建了一个多视图的图卷积神经网络,其中,每一个视图对应一个类别的词对关系标签,并以词嵌入为图的节点,对应的词对关系建模结果为图的边。基于此,本文将词对复杂依赖关系建模任务形式化为模型基于学到的某类别词对关系自动构建对应视图的词对邻接矩阵,实现了对复杂依赖关系更深层次的建模。(3)针对结构化情感分析任务中存在的词对复杂句法依赖关系的问题,本文基于外部解析器的依存句法分析和成分句法分析的结果,构建了词对关系形式的句法标签,从而将句法知识引入模型隐层的词对复杂依赖关系建模任务中,实现句法知识和模型更充分的融合,缓解模型在处理跨从句元素间关系问题时的局限性。(4)针对句法知识融合过程中存在的误差传播和任务目标不一致问题,本文对模型做了进一步的改进,首先,本文提出了词对自适应阈值技术,在词对句法知识建模的过程中实现自适应的筛选,提升模型的鲁棒性;然后,本文在隐层构建了一个多视图图注意力神经网络以更好的区分不同词对关系的重要性。综上所述,本文进一步探索了词对关系建模方法在结构化情感分析任务中的应用,针对该任务数据中存在的词对复杂依赖关系的特点,提出了新的词对标签设计方案和相应的模型架构,最后,本文在五个结构化情感分析相关数据集上进行了实验,本文提出的模型均取得了当时最好的效果,推动了结构化情感分析领域的前沿发展,可为相关研究者提供参考和借鉴。
【Abstract】 In recent years,with the rapid development of social media,people’s expressions of emo-tions and opinions have become increasingly diverse and complex.In order to accurately and timely grasp public opinions on political events and public topics,fine-grained sentiment analy-sis has received increasing attention from researchers.The task of fine-grained sentiment anal-ysis aims to accurately understand the expression of emotions and opinions in text,generally involving the analysis of four types of sentiment elements,including holders,expressions,tar-gets,and sentiment polarity.However,for these four types of elements,the sentiment analysis field has developed more and more subtasks,each of which usually only focuses on researching a part of the sentiment elements.This trend hinders the study of problems and challenges in the sentiment analysis field from a more holistic perspective.In response to the trend of task subdivision in sentiment analysis,the task of structured sentiment analysis has been proposed.It involves extracting four categories of sentiment ele-ments from the text and analyzing their semantic relationships to generate sentiment quadruples.However,traditional sequence labeling methods struggle to recognize the long text segments and overlapping nested sentiment elements in structured sentiment analysis tasks.Inspired by the impressive performance of word-pair relation modeling method in information extraction,some researchers have explored its application in structured sentiment analysis tasks.Based on a specially designed word-pair relation labeling scheme,sentiment quadruple extraction tasks can be transformed into word-pair relation classification and decoding tasks.However,due to the more complex dependency of word-pair relationships in structured sentiment analysis data,caused by the presence of overlapping nested elements and elements spanning multiple clauses,the design of labels and models faces significant challenges.Previous related methods suffer from insufficient modeling of word-pair complex dependency relationships and issues related to unbalanced word-pair label design affecting the prediction of element relationships.To address these challenges,this paper will focus on modeling word-pair complex dependency relationships,and further explore word-pair label design schemes and corresponding model ar-chitectures that are more suitable for structured sentiment analysis tasks.Specifically,the main proposed solutions of this paper are as follows:(1)To solve the problem of unbalanced word-pair labeling design in existing relevant work,this paper proposes a decoupling labeling design scheme that decouples sentiment quadruple prediction and word-pair complex dependency modeling into two tasks by designing two sets of word-pair labeling schemes.One set of sparser but more balanced label sets is used to solve the overlapping sentiment quadruple extraction task at the model prediction layer,making the model’s ability to predict element fragments and element relationship more balanced.Another set of richer label sets fully utilizes task-related knowledge already labeled in the dataset,used as the learning objective for the model’s hidden layer to model the word-pair complex dependency.(2)To solve the problem of insufficient word-pair complex dependency modeling in the existing model design,this paper constructs a multi-view graph convolutional neural network in the model’s hidden layer,with each view corresponding to a category of word-pair relationship labels,and each word represented as a node in the graph and the corresponding word-pair rela-tionship modeling results as the graph’s edges.Therefore,this paper formalizes the word-pair complex dependency modeling task as the model automatically constructing the correspond-ing word-pair adjacency matrix based on the learned class of word-pair relationships,thereby achieving deeper modeling of word-pair complex dependency.(3)To solve the problem of complex word-pair syntactic dependency relationships,this pa-per constructs a syntactic labeling scheme for word pairs based on the results of external parsers’dependency and constituent syntax analysis,introducing syntactic knowledge into the model’s hidden layer’s word-pair complex dependency modeling task,realizing a more comprehensive fusion of syntactic knowledge and the model,and mitigating the model’s limitations in dealing with element relationships between clauses.(4)To solve the problem of error propagation and inconsistent task goals in the process of word-pair syntactic knowledge fusion,this paper further improves the model.First,this paper proposes word-pair adaptive threshold technology,achieving adaptive filtering during word-pair syntactic knowledge modeling,improving the model’s robustness.Then,this paper constructs a multi-view graph attention neural network in the hidden layer to better distinguish the importance of different word-pair relationships.In summary,this paper further explores the application of word-pair relation modeling in structured sentiment analysis tasks.To address the complex dependency relationships among word pairs in the task data,a new word-pair labeling scheme and corresponding model architec-ture are proposed.Finally,experiments on five structured sentiment analysis datasets demon-strate that the proposed model achieves state-of-the-art performance at the time of publication,which advances the forefront of the structured sentiment analysis field and provides a reference and inspiration for related researchers.
- 【网络出版投稿人】 武汉大学 【网络出版年期】2026年 07期
- 【分类号】TP391.1