节点文献

基于Word2Vec的情感词典自动构建与优化

Automatic Construction and Optimization of Sentiment Lexicon Based on Word2Vec

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 杨小平张中夏王良张永俊马奇凤吴佳楠张悦

【Author】 YANG Xiao-ping;ZHANG Zhong-xia;WANG Liang;ZHANG Yong-jun;MA Qi-feng;WU Jia-nan;ZHANG Yue;School of Information,Renmin University of China;

【机构】 中国人民大学信息学院

【摘要】 情感词典的构建是文本挖掘领域中重要的基础性工作。近几年,情感词典的极性标注从二元褒贬标注向多元情绪标注发展,词典的领域特性也日趋明显。但是情感类别的手工标注不但费时费力,而且情感强度难以得到准确量化,同时对领域性的过分关注也大大限制了情感词典的适用性[1]。通过神经网络语言模型对大规模中文语料进行统计训练,并在此基础上提出了基于转换约束集的多维情感词典自动构建方法;然后研究了基于词分布密度的感情色彩消歧方法,对兼具褒贬意味词语的感情极性进行区分和识别,并分别计算两种感情色彩下的情感类别与强度;最后提出基于多个语义资源的全局优化方案,得到包含10种情绪标注的多维汉语情感词典SentiRuc。实验证实该词典1)在类别标注检验、强度标注检验、情感消歧效果及情感分类任务中均具有良好的效果,其中的情感强度检验证实该词典具有极强的情感语义描述力。

【Abstract】 The construction of sentiment lexicon plays an important role in text mining.In recent years,the lexicon annotating format gradually evolves from binary annotation to multiple annotation,and sentiment lexicons of a single specific domain have caught more and more attentions of researchers.However,manual annotation costs too much labor work and time,and it is also difficult to get accurate quantification of emotional intensity.Besides,the excessive emphasis on one specific field has greatly limited the applicability of domain sentiment lexicons[1].This paper implemented statistical training for large-scale Chinese corpus through neural network language model,and proposed an automatic method of constructing a multidimensional sentiment lexicon based on constraints of Euclidean distance group.In order to distinguish the sentiment polarities of those words which may express either positive or negative meanings in different contexts,we further presented a sentiment disambiguation algorithm to increase the flexibility of our lexicon.Lastly,we presented a global optimization framework that provides a unified way to combine several human-annotated resources for learning our 10-dimensional sentiment lexicon SentiRuc.Experiments show the superior performance of SentiRuc lexicon in category labeling test,intensity labeling test and sentiment classification tasks.It is worth mentioning that in intensity label test,SentiRuc outperforms the second place by 23%.

【基金】 国家自然科学基金(71271209);北京市自然科学基金(4132067);教育部人文社会科学青年基金(11YJC630268);数字出版技术国家重点实验室开放课题资助
  • 【文献出处】 计算机科学 ,Computer Science , 编辑部邮箱 ,2017年01期
  • 【分类号】TP391.1
  • 【被引频次】127
  • 【下载频次】1595
节点文献中: 

本文链接的文献网络图示:

本文的引文网络