节点文献

面向新媒体的中文意见挖掘研究

Chinese Opinion Mining for New Media

【作者】 张鑫;

【导师】 张梅山; 谢朋峻;

【作者基本信息】 天津大学 , 电子信息, 2022, 硕士

【摘要】 互联网的发展促使了各种新媒体的诞生,人人都可参与表达和传播,分享自己的观点和意见,以微信、微博等为代表的新媒体社交平台的海量数据中蕴含了丰富的信息和价值。对于文本的意见挖掘是一种分析和抽取文本中的意见相关信息的技术,是近年来较为火热的研究方向,但是针对于中文新媒体的研究还较少。本文面向于新媒体数据上的中文意见挖掘,以微博文本数据为切入点,探索面向新领域新话题的中文意见挖掘系统的快速构建方法。本文主要研究了两个意见挖掘的相关任务,分别是意见表达式识别和细粒度联合意见挖掘。本文主要工作分为两个层面,在数据上,基于微博文本构建了意见表达式识别和细粒度意见挖掘数据集,在方法上,提出了基于众包学习的意见表达式识别模型和跨语言跨领域的细粒度联合意见挖掘模型。具体工作如下:首先,基于新冠肺炎疫情期间微博文本,探索了众包数据标注方法,以较低的成本和较快的速度构建了带有平行的众包标签的意见表达式识别数据集,然后又标注了相应的黄金标签,用于模型研究。基于上述成果,以专家标注的方式进行二阶段的意见目标和持有者标注,构建了金标细粒度意见挖掘数据集。其次,针对意见表达式识别问题的众包数据特性,提出了将众包学习建模为领域自适应问题的思想,并应用此思想提出了标注者表示学习模型Annotator-adapter以解决序列标注方法的众包学习,然后提出了标注者混合策略Annotator-mixup对表示学习模型训练过程进行增强,在构建的中文众包意见表达式识别数据集上取得了领先的性能,且与金标模型的差距是可以接受的。最后,针对无标注数据的细粒度联合意见挖掘问题,提出了跨语言跨领域的迁移学习方法,基于MPQA英文数据的自动语料翻译和mBERT多语言表示进行跨语言迁移,基于无标注疫情微博文本构建自学习伪语料和参数生成网络多语料融合进行跨领域迁移,最终达成跨语言跨领域的联合迁移,实现无监督的细粒度意见挖掘,在构建的细粒度意见挖掘数据集上取得了较为可观的性能。通过以上研究,本文完成了对低成本快速构建新领域意见挖掘系统的初步探索,为未来的研究奠定了语料基础和方法基础。

【Abstract】 The development of the Internet has led to the birth of various new media,where everyone can participate in expression and communication,and share their views and opinions.The massive data of new media social platforms represented by We Chat and Weibo contain rich information and values.Opinion mining for text is a technique to analyze and extract opinion-related information from text,which is a hot research direction in recent years,but there are few studies for Chinese new media.This paper takes Weibo data as the starting point to explore the fast construction of Chinese opinion mining systems for new topics in new fields.This paper focuses on two opinion mining related tasks,which are opinion expression identification and fine-grained joint opinion mining.The main work of this paper is divided into two levels,in terms of data,the opinion expression recognition and fine-grained opinion mining datasets are constructed based on Weibo texts;and in terms of methods,the opinion expression recognition model based on crowdsourcing learning and the cross-lingual cross-domain fine-grained joint opinion mining model are proposed.The specific works are as follows:First,based on the Weibo text during the COVID-19 epidemic,a crowdsourced data annotation method is explored to construct an opinion expression recognition dataset with parallel crowdsourced labels at a low cost and fast speed.Then the corresponding gold-standard labels are annotated for research.Based on the above results,the secondstage opinion target and holder annotating with expert labeling is performed to construct the gold label fine-grained opinion mining dataset.Second,for the crowdsourcing data characteristics of the opinion expression recognition problem,the idea of treating crowdsourcing learning as a domain adaptive problem is proposed.The annotator representation learning model Annotator-adapter is applied to solve the crowdsourcing learning of sequence labeling method.Then the Annotator-mixup strategy is proposed to enhance the training of the representation learning model.These methods achieve leading performance on the constructed Chinese crowdsourced opinion expression identification dataset.The performance gap with the gold standard model is acceptable.Finally,for the fine-grained joint opinion mining problem in unsupervised setting,a cross-lingual and cross-domain migration learning method is proposed.The cross-lingual transfer is based on the automatic corpus translation of the English MPQA dataset and m BERT multilingual representation.The cross-domain transfer is based on the construction of self-learning pseudo-corpus in unlabeled COVID-19 texts and the parameter generation network for the multi-corpus fusion.Finally,the cross-lingual and cross-domain joint transfer is performed for the unsupervised fine-grained opinion mining.It achieves a considerable performance on the constructed fine-grained opinion mining dataset.Through the above research,this paper completes the preliminary exploration of low-cost and fast construction of opinion mining systems for new domains.It lays the corpus basis and methodological foundation for future research.

  • 【网络出版投稿人】 天津大学
  • 【网络出版年期】2025年 07期
  • 【分类号】TP391.1
节点文献中: 

本文链接的文献网络图示:

本文的引文网络