节点文献

面向非平行数据的文本情感迁移及生成研究与实现

Research and Implementation of Sentiment Transfer and Generation on Non-parallel Data

【作者】 刘悦;

【导师】 詹志强;

【作者基本信息】 北京邮电大学 , 信息与通信工程, 2020, 硕士

【摘要】 在深度学习领域中,相比于图像的生成和风格迁移,文本的风格迁移和语义控制还存在很多挑战。文本风格迁移即针对文本表达的情感倾向、时态或话题等属性进行改写从而生成符合目标特性的文本。这其中涉及的语义控制、风格控制等技术对于可控文本生成至关重要。本文针对情感这一广受关注的属性研究并实现文本改写算法。然而平行语料的标注耗费资源且依赖于此的有监督模型效果受制于标注数据的数量和质量,因此本论文聚焦于非平行语料的文本情感迁移与生成算法。这其中文本情感和语义内容的拆分和重组、不同情感间映射关系的建模都颇具挑战。本文针对现存方法在情感转换准确性及内容生成方面的不足提出以下两种改进的深度学习模型。1)基于注意力机制及生成对抗原理的CAM-DAE模型,其通过注意力机制完成情感和内容的拆分,同时通过对抗学习完成无监督情感改写。实验证明该模型性能相比基线模型有明显提升。2)基于无监督回译原理的BTM模型,该模型对两类文本间的转换关系进行建模,通过回译产生伪平行语料进而迭代优化模型。该模型相比基线模型性能明显提升,尤其在文本语义保留方面取得了突破性进展。此外,本文基于BTM模型设计并开发了一款自动评论生成系统。此系统能够自动生成特定主题及情感的评论集,辅助用户进行评论发表。

【Abstract】 In the field of deep learning,compared to the generation and style transfer of images,there are still many challenges in style transfer and semantic control in text generation.Text style transfer is rewriting a text to alter a specific attribute,such as sentiment,tense and topic,while preserving the content.In this problem,technologies involved such as semantic control and style control are of great significance for further research on controllable text generation.This paper focuses on the text rewriting algorithm for sentiment,a widely-attended attribute of text.However,parallel data often consumes a lot of resources and models relying on it is subject to the quantity and quality of labeled data.Therefore,this paper focuses on sentiment transfer and text generation using non-parallel data which is quite challenging to separate and reconstruct sentiment and content.This paper introduces two improved models based on non-parallel data.1.CAM-DAE model:based on the attention mechanism and generative adversarial nets.It separates sentiment and content more precisely through the Attention mechanism and guides the decoder to complete unsupervised sentiment rewriting through the adversarial network.Experimental results show that the performance of this model is significantly improved compared to the baseline models.2.BTM model:based on unsupervised back translation,which abstracts sentiment transfer into a translation problem between two sets of text.Pseudo-parallel corpora are continuously generated through an iterative translation process,which is used for training the back-translation model.Experimental results prove that the model have achieved more advanced results than the baseline models,especially in terms of content preservation.In addition,based on the BTM model,an automatic comment generation system has been designed and developed in this paper.This system can automatically generate comment sets on specific topics and sentiments to assist users in posting comments.

节点文献中: 

本文链接的文献网络图示:

本文的引文网络