节点文献

基于选择性编码模型的自动文本摘要技术研究

Research of Automatic Text Summarization Based on Selective Encoding Model

【作者】 陈洪

【导师】 王玉明;

【作者基本信息】 华中科技大学 , 信息与通信工程, 2020, 硕士

【摘要】 自动文本摘要技术旨在从文本中提取重要信息来自动生成摘要,它能够对文本的信息进行有效压缩与提炼。这在信息急剧增长的互联网时代,可以有效地解决信息过载问题,从而极大地提高人们浏览和处理信息的效率。本文在对生成式摘要方法进行研究时发现,现有模型主要采用编码-解码的方式生成摘要,而这种方式缺少了对文本信息的选择过程,导致有大量与摘要无关的冗余信息对生成摘要造成干扰。因此,本文认为目前的主要挑战在于从原文本中有效地选择出重要信息,并忽略掉非关键信息。针对上述挑战,本文基于选择性编码模型(Selective Encoding for Abstractive Sentence Summarization,SEASS)设计了一个关注主题的选择性编码模型(Topic-aware Selective Encoding Model,TSEM)。TSEM 模型将文本的主题信息作为先验知识分别融入到编码器和选择门网络中,来提高模型对文本信息的理解及选择能力。此外,本文还提出了一个对照机制,该机制可以使模型在训练的过程中充分考虑到摘要与原文内容之间的差异,从而进一步提升模型对原文本信息的选择能力。为了验证模型的有效性,本文在广泛应用于自动文本摘要技术的Gigaword公开数据集上进行了对比实验,并使用ROUGE评价方法对实验结果进行了评测。同时,对模型生成的摘要进行了摘要重复率和主题相似度分析,并结合具体的案例定性分析了模型的效果。实验结果表明,融入了主题信息并且使用了对照机制的TSEM模型在ROUGE评分上能够取得更好的结果,相对于原SEASS模型在ROUGE-1,ROUGE-2和ROUGE-L评分上分别提升了 1.39%、2.12%和1.22%,其生成的摘要可以包含更多原文的关键信息,和原文的主旨更相符。

【Abstract】 Automatic text summarization technology aims at extracting important information from text to automatically generate summary,which can effectively compress and refine the information of text.This can effectively solve the problem of information overload in the Internet age with the rapid growth of information,thus greatly improving the efficiency of people’s browsing and processing of information.In this paper,we found that existing models mainly use the encoding-decoding method to generate summary,however,this method lacks the process of selecting text information,which results in a large number of redundant information unrelated to summary that interfere with the generation of summary.Therefore,the main challenge is to effectively select the important information from the original text and ignore the non-critical information.In this paper,a Topic-aware Selective Encoding Model(TSEM)is designed based on the Selective Encoding Model(SEASS)to address the above challenge.The TSEM model integrates the topic information of the text into the encoder and the selective gate network as the prior knowledge to improve the model’s understanding and selection of the text information.In addition,a comparison mechanism is proposed in this paper,which can make the model fully consider the difference between the summary and the original text in the process of training,so as to further improve the model’s ability to select the original text information.In order to verify the validity of the model,this paper conducts a comparison experiment on Gigaword public data sets widely used in automatic text summarization technology,and evaluates the experimental results using ROUGE evaluation method.At the same time,the thesis makes repetition rate and topic similarity analysis of the summary generated by the model,and makes a qualitative analysis of the effect of the model based on specific cases.The experimental results showed that the TSEM model incorporating the topic information and using the control mechanism could achieve better results in ROUGE scores.Compared with the original SEASS model,the TSEM model improved by 1.39%,2.12%and 1.22%in rouge-1,rouge-2 and rouge-1 scores,respectively.The generated summary could contain more key information of the original text and be more consistent with the main idea of the original text.

  • 【分类号】TP391.1
  • 【下载频次】35
节点文献中: 

本文链接的文献网络图示:

本文的引文网络