节点文献
基于深度学习的中文单文档自动文摘方法研究
Research on Chinese Single Document Automatic Summarization Based on Deep Learning
【作者】 王炜;
【导师】 杨海彤;
【作者基本信息】 华中师范大学 , 计算机技术, 2018, 硕士
【摘要】 自动文摘就是通过编写程序让计算机自动的从原始文档中提取摘要,所提取的摘要必须是全面准确的反映原始文档中心内容并且形式上是简单连贯的短文。基于神经网络的生成式文本摘要一般是通过将原始文档的内容加以“理解”,和抽取式摘要相比,它可以言简意赅的概括文本摘要,语法也很简洁且可读性强。然而在实际应用中,由于技术的限制,现如今一些主流的基于神经网络的生成式文本摘要方法生成的摘要中经常会出现OOV(Out of Vocabulary)问题以及原文中某些重要的语义单元不断地重复于最后的摘要中的问题。造成这种现象的原因主要有:第一,原文中出现次数少但是又极其重要的词、短语等语义单元很难被很好的捕捉到并作为摘要的一部分输出;第二,由于人工神经网络自身的弊端导致生成通顺的语句比较困难。本文以提高中文单文档摘要的生成质量为目的,针对上述自动摘要所面临的问题做了以下两个方面的研究:1.提出了一种融合词抽取的策略来改善一些在原文中极其重要的低频词无法被很好的生成在最后的摘要中。传统的注意力机制只能关注到哪些输入对输出有着更加大的影响,本文的策略通过增加一个词表,该词表在原有语料库的词表的基础上加上所有原文中包含的词但是初始词表中没有包含的词,这样在生成词的时候就可以考虑到原文中低频词的概率分布并生成这些词作为最后的摘要。实验结果表明该策略能在LCSTS以及NLPCC2017两个数据集上相较传统的抽取式方法以及基于基础的端到端的神经网络模型更好地结果。2.提出了一种消重策略来改善摘要中单个词的重复出现的问题。每次生成当前单词的时候都会将前一个生成摘要单词作为输入,所以在解码过程中,会出现注意力过分其中在编码器的某一部分,从而造成了错误,然后就出现无休止的短语重复,基于这个问题,我们加入了新的融合机制,在每次生成词的时候对之前“关注过”的词在这一轮给予一定的“惩罚”,这样就可以避免之前由于生成过的单词在这一轮再次受到较高的“关注度”。实现表明该策略在生成的摘要中能有效地避免重复出现某个重要的单词,使生成的语句可读性更好。
【Abstract】 Automatic summarization is the use of computers to automatically extract summarization from original documents by programming,the summarization must be simple and consistent,it also can reflect the contents of documents fully and accurately.The abstractive summaries based on neural network is used to "understand" the main content of original articles.Compared with the extractive summaries,it can summarize texts concisely,and the grammar is also very simple,the summarization is very readable.However,in practical applications,due to technical limitations,the OOV(Out of Vocabulary)problem often occurs in some abstractive summaries which generated based on neural network,furthermore,some important semantic units in the original article are constantly repeat themselves in the final summary.There are two main reasons for this phenomenon:First,there are some low frequency but extremely important words in the original text,these words can hardly be captured and output as part of the summarization;secondly,due to artificial nerves the drawbacks of the neural network,it’s difficult to generate fluent sentences.This paper aims to improve the generation quality of Chinese single document summarization.In view of the problems faced by the above automatic summarization,the following two aspects are studied:1.A strategy of fusion word extraction is proposed to improve some low-frequency words which are extremely important in the original text,which can not be well generated in the final summary.The traditional attention mechanism can only focus on which input has a greater impact on the output.The strategy of this article is by adding a word list,which is added to all the words contained in the original text on the basis of the original corpus of the corpus,but the words are not included in the initial word list,which can be considered when the word is generated.The probability distribution of low-frequency words in the original text is generated and these words are generated as final summaries.The experimental results show that the strategy can better result than the traditional decimation method and the base-to end-based neural network model on the two data sets of LCSTS and NLPCC2017.2.A strategy of weight elimination is proposed to improve the repetition of single words in summarization.Each time the current word is generated,the first generated summary word is used as input,so during the decoding process,there will be an excessive attention in one part of the encoder,resulting in the error,and then the endless phrase repetition.Based on this problem,we add a new fusion mechanism.Each time a word is generated,the word "concerned" is given a certain "punishment"in this round,so that it can avoid a higher "attention" in this round because of the generated words.Implementation shows that the strategy can effectively avoid duplication of an important word in the generated summary,making readability of the generated statement better.
【Key words】 Automatic summary; Word extraction; Neural Networks; attention; decode;
- 【网络出版投稿人】 华中师范大学 【网络出版年期】2019年 01期
- 【分类号】TP391.1;TP181
- 【被引频次】7
- 【下载频次】343