节点文献
基于自适应聚类的中文自动文摘研究
A Study of Chinese Text Summarization Based on Adaptive Clustering Algorithm
【作者】 胡珀;
【导师】 何婷婷;
【作者基本信息】 华中师范大学 , 计算机软件与理论, 2005, 硕士
【摘要】 自动文摘是自然语言处理领域的一项重要的研究课题。作为解决目前信息过载问题的一种辅助手段,它能在一定程度上弥补传统的信息检索技术在应对信息过载危机时所表现出来的种种缺憾,帮助用户提高信息检索的速度,节省重要信息的浏览时间。 中文自动文摘的研究如火如荼地开展了近20年,令人鼓舞的成果层出不穷。然而,在欣然地享受着这些精彩成果的同时,若干可能会影响自动文摘效果的潜在问题正逐渐被越来越多的研究人员所重视。 以句子权值排序作为抽取依据的文摘方法是中文自动文摘领域广泛采纳的一种典型方法,它简单易行且适用面宽。然而由于摘要对象的多样性,它的缺陷也正变得日趋明显。其主要表现在它所产生的文摘往往很难在主题覆盖度与冗余之间达到某种平衡,常常出现主题遗漏或内容冗余等问题。因此,针对不同题材文本具有不同的潜在主题结构这一现象,如何自适应地发现不同文本潜在的主题将会对现有文摘方法的摘要效果产生积极的影响。此外,在采用统计学方法构造自动文摘系统的过程中,建立各级语言单元的特征向量往往是一个基础性环节。而在实际的摘要实验中,我们发现建立的特征向量的维数常常偏大,达到几百维甚至上千维,而这无疑会制约后续摘要算法的效率。因此,对这些特征向量进行一定程度的降维处理将必不可少。 致力于对上述问题的解决,我们尝试性地提出了一种基于自适应聚类的中文自动文摘方法。在该方法中,我们采用了如下四种关键技术: 关键技术1 基于无监督特征抽取的文本各级语言单元的特征向量表达 关键技术2 基于自适应段落聚类的文本潜在主题的自动发现 关键技术3 基于主题语义相似度计算的文本主题代表句的自动选取 关键技术4 基于表达熵的文摘冗余的量化评价 为了验证提出的中文自动文摘方法的可行性和有效性,我们从国家语委现代汉语语料库中随机选取了30篇不同题材的文本作为实验文本,分别采用提
【Abstract】 Automatic summarization is an important research issue in natural language processing. Now, more and more researchers over the world are paying attention to this area. For one thing, automatic summarization technology can compensate the pitfalls of traditional information retrieval technology in a certain degree when dealing with information overload problem; for another, automatic summarization technology can release users’ browsing pressure.There are still a lot of problems in the research of Chinese document summarization. For instance, a lot of researchers are adopting the traditional summarization method, which extracts relevant sentences from the entire text according to each sentence’s score. However, these methods do not take the document’s thematic structure into account, so the generated summaries using these methods will cover only those main themes while neglecting the others, and sometimes have a high level of redundancy. In addition, in the course of developing a practical automatic summarization system, dimensionality reduction of various linguistic units will be a fundamental and important step.In this paper, we propose a Chinese summarization method based on adaptive clustering algorithm. Four key technologies are adopted in this method:The key technology one: Feature vector representations of various linguistic units based on unsupervised feature extractionThe key technology two: Discovery of latent themes based on adaptive clustering algorithmThe key technology three: Selection of representative sentences from different themes using theme-sentence similarity calculationThe key technology four: Quantitative evaluation of summary’s redundancy based on representation entropyWe choose thirty different genres of documents as experimental samples from the Modern Chinese Corpus of State Language Commission. By using the proposed method and traditional baseline method, we get the relevant results. And the experimental results indicate that the proposed method is more effective and efficient when dealing with various genres of documents, for it can balance the generated summary’s thematic coverage and redundancy in a certain degree.
【Key words】 automatic summarization; thematic discovery; unsupervised feature extraction; clustering; representation entropy;
- 【网络出版投稿人】 华中师范大学 【网络出版年期】2005年 05期
- 【分类号】TP391.1
- 【被引频次】6
- 【下载频次】342