节点文献

汉语句子相似度计算方法及其应用的研究

Study and Application on Chinese Sentence Similarity Computation

【作者】 周舫

【导师】 郑逢斌;

【作者基本信息】 河南大学 , 应用数学, 2005, 硕士

【摘要】 在中文信息处理中,汉语句子的相似度计算是一项基础而又重要的工作,它直接决定着某些领域的研究发展状况。例如,自动问答系统、基于实例的机器翻译、信息检索、自动文摘等领域,句子相似度计算都是一个非常关键的问题,长期以来一直是人们研究的一个热点和难点。本文在研究汉语句子相似度的过程中,针对汉语由语素构成词语,由词语构成语句的特点,分别对汉语中的义原、词语、句子三个层次的相似度计算进行了研究。这三者层次不同,但是联系密切,由部分构成一个有机的整体,整个计算过程每一步都利用上一步的计算结果。本文主要有如下几个方面的研究成果:1.研究了汉语语句的问句意图,并提出问句意图的提取方法。问句意图的提取是以疑问句为研究对象的,问句类型不同,提取意图的方法也会有所差异。本文的研究是处于问答系统背景下,分析语料根据不同疑问句出现的频率,把问句类型分为三类:特指问句,正反问句,句末语助词问句,从而根据问句的类型针对性的提出相对应的问句意图提取方法。2.研究了汉语词语语义相似度及其计算方法,利用知网提供的丰富语义信息,计算义原相似度,进一步计算基于知网的词语语义相似度。词语相似度是本文句子相似度计算的基础。3.提出多层次多种特征融合的汉语句子相似度计算方法。该算法从多个角度考察语句的相似性,充分利用句子含有的目标层、结构层、语义层等丰富信息,从句子中提取问句意图、关键词集、句子长度、名词个数、动词个数、专有名词个数等多种特征。运用一种简单有效的融合手段,进而获取综合特征,利用综合特征确定句子相似度的值。4.以金融领域自然语言问答系统的模型为实例,体现句子相似度计算在具体应用领域的重要性。这一课题的研究及其成果对于中文信息处理中的多种领域,都将具有一定的参考价值和良好的应用前景。

【Abstract】 Chinese sentence similarity computation is an essential task and widely used in theChinese information processing. It can decide the development of certain relatedresearch directions. For example, in the area of automatic question-answering, EMBT,information retrieval etc, how to compute the sentence similarity is one of the mostimportant problem which is also a hotspot and very difficulty that people study for along time. During the research of Chinese sentence similarity computation, the similaritycomputation that we have studied is focus on three levels: sememe, word and sentence.It is based on the feature of Chinese,that is the word is composed of morphemes, andthe sentence is composed of words. Although three levels are different, from thesimilarity computation to its applications, it is a gradually process with closerelationship as a whole. The main innovative achievements of this paper are as follows: First, the extraction method of question intention is presented which is based onthe research of question intention. Question intention is the surface meaning which thequestion wants to express, and equals to the feature of sentence object layer. Analyzingmuch corpus, the question is divided into three types: question-word questions, A-not-Aquestions, sentence-final particle sentences. Different question types have differentways of extract intention, according to the question type, intention extraction method isput forward. Secondly, we have studied the method of computing semantic similarity betweenChinese words. Using the abundant semantic information supplied by HowNet semanticconcept relation net, we compute the HowNet-based Chinese words semantic similarity. Thirdly, the Chinese sentence similarity computation of multi-levels andmulti-features fusion is presented. This method makes the best use of the sentenceinformation about object level, structure level and semantic level. Several features suchas question intention, keywords set, sentence length, noun number, verb number, propernoun number etc are extracted. And it gains an integrated feature as value of sentencesimilarity computation using fusion algorithm with simplicity and effect. Fourthly, taking the natural language question answering system in financial fieldas the examples, we show the important roles that the Chinese sentence similaritycomputation has been in practice. This research can contribute to some domains in Chinese information processing, itwill be valuable and have good prospect to a certain extend.

  • 【网络出版投稿人】 河南大学
  • 【网络出版年期】2005年 05期
  • 【分类号】TP391.1
  • 【被引频次】65
  • 【下载频次】1727
节点文献中: