节点文献
中文自动分词系统的研究
Study on the System of Chinese Automatic Word Segmentation
【作者】 朱珣;
【导师】 何婷婷;
【作者基本信息】 华中师范大学 , 计算机应用技术, 2004, 硕士
【摘要】 中文自动分词系统是利用计算机对中文文章进行自动分词、识别的计算机应用系统,它包括基本的自动分词方法、歧义处理和命名实体的识别等基本模块,其各部分相互依赖,共同决定该系统的质量、价值和应用水平。 中文自动分词方法分为机械分词方法和非机械分词方法。最大正向匹配法、逆向最大匹配法和逐词遍历法是三种最基本的机械分词方法。另外八种机械分词法只是在基本分词方法的基础上采用了一些技巧,它们不是纯粹意义的机械分词方法。专家系统方法是一种基于规则的分词方法,而神经元网络方法则将人工神经网络的基本原理应用于计算机汉语分词。 根据国内外对自动分词方法的研究和一些实用系统的设计,本文给出了自动分词系统的理论模型CWSM:M(F,W,T,K)的概念,即机械分词方法+分词词典+汉语言文本+知识库,并介绍了自动分词系统的评价标准。 分词过程中歧义的产生主要是由计算机分词产生的特有歧义、自然语言中的二义性歧义和由分词词库大小引起的歧义等三类组成。歧义字段可从三个方面进行分类。从分词的切分结果可分为两类:真歧义和伪歧义;从切分歧义所需的知识层次,可分为三类:语法歧义、语义歧义和语用歧义;从歧义字段的结构可分为交集型歧义字段和多义型歧义字段。交集型歧义字段的切分可采用基于统计的方法和基于规则(词性)方法。对多义型歧义字段的处理分别从句法歧义、语义歧义和语用歧义三个方面进行。 中文信息处理中,处理的最多的就是名词。特别是对专有名词的处理是中文自动分词中的又一个难点。本文分析了中文姓名中姓和名的各自特点,给出了中文姓名的自动识别技术。对地名的识别则利用知识库和规则库,采用推理机制技术进行分析;对机构名称的识别技术以高校名称为例,从其语法性质、语义特性和组织规律等特征入手,给出了高校名称识别的基本规则。同时,简要分析了机构名称与人名、地名的关系。
【Abstract】 Chinese automatic word segmentation system is a computer application system ,which make use of computer to conduct the word segmentation and identification for Chinese articles. The system mainly includes automatic word segmentation module, ambiguous word segmentation module and special word identification module, and the quality, value and application level of the system are determined by all these modules which depend on each other.Chinese automatic word segmentation method is made up of mechanical word segmentation method and non-mechanical word segmentation. Maximum positive match method , Maximum negative match method and word by word travel method is the basal mechanical word segmentation, and other eight types, which is not true mechanical word segmentation, are only take some skills on the basal word segmentation method. The specialist system method is a word segmentation method based on the regularity, while the nerve fiber network method is a computer Chinese word segmentation technology based on the fundamental of the artifical nerve network.According to the research and system design about automatic word segmentation method at home and abroad, this paper puts forward the conception of the academic model CWSM:M(F,W,T,K) for the automatic word segmentation system, which includes mechanical word segmentation method, word segmentation dictionary, Chinese text and repository. Further more this paper introduces the evaluation standard of the automatic word segmentation.The ambiguous meaning emerging in the process of word segmentation is mainly made up of special ambiguous meaning caused by computer word segmentation, ambiguous duality meaning caused by natural language and ambiguous meaning caused by the magnitude of .word segmentation library. The ambiguous fields can be classified into three aspects. From the result ofsegmentation, it can be sorted to true ambiguous meaning and false ambiguous meaning. From the acknowledge hiberarchy needed by the segmentation of the ambiguous field, it can be sorted to ambiguity of syntax, the ambiguity of language’s meaning and the ambiguity of language’s application. From the structure of the ambiguous field it can be sorted to intersection field and multi-meanings field. The method of segmentation’ of ambiguous intersection field includes statistic method and part of speech method. The treatment of ambiguous multi-fields can be conducted from three different aspects: ambiguity of syntax, the ambiguity of language’s meaning and the ambiguity of language’s application.In Chinese information system, the use of noun is the most frequent. Especially, it is very difficult to deal with special noun in Chinese automatic word segmentation. First, this paper analyses the character of the surname and firstname in Chinese name, and bring forward the automatic identification technology of Chinese name. Second, this paper takes the repository and rule library to identify the placename by the deduce mechanism. Thirdly, this paper uses college name as a example of organization name identification. According to the characters of grammar, the meaning of language and organization, it brings forward the rule of college name identification. In additional, this paper analyses the relation between the organization name, people name and placename in brief.
- 【网络出版投稿人】 华中师范大学 【网络出版年期】2004年 04期
- 【分类号】TP391.1
- 【被引频次】45
- 【下载频次】1653