节点文献
一种规则与统计相结合的汉语分词方法
A Method Combining Rule-based and Statistics-based Approaches for Chinese Word Segmentation
【摘要】 汉语自动分词是中文信息处理领域的一项基础性课题,对现有的汉语分词方法作了简单的概述和分析,然后提出了一种新的分词方法,该方法基于一个标注好了的语料库,并且结合了规则和语料库统计两种分词方法。
【Abstract】 Chinese automatic word segmentation is a basic task in the area of Chinese NLP.After summarizing and analyzing current techniques used in Chinese word segmentation,this paper presents a new method for word segmentation which is based on a marked corpus base.The method combines rule-based and corpus-based statistical methods.
【关键词】 中文信息处理;
分词;
语料库;
交集型歧义;
【Key words】 Chinese NLP; Word Segmentation; Corpus; Crossing Ambiguities;
【Key words】 Chinese NLP; Word Segmentation; Corpus; Crossing Ambiguities;
【基金】 国家"863"基金资助项目(2001AA114102)
- 【文献出处】 计算机应用研究 ,Application Research of Computers , 编辑部邮箱 ,2004年03期
- 【分类号】TP391.1
- 【被引频次】140
- 【下载频次】568