节点文献

基于规则和统计的日语分词和词性标注的研究

Study on Japanese Word Segmentation and POS Tagging Based on Rules and Statistics

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 姜尚仆陈群秀

【Author】 JIANG Shangpu,CHEN Qunxiu Computer Science and Artificial Intelligence Division,National Laboratory for information Science and Technology, Tsinghua University Department of Computer Science and Technology,Tsinghua University,Beijing,100084

【机构】 清华大学信息科学与技术国家实验室计算机科学与人工智能研究部清华大学计算机系

【摘要】 和中文类似,日语的词法分析需要首先进行分词。基于词的方法是日语分词的主流方法。同时,对中文的研究结果表明,词性标注对分词结果的正确性有帮助,这点在日语中也得到了证实。我们提出了一种基于规则和统计的日语分词和词性标注方法,使用基于单一感知器的联合分词和词性标注算法进行训练和解码,并加入了词语的邻接属性特征。实验结果表明,这种方法无论是分词准确率还是分词加词性标注的准确率都比原有的基于字和词的混合HMM算法更高。我们已将这种方法应用到我们的日汉机器翻译系统中。

【Abstract】 Like that of Chinese,Japanese morphological analysis starts with word segmentation.Word-based approach is the mainstream on Japanese word segmentation.Meanwhile,according to the study on Chinese,POS tagging results are helpful to the correctness of word segmentation.This conclusion is also substantiated on Japanese.We propose a Japanese word segmentation and POS tagging approach based on rules and statistics,which uses a single perceptron based joint word segmentation and POS tagging algorithm for training and decoding,and is added with the features of adjacency attribute. The experiment shows that the new approach is better performed than the hybrid character and word based HMM algorithm.We have already applied this approach into our Japanese-Chinese machine translation system.

【基金】 国家863计划重点项目(项目号:2006AA010109)资助
  • 【会议录名称】 中国计算机语言学研究前沿进展(2007-2009)
  • 【会议名称】第十届全国计算语言学学术会议
  • 【会议时间】2009-07-24
  • 【会议地点】中国山东烟台
  • 【分类号】TP391.2
  • 【主办单位】中国中文信息学会
节点文献中: