节点文献

基于统计的汉语词性标注方法的分析与改进

Analysis and Improvement of Statistics Based Chinese Part of Speech Tagging

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 魏欧吴健孙玉芳sonata.iscas.ac.cn

【Author】 WEI Ou WU Jian SUN Yu fang(Institute of Software The Chinese Academy of Sciences Beijing 100080)

【机构】 中国科学院软件研究所!北京100080E-mail:yfsun

【摘要】 从词性概率矩阵与词汇概率矩阵的结构和数值变化等方面 ,对目前常用的基于统计的汉语词性标注方法中训练语料规模与标注正确率之间所存在的非线性关系作了分析 .为了充分利用训练语料库 ,提高标注正确率 ,从利用词语相关的语法属性和加强对未知词的处理两个方面加以改进 ,提高了标注性能 .封闭测试和开放测试的正确率分别达到 96.5%和 96% .

【Abstract】 In this paper, a popular statistics\|based training and tagging method for Chinese texts is studied, and the nonlinear relation between training set and tagging accuracy is analyzed from the aspects of the structure and numerical value of the matrix of transition probabilities and the matrix of symbol probabilities. In order to make use of training corpus sufficiently and get the higher tagging accuracy, the training and tagging method is improved from two aspects: using other grammatical attributes of words, and strengthening the processing of unknown words. With the improved method, open test and close test showed that the overall accuracies are about 96.5% and 96% respectively.

【关键词】 词性标注n元语法语料语法属性
【Key words】 Part-of-Speech taggingn-gramcorpusgrammatical attribute.
【基金】 国家“九五”重点科技攻关项目基金!(Nos.96-B08-1-3,98-779-0 1-02)资助
  • 【文献出处】 软件学报 ,JOURNAL OF SOFTWARE , 编辑部邮箱 ,2000年04期
  • 【分类号】TP391.1
  • 【被引频次】97
  • 【下载频次】713
节点文献中: 

本文链接的文献网络图示:

本文的引文网络