节点文献

融合无监督特征的藏文分词方法研究

Study on Fusion of Unsupervised Features for Tibetan Word Segmentation

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 李亚超加羊吉江静何向真于洪志

【Author】 LI Yachao;JIA Yangji;JIANG Jing;HE Xiangzhen;YU Hongzhi;Key Lab of Chinese National Linguistic Information Technology,Northwest University for Nationalities;

【机构】 西北民族大学中国民族语言文字信息技术重点实验室

【摘要】 藏文分词是藏文信息处理的基础性关键问题,目前基于序列标注的藏文分词方法大都采用音节位置特征和类别特征等。该文从无标注语料中抽取边界熵特征、邻接变化数特征、无监督间隔标注等无监督特征,并将之融合到基于序列标注的分词系统中。从实验结果可以看出,与基线藏文分词系统相比,分词F值提高了0.97%,并且未登录词识别结果也有较大的提高。说明,该文从无标注数据中提取出的无监督特征较为有效,和有监督的分词模型融合到一起显著提高了基线分词系统的效果。

【Abstract】 Tibetan word segmentation(TWS)is an important problem in Tibetan information processing,while the current TWS features are mostly adopt the syllable position and syllable categories.The paper extracted unsupervised features,for example,boundary entropy,accessorvariety and unsupervised gap tagging,from unlabeled corpus,and studied the TWS merged with unsupervised features.The experimental results show that,F score increase of 0.97% compare to the baselinesystem,the method get a good performance on out of vocabulary words.From the above,we can conclude that this method can effectively distracted from unlabeled corpus,which can be combined easily with the supervised segmentation model. The method can significantly increases the effect of the baseline TWS.

【关键词】 藏文分词序列标注
【Key words】 Tibetanword segmentationsequence labeling
【基金】 国家社科基金青年项目(15CYY043);国家自然科学基金(61262054);甘肃省高等学校科研项目(2016B—007);甘肃省民族语言智能处理重点实验室开放基金;西北民族大学中央高校基本科研业务费专项资金(31920140064,31920150089)
  • 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2017年02期
  • 【分类号】TP391.1
  • 【被引频次】16
  • 【下载频次】210
节点文献中: 

本文链接的文献网络图示:

本文的引文网络