节点文献

基于深度学习的藏文分词方法

Tibetan word segmentation based on deep learning

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 李博涵刘汇丹龙从军吴健

【Author】 LI Bo-han;LIU Hui-dan;LONG Cong-jun;WU Jian;Institute of Software,Chinese Academy of Sciences;School of Computer and Control Engineering,University of Chinese Academy of Sciences;Institute of Ethnology and Anthropology,Chinese Academy of Social Sciences;

【机构】 中国科学院软件研究所中国科学院大学计算机与控制学院中国社会科学院民族学与人类学研究所

【摘要】 重点研究将深度学习技术应用于藏文分词任务,采用多种深度神经网络模型,包括循环神经网络(RNN)、双向循环神经网络(Bi RNN)、层叠循环神经网络(Stacked RNN)、长短期记忆模型(LSTM)和编码器-标注器长短期记忆模型(Encoder-Labeler LSTM)。多种模型在以法律文本、政府公文、新闻为主的分词语料中进行实验,实验数据表明,编码器-标注器长短期记忆模型得到的分词结果最好,分词准确率可以达到92.96%,召回率为93.30%,F值为93.13%。

【Abstract】 The application of deep learning on Tibetan word segmentation was studied.Several models of deep neural network were implemented,including recurrent neural network,bi-directional recurrent neural network,stacked recurrent neural network,long short-term memory network and encoder-labeler long short-term memory network.These models were performed on written style corpus,including legal text,government documents and news.Experimental results show that the encoder-labeler long shortterm memory network achieves the best results,the precision,recall and F value reach 92.96%,93.30% and 93.13% respectively.

【基金】 国家自然科学基金项目(61303165、61540057、61132009);青海省自然科学基金项目(2016-ZJ-Y04、2016-ZJ-740);国家语委重点基金项目(ZDI135-17)
  • 【文献出处】 计算机工程与设计 ,Computer Engineering and Design , 编辑部邮箱 ,2018年01期
  • 【分类号】TP18;TP391.1
  • 【被引频次】29
  • 【下载频次】488
节点文献中: 

本文链接的文献网络图示:

本文的引文网络