节点文献

基于感知器的中文分词增量训练方法研究

An Incremental Learning Scheme for Perceptron Based Chinese Word Segmentation

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 韩冰刘一佳车万翔刘挺

【Author】 HAN Bing;LIU Yijia;CHE Wanxiang;LIU Ting;Research Center for Social Computing and Information Retrieval,Harbin Institute of Technology;

【机构】 哈尔滨工业大学计算机学院社会计算与信息检索研究中心

【摘要】 该文提出了一种基于感知器的中文分词增量训练方法。该方法可在训练好的模型基础上添加目标领域标注数据继续训练,解决了大规模切分数据难于共享,源领域与目标领域数据混合需要重新训练等问题。实验表明,增量训练可以有效提升领域适应性,达到与传统数据混合相类似的效果。同时该文方法模型占用空间小,训练时间短,可以快速训练获得目标领域的模型。

【Abstract】 In this paper,we propose an incremental learning scheme for perceptron based Chinese word segmentation.Our method can perform continuous training over a fine tuned source domain model,enabling to deliver model without annotated data and re-training.Experimental results shows the scheme proposed can significantly improve adaptation performance on Chinese word segmentation and achieve comparable performance with traditional method.At the same time,our method can significantly reduce the model size and the training time.

  • 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2015年05期
  • 【分类号】TP391.1
  • 【被引频次】11
  • 【下载频次】275
节点文献中: 

本文链接的文献网络图示:

本文的引文网络