节点文献

基于单语语料和词向量对齐的蒙汉神经机器翻译研究

Mongolian-Chinese Neural Machine Translation Based on Monolingual Corpora and Word Embedding Alignment

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 曹宜超高翊李淼冯韬王儒敬付莎

【Author】 CAO Yichao;GAO Yi;LI Miao;FENG Tao;WANG Rujing;FU Sha;Institute of Intelligent Machines,Chinese Academy of Sciences;University of Science and Technology of China;Yunnan Minority Languages Guidance Committee Office;

【通讯作者】 李淼;

【机构】 中国科学院合肥智能机械研究所中国科学技术大学云南省少数民族语文指导工作委员会办公室

【摘要】 近年来,随着人工智能和深度学习的发展,神经机器翻译在某些高资源语言对上取得了接近人类水平的效果。然而对于低资源语言对如汉语和蒙古语,神经机器翻译的效果并不尽如人意。为了提高蒙汉神经机器翻译的性能,该文基于编码器—解码器神经机器翻译架构,提出一种改善蒙汉神经机器翻译结果的方法。首先将蒙古语和汉语的词向量空间进行对齐并用它来初始化模型的词嵌入层,然后应用联合训练的方式同时训练蒙古语到汉语的翻译和汉语到蒙古语的翻译。并且在翻译的过程中,最后使用蒙古语和汉语的单语语料对模型进行去噪自编码的训练,增强编码器的编码能力和解码器的解码能力。实验结果表明该文所提出方法的效果明显高于基线模型,证明该方法可以提高蒙汉神经机器翻译的性能。

【Abstract】 To improve the Mongolian-Chinese neural machine translation performance, this paper proposes a method based on monolingual corpora and word embedding alignment.First, the Mongolian and Chinese word embedding spaces are aligned to initialize the embedding layers of the model.Second, jointly training is employed to train Mongolian-to-Chinese translation and Chinese-to-Mongolian translation at the same time.Finally, Mongolian and Chinese monolingual corpora are utilized to train the model as a denoising autoencoder.Experimental results show that the proposed method outperforms the baseline approach and improves the performance of Mongolian-Chinese neural machine translation.

【基金】 国家自然科学基金(61572462);中国科学院“十三五”信息化专项科学大数据工程(XXH13505-03-203)
  • 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2020年02期
  • 【分类号】TP391.1;H085
  • 【被引频次】13
  • 【下载频次】164
节点文献中: 

本文链接的文献网络图示:

本文的引文网络