节点文献
基于单语语料和词向量对齐的蒙汉神经机器翻译研究
Mongolian-Chinese Neural Machine Translation Based on Monolingual Corpora and Word Embedding Alignment
【摘要】 近年来,随着人工智能和深度学习的发展,神经机器翻译在某些高资源语言对上取得了接近人类水平的效果。然而对于低资源语言对如汉语和蒙古语,神经机器翻译的效果并不尽如人意。为了提高蒙汉神经机器翻译的性能,该文基于编码器—解码器神经机器翻译架构,提出一种改善蒙汉神经机器翻译结果的方法。首先将蒙古语和汉语的词向量空间进行对齐并用它来初始化模型的词嵌入层,然后应用联合训练的方式同时训练蒙古语到汉语的翻译和汉语到蒙古语的翻译。并且在翻译的过程中,最后使用蒙古语和汉语的单语语料对模型进行去噪自编码的训练,增强编码器的编码能力和解码器的解码能力。实验结果表明该文所提出方法的效果明显高于基线模型,证明该方法可以提高蒙汉神经机器翻译的性能。
【Abstract】 To improve the Mongolian-Chinese neural machine translation performance, this paper proposes a method based on monolingual corpora and word embedding alignment.First, the Mongolian and Chinese word embedding spaces are aligned to initialize the embedding layers of the model.Second, jointly training is employed to train Mongolian-to-Chinese translation and Chinese-to-Mongolian translation at the same time.Finally, Mongolian and Chinese monolingual corpora are utilized to train the model as a denoising autoencoder.Experimental results show that the proposed method outperforms the baseline approach and improves the performance of Mongolian-Chinese neural machine translation.
【Key words】 Mongolian-Chinese neural machine translation; monolingual corpora; word embedding alignment;
- 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2020年02期
- 【分类号】TP391.1;H085
- 【被引频次】13
- 【下载频次】164