节点文献

基于循环神经网络的藏语语音识别声学模型

The Acoustic Model for Tibetan Speech Recognition Based on Recurrent Neural Network

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 黄晓辉李京

【Author】 HUANG Xiaohui;LI Jing;College of Computer Science and Technology,University of Science and Technology of China;PLA University of Foreign Language;

【机构】 中国科学技术大学计算机科学与技术学院解放军外国语学院

【摘要】 探索将循环神经网络和连接时序分类算法应用于藏语语音识别声学建模,实现端到端的模型训练。同时根据声学模型输入与输出的关系,通过在隐含层输出序列上引入时域卷积操作来对网络隐含层时域展开步数进行约简,从而有效提升模型的训练与解码效率。实验结果显示,与传统基于隐马尔可夫模型的声学建模方法相比,循环神经网络模型在藏语拉萨话音素识别任务上具有更好的识别性能,而引入时域卷积操作的循环神经网络声学模型在保持同等识别性能的情况下,拥有更高的训练和解码效率。

【Abstract】 The recurrent neural network and the connectionist temporal classification algorithm are applied to the acoustic modeling of Tibetan speech recognition,so as to achieve end-to-end model training.According to the relationship between the input and output of the acoustic model,the time domain convolution operation on the output sequence of the hidden layer is introduced to reduce the time domain expansion of the network’s hidden layers.Experimental results show that the recurrent neural network model achieves better recognition performance in Tibetan Lhasa phoneme recognition compared with the traditional acoustic models based on Hidden Markov Model,while the acoustic model based on recurrent neural network with time-domain convolution possesses higher training and decoding efficiency while maintaining the same recognition performance.

【基金】 国家重点研发计划项目(2016YFB0201402)
  • 【文献出处】 中文信息学报 ,Journal of Chinese Information Processing , 编辑部邮箱 ,2018年05期
  • 【分类号】TN912.34;TP183
  • 【被引频次】35
  • 【下载频次】431
节点文献中: