节点文献
一种RNN-T与BERT相结合的端到端语音识别模型
An end-to-end speech recognition model combining RNN-T and BERT
【摘要】 端到端语音识别模型由于结构简单且容易训练,已成为目前最流行的语音识别模型。然而端到端语音识别模型通常需要大量的语音-文本对进行训练,才能取得较好的识别性能。而在实际应用中收集大量配对数据既费力又昂贵,因此其无法在实际应用中被广泛使用。本文提出一种将RNN-T (Recurrent Neural Network Transducer,RNN-T)模型与BERT(Bidirectional Encoder Representations from Transformers,BERT)模型进行结合的方法来解决上述问题,其通过用BERT模型替换RNN-T中的预测网络部分,并对整个网络进行微调,从而使RNN-T模型能有效利用BERT模型中的语言学知识,进而提高模型的识别性能。在中文普通话数据集AISHELL-1上的实验结果表明,采用所提出的方法训练后的模型与基线模型相比能获得更好的识别结果。
【Abstract】 The end-toend speech recognition model has become one of the most popular speech recognition models due to its simple structure and easy training. However,it usually needs a large number of speech-text pairs for the training of an end-to-end speech recognition model to achieve a better performance. In practical applications,it is very laborious and expensive to collect a large number of the paired data,resulting in the model cannot be widely used. This paper proposes a method of combining the Recurrent Neural Network Transducer( RNN-T) model with the Bidirectional Encoder Representations from Transformers( BERT)model to solve the above problems. It replaces the prediction network part in the RNN-T with the BERT model and fine-tunes the entire network,thus the RNN-T model effectively uses linguistic information to improve model recognition performance. The experimental results on the Chinese mandarin data set AISHELL-1 show that,compared with the baseline system,the system using the proposed expansion method achieves better recognition results.
- 【文献出处】 智能计算机与应用 ,Intelligent Computer and Applications , 编辑部邮箱 ,2021年02期
- 【分类号】TN912.34
- 【被引频次】1
- 【下载频次】229