节点文献

融合音字特征转换的非自回归Transformer中文语音识别

Non-autoregressive Transformer Chinese Speech Recognition Incorporating Pronunciation-Character Representation Conversion

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 滕思航王烈李雅

【Author】 TENG Sihang;WANG Lie;LI Ya;School of Computer,Electronics and Information,Guangxi University;

【通讯作者】 王烈;

【机构】 广西大学计算机与电子信息学院

【摘要】 基于自注意力机制的Transformer模型在语音识别任务中展现出了强大的模型性能,其中非自回归Transformer自动语音识别模型与自回归模型相比解码速度更快,然而语音识别速度的提升却造成了准确度的大幅降低。为提升非自回归Transformer语音识别模型的识别准确度,首先引入基于连续时间分类(Connectionist Temporal Classification, CTC)的帧信息合并,在帧宽范围内对语音高维表示向量进行融合,改善非自回归Transformer decoder输入序列的特征信息不完整问题;其次对模型输出进行音字特征转换,在decoder的输出读音特征中融合上下文信息,然后转换为包含更多字符特征的输出,从而改善模型同音不同字的识别错误问题。在中文语音数据集AISHELL-1上的实验结果显示,所提模型实现了实时性因子(Real Time Factor, RTF)0.002 8的识别速度与字符错误率(Character Error Rate, CER)8.3%的识别精度,在众多主流中文语音识别算法中展现出较强的竞争力。

【Abstract】 The Transformer based on self-attention mechanism shows powerful model performance in speech recognition tasks, where the non-autoregressive Transformer automatic speech recognition model has a faster decoding speed compared with the autoregressive model.However, the increase in speech recognition speed causes a larger decrease in accuracy.To improve the accuracy of the non-autoregressive Transformer speech recognition model, the frame information merging based on connectionist temporal classification(CTC) is introduced firstly, which fuses the speech high-dimensional representation in the frame width range to improve the problem of incomplete feature information in the non-autoregressive Transformer decoder input sequences.Secon-dly, pronunciation-character representation conversion is performed on the model output, and the pronunciation representation is converted into an output containing more character features by fusing contextual information on the pronunciation features of the decoder output, thus improving the recognition error problem of the model with different characters in the same pronunciation.Experiments on the Chinese speech dataset AISHELL-1 show that the proposed model achieves a recognition speed of real time factor(RTF) 0.0028 and recognition accuracy of 8.3% character error rate(CER),demonstrating strong competitiveness among many mainstream Chinese speech recognition algorithms.

【基金】 广西科技重大专项(桂科AA21077007-1)~~
  • 【文献出处】 计算机科学 ,Computer Science , 编辑部邮箱 ,2023年08期
  • 【分类号】TN912.34
  • 【下载频次】33
节点文献中: 

本文链接的文献网络图示:

本文的引文网络