节点文献

基于Transformer的中文语音识别研究

Mandarin Automatic Speech Recognition Based on Transformer

【作者】 张淳;

【导师】 张伟彬;

【作者基本信息】 华南理工大学 , 信号与信息处理, 2021, 硕士

【摘要】 语音识别发展迅速,端到端语音识别更因其结构简洁、目标统一等优点,已达到可以和传统语音识别媲美的程度。其中,基于Transformer的端到端语音识别框架因其优秀的建模能力已广泛应用于离线语音识别领域,但目前的研究仍存在一些问题。一方面Transformer模型的优秀性能得益于自注意力模块的全局建模能力,但中文由于其同音异形字、词组等特殊结构,自注意力模块缺乏对其局部建模能力;同时训练过程中自注意力机制可以大规模并发运算从而提升训练效率,但也加重了曝光误差。因此如何提升局部建模能力、改善曝光误差现象尤为重要。另一方面,Transformer解码耗时过长,尤其句子变长时解码时间会明显激增,从而导致性能下降。本文充分考虑中文语音的特点,针对以上的问题,主要研究内容和成果如下:(1)提出基于局部时序依赖的Transformer模型。针对模型编码器缺乏语音特征序列局部建模能力的问题,提出局部密集合成注意力算法,重点关注局部范围帧,同全局自注意力机制结合能有效改善模型建模的能力;针对模型解码器缺乏对目标文字序列的局部建模问题,提出损失自适应局部掩蔽采样方法,降低曝光误差并加强对常见汉语组词方式的局部建模。将上述两种算法结合进Transformer模型,类比基础Transformer模型在中文数据集Aishell1、Aishell2上可获得约13.8%、9.3%的精度提升。(2)提出基于Transformer的语音识别解码速度优化算法,包括模型推理加速、搜索优化两部分。其中模型推理加速主要涉及解码器不同注意力模块,包含自注意力模块加速、编-解码注意力模块加速,类比基础Transformer模型可在模型性能无精度损失的基础上降低相对25%解码耗时。搜索优化则包括集束搜索算法优化、非自回归解码方法两部分。集束搜索算法优化包含动-静态阈值解码加速,类比基础Transformer模型解码流程,有效裁剪置信度较低的解码路径,同模型推理加速结合可相对降低45%解码耗时。最后在模型中引入连接时序分类损失函数,并在其前缀预测结果中融入Transformer解码分数,提出的非自回归解码算法可替代Transformer自回归解码方式,相较于集束搜索优化效果与模型推理加速融合的方法可提升近一倍的模型解码速度,使性能达到更优。

【Abstract】 Speech recognition is developing rapidly at present.Due to simpler structure and unified objective function,end-to-end speech recognition has reached a level which is comparable to traditional speech recognition systems.Among them,the Transformer-based end-to-end speech recognition framework has been widely used in the field of offline speech recognition because of its excellent modeling capabilities,but there are still some problems in current research.The excellent performance of the Transformer-based framework benefits from its global modeling capability of the self-attention module,but the global attention mechanism does not have monotonicity and lacks the modeling capability of local timing dependence of timing signals.Also,its unique batch processing structure,large-scale concurrent operations can increase not only the training rate at training time,but also increases the exposure error.Therefore,how to improve the local modeling ability and improve the robustness is particularly important.In addition,decoding with Transformer-based framework results in large fluctuations in duration and a significant variation in sentence length and delay,which leads to performance degradation.To deal with the above problems,our main research contents and results are as follows:1.A Transformer model based on local timing dependence is proposed.Aiming at the problem that the encoder part lacks the local modeling ability of the speech feature sequence,we propose to use a local dense synthesis algorithm to limit the range of attention to local.When combined with the global attention algorithm,this method effectively improves modeling capabilities.Aiming at the lack of local modeling of the target text sequence in the decoder,a loss-adaptive partial masking sampling algorithm is proposed.It reduces the exposure error and strengthen the local modeling of common Chinese word formation.Incorporating the modules above into the Transformer structure,an accuracy improvement of about 13.8% and 9.3% is obtained on the Chinese data set Aishell1 and Aishell2.2.A Transformer-based speech recognition decoding speed optimization algorithm is proposed,including model inference acceleration and search optimization.The model inference acceleration involves the acceleration of the encoder-decoder attention module and the selfattention module.This can reduce the decoding delay by 25% relatively,without any loss of accuracy.Search optimization includes two parts: beam search algorithm optimization and nonautoregressive decoding acceleration.The beam search algorithm includes both static and dynamic threshold decoding algorithm,which is analogous to the traditional beam search decoding process and effectively cuts the decoding path with lower confidence.Combined with the above acceleration algorithms,the decoding time can be reduced by at least 45%.Secondly,the time-series connection time classification loss function is introduced in the Transformer framework,and integrates the Transformer decoding score into its prefix prediction results.The proposed non-autoregressive decoding algorithm can replace the Transformer autoregressive decoding method.Compared with model inference acceleration and beam search algorithm optimization fusion method,the algorithm can further increase the model decoding speed by nearly doubled,making the performance better.

节点文献中: