节点文献

基于CNN的端到端中文语音识别算法设计与FPGA验证

The Design and FPGA Verification of End-to-end Mandarin Speech Recognition Based on CNN

【作者】 王亮;

【导师】 陆生礼; 禹胜林;

【作者基本信息】 东南大学 , 集成电路工程(专业学位), 2021, 硕士

【摘要】 语音识别作为人机交互的第一接口,广泛应用于智能音箱、智能家居、汽车电子等领域。卷积神经网络凭借其强大的非线性表达和特征提取能力,被广大研究者应用到语音识别算法声学模型的研究。然而相对于传统语音识别算法,基于卷积神经网络的语音识别算法拥有更多的参数量和计算量,对硬件条件要求更高,使得其部署在移动终端存在巨大的困难。因此,基于软硬件协同设计,实现高效快速的语音识别算法具有重要的现实意义。本文基于卷积神经网络设计了一种端到端语音识别算法,通过引入卷积注意力模块,提高了卷积神经网络对声学模型的建模能力。通过使用复杂度更低的语谱图,在节省了特征提取计算时间的同时,保留了输入语音的大部分信息。通过优化网络结构和使用连接时序分类,在不增加模型参数量的前提下,提升了模型性能。使用数据增强,扩大了小数据集的数据多样性,大幅提升了模型的识别准确率。在卷积神经网络加速器方面,设计了卷积计算模块、模式控制器、数据缓存模块、中间缓存区模块和结果处理模块,并完成各模块的功能仿真。最后,搭建FPGA验证系统,进行算法移植,验证了语音识别算法的有效性。本文基于卷积神经网络设计的语音识别算法,在thchs-30数据集上达到了82.4%的准确率。为对算法进行验证,本文基于FPGA平台搭建了验证系统。实验结果表明,在时钟频率为100MHz,卷积神经网络加速器有效计算能力达到53.2GOPS,性能功耗比为9.9GOPS/W。从语音输入结束到识别完成,延迟时间约274ms。本文的研究对未来高准确率低延迟的语音识别系统的实现具有一定的参考意义。

【Abstract】 As the first interface of human-computer interaction,voice recognition is widely used in fields such as smart speakers,smart homes,and automotive electronics.With its powerful nonlinear expression and feature extraction capabilities,convolutional neural networks have been applied to the study of acoustic models of speech recognition algorithms by a large number of researchers.However,compared with traditional speech recognition algorithms,speech recognition algorithms based on convolutional neural networks have more parameters and calculations,and require higher hardware conditions,making it difficult to deploy in mobile terminals.Therefore,based on the software and hardware co-design,the realization of efficient and fast speech recognition algorithms has important practical significance.In this thesis,an end-to-end speech recognition algorithm is designed based on convolutional neural networks.By introducing the convolutional attention module,the convolutional neural network’s ability to model acoustic models is improved.By using a less complex spectrogram,while saving the calculation time for feature extraction,most of the information of the input speech is preserved.By optimizing the network structure and using connection timing classification,the model performance is improved without increasing the amount of model parameters.The use of data enhancement expands the data diversity of small data sets and greatly improves the recognition accuracy of the model.In terms of the convolutional neural network accelerator,the convolution calculation module,the mode controller,the data buffer module,the intermediate buffer module and the result processing module are designed,and the functional simulation of each module is completed.Finally,an FPGA verification system was built and the algorithm was transplanted to verify the effectiveness of the speech recognition algorithm.The speech recognition algorithm designed in this thesis based on convolutional neural network has achieved 82.4% accuracy on the thchs-30 data set.In order to verify the algorithm,this thesis builds a verification system based on the FPGA platform.The experimental results show that at a clock frequency of 100 MHz,the effective computing power of the convolutional neural network accelerator reaches 53.2 GOPS,and the performance-to-power ratio is 9.9 GOPS/W.From the end of voice input to the completion of recognition,the delay time is about 274 ms.The research in this paper has certain reference significance for the realization of high accuracy and low delay speech recognition system in the future.

  • 【网络出版投稿人】 东南大学
  • 【网络出版年期】2022年 06期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络