节点文献
结合发音特征的抗噪语音识别
An English Speech Recognition System Using Articulatory Features and Hearing Mechanism
【Author】 Zhuanling Zha;Jin Hu;Yahui Shan;Xiang Xie;Jing Wang;School of Information and Electronics,Beijing Institute of Technology(BIT);
【机构】 北京理工大学,信息与电子学院;
【摘要】 本文引入发音与听觉感知的非线性关系到语音识别中,通过实验验证了改进系统具有更好的鲁棒性。实验在MOCHATIMIT数据库上,训练得到将声学特征映射成发音参数的ELM神经网络,通过该网络得到TIMIT数据集的发音特征,与声学参数(MFCC)相融合,分别在DNN-HMM与GMM-HMM系统上训练,相比只用MFCC训练的语音识别系统有更好的抗噪性能。干净语音下,特征融合的GMM-HMM系统比基线系统词错误率(WER)下降了2.2个百分点,在10dB街道噪声信噪比时,GMM-HMM系统的WER也下降了3.4个百分点,DNN-HMM的基于MFCC与MFCC+EMA的识别性能也有5个百分点的绝对提升。
【Abstract】 In this paper,the nonlinear relationship between pronunciation and auditory perception is introduced into speech recognition and better robustness is shown in the experimental results.The ELM neural network map ping the relation is trainedthrough MOCHA TIMIT database.EMA which are obtained by the network and acoustic parameters(MFCC) are fusion for training acoustic model DNN-HMMandGMM-HMMin this experiment.Theproposed system has a better recognition performance andanti-noise performance compared to the baseline system which is trained using only MFCC.Theproposed system based GMM-HMM has a decrease of 2.2 percentage points on the word error rate(WER) than the baseline system in the clean speech,and has adecrease of 3.4 percentage points in the lOdB signal noise ratio AURORA street noise.While the system DNN-HMM based on MFCC also has a decrease nearly 5 percentage points than that based MFCC+EMA.
【Key words】 speech recognition; articulatory features; ELM; robustness;
- 【会议录名称】 第十四届全国人机语音通讯学术会议(NCMMSC’2017)论文集
- 【会议名称】第十四届全国人机语音通讯学术会议
- 【会议时间】2017-10-11
- 【会议地点】中国江苏连云港
- 【分类号】TN912.34
- 【主办单位】中国中文信息学会语音信息专业委员会