节点文献

语音识别中神经网络声学模型的说话人自适应研究

SPEAKER ADAPTATION RESEARCH OF NEURAL NETWORK ACOUSTIC MODEL IN SPEECH RECOGNITION

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 金超龚铖李辉

【Author】 Jin Chao;Gong Cheng;Li Hui;School of Electronic Science and Technology,University of Science and Technology of China;

【机构】 中国科学技术大学信息科学技术学院

【摘要】 针对语音识别系统中测试的目标说话人语音和训练数据的说话人语音存在较大差异时,系统识别准确率下降的问题,提出一种基于深度神经网络DNN(Deep Neural Network)的说话人自适应SA(Speaker Adaptation)方法。它是在特征空间上进行的说话人自适应,通过在DNN声学模型中加入说话人身份向量I-Vector辅助信息来去除特征中的说话人差异信息,减少说话人差异的影响,保留语义信息。在TEDLIUM开源数据集上的实验结果表明,该方法在特征分别为fbank和f MLLR时,系统单词错误率WER(Word Error Rate)相对基线DNN声学模型提高了7.7%和6.7%。

【Abstract】 Aiming at the problem that the accuracy of speech recognition system is degraded when there is a large difference between the speaker’s speech of the target speaker and the training data tested in the speech recognition system,this paper proposed speaker adaptation method which was based on deep neural network. It performed featuresspace speaker adaptation. By adding the I-Vector auxiliary information of the speaker’s identity vector to the DNN acoustic model,the speaker’s difference information in the feature was removed. The influence of the speaker’s difference was reduced,and the semantic information was retained. The experimental results on the open source dataset TEDLIUM showed that the Word Error Rate( WER) was 7. 7% and 6. 7% higher than the baseline DNN acoustic model when the characteristics were fbank and fMLLR in the SA-DNN acoustic model.

  • 【文献出处】 计算机应用与软件 ,Computer Applications and Software , 编辑部邮箱 ,2018年02期
  • 【分类号】TN912.34;TP183
  • 【被引频次】21
  • 【下载频次】366
节点文献中: 

本文链接的文献网络图示:

本文的引文网络