节点文献
手语到情感语音的转换
Converting sign language to emotional speech
【摘要】 为了解决语言障碍者与健康人之间的交流障碍问题,提出了一种基于神经网络的手语到情感语音转换方法。首先,建立了手势语料库、人脸表情语料库和情感语音语料库;然后利用深度卷积神经网络实现手势识别和人脸表情识别,并以普通话声韵母为合成单元,训练基于说话人自适应的深度神经网络情感语音声学模型和基于说话人自适应的混合长短时记忆网络情感语音声学模型;最后将手势语义的上下文相关标注和人脸表情对应的情感标签输入情感语音合成模型,合成出对应的情感语音。实验结果表明,该方法手势识别率和人脸表情识别率分别达到了95.86%和92.42%,合成的情感语音EMOS得分为4.15,合成的情感语音具有较高的情感表达程度,可用于语言障碍者与健康人之间正常交流。
【Abstract】 In order to solve the problem of communication between speech-impaired people and healthy people, a neural network-based sign language-to-emotional speech conversion method is proposed. Firstly, a gesture corpus, a facial expression corpus, and an emotional speech corpus are established. Then, a deep convolution neural network is used to realize the recognition of gestures and facial expression. Mandarin vowels and consonants are used as synthesis units to train the deep neural network emotional speech acoustic model based on speaker adaptation and the mixed long short-term memory network emotional speech acoustic model based on speaker adaptation. Finally, the context-dependent labels of gesture semantics and the emotion labels corresponding to facial expression are input into the emotional speech synthesis model to synthesize the corresponding emotional speech. The experimental results show that gesture recognition accuracy and the facial expression recognition accuracy are 95.86% and 92.42%, respectively, and the average mean score of the synthesized emotional speech is 4.15. Meanwhile, the synthesized emotional speech has a high degree of emotional expression, which can be used for communication between speech-impaired people and healthy people.
【Key words】 gesture recognition; facial expression recognition; emotional speech synthesis; neural network; sign language to speech conversion; speech-impaired people;
- 【文献出处】 计算机工程与科学 ,Computer Engineering & Science , 编辑部邮箱 ,2022年10期
- 【分类号】TP391.41;TN912.3
- 【下载频次】102