节点文献

基于BLSTM-RNN的语音驱动逼真面部动画合成

Speech-driven video-realistic talking head synthesis using BLSTM-RNN

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 阳珊樊博谢磊王丽娟宋謌平

【Author】 YANG Shan;FAN Bo;XIE Lei;WANG Lijuan;SONG Geping;School of Computer Science, Northwestern Polytechnical University;Microsoft Research Asia;

【机构】 西北工业大学计算机学院陕西省语音与图像处理重点实验室微软亚洲研究院

【摘要】 本文提出了一种基于深度BLSTM(bidirectional long short-term memory)的语音驱动面部动画合成方法。BLSTM是一种特殊的递归神经网络(recurrent neural network,RNN),能够有效地对语音的长时上下文进行建模。本文利用说话人的音视频双模态信息训练BLSTM-RNN神经网络,采用主动外观模型(active appearance model,AAM)对人脸图像进行建模,将AAM模型参数作为网络输出。本文研究了网络结构、不同语音特征输入对动画合成效果的影响。基于LIPS2008标准评测库的实验表明,具有BLSTM层的网络效果明显优于前向网络,基于BLSTM-前向-BLSTM 256节点(BFB256)的三层模型结构的效果最佳,FBANK和基频、能量组合可以进一步提升动画合成效果。

【Abstract】 We propose a deep Bidirectional Long Short Term Memory(BLSTM) approach to speech-driven photo-realistic talking head animation. Long short-term memory(LSTM) is a specific recurrent neural network(RNN) architecture that is designed to model temporal sequences and their long-range dependencies more accurately than conventional RNNs. In this paper, we train deep BLSTM using a speaker’s audio-visual bimodal data. Active appearance model(AAM) is used to model the facial movements and AAM parameters are used as the prediction targets of the neural network. Experiments on LIPS2008 audio-visual corpus show that networks with BLSTM layer(s) consistently outperform those only have feed-forward layers. We interestingly find that the best network is a feed-forward layer inserting into two BLSTM layers(BFB) on our dataset. The combination of FBANK, pitch and energy is the best performed feature set for the speech-driven taking head animation task.

【基金】 国家自然科学基金项目(61175018)
  • 【会议录名称】 第十三届全国人机语音通讯学术会议(NCMMSC2015)论文集
  • 【会议名称】第十三届全国人机语音通讯学术会议(NCMMSC2015)
  • 【会议时间】2015-10-25
  • 【会议地点】中国天津
  • 【分类号】TP391.41
  • 【主办单位】中国中文信息学会语音信息专业委员会
节点文献中: 

本文链接的文献网络图示:

本文的引文网络