节点文献

改进的跨语种语音合成模型自适应方法

An Improved Cross-Language Model Adaptation Method for Speech Synthesis

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 刘航凌震华郭武戴礼荣

【Author】 LIU Hang, LING Zhen-Hua, GUO Wu, DAI Li-Rong (iFLYTEK Speech Laboratory, Department of Electronic Engineering and Information Science,University of Science and Technology of China, Hefei 230027)

【机构】 中国科学技术大学电子工程与信息科学系讯飞语音实验室

【摘要】 统计参数语音合成中的跨语种模型自适应主要应用于目标说话人语种与源模型语种不同时,使用目标发音人少量语音数据快速构建具有其音色特征的源模型语种合成系统.本文对传统的基于音素映射和三音素模型的跨语种自适应方法进行改进,一方面通过结合数据挑选的音素映射方法以提高音素映射的可靠性,另一方面引入跨语种的韵律信息映射以弥补原有方法中三音素模型在韵律表征上的不足.在中英文跨语种模型自适应系统上的实验结果表明,改进后系统合成语音的自然度与相似度相对传统方法都有了明显提升.

【Abstract】 Cross-language model adaptation in statistical parametric speech synthesis is used for rapidly constructing a text-to-speech (TTS) system with the target speaker’s characteristics when the source and the target speakers’ languages are different. In this paper, the conventional cross-language adaptation method based on phone-mapping and triphone models is improved by two means. Firstly, phone mapping combined with data-selection is adopted to improve its reliability. Secondly, cross-language prosodic information mapping is introduced to make use of prosodic information, which is ignored in the triphone model. Experiments on Chinese-to-English adaptation show that the synthesized speech using the improved method has much better naturalness and speaker similarity compared with the result of conventional method.

【基金】 中央高校基本科研业务费专项资金资助项目
  • 【文献出处】 模式识别与人工智能 ,Pattern Recognition and Artificial Intelligence , 编辑部邮箱 ,2011年04期
  • 【分类号】TN912.33
  • 【被引频次】1
  • 【下载频次】120
节点文献中: 

本文链接的文献网络图示:

本文的引文网络