节点文献

基于说话人自适应训练的汉藏双语语音合成

Realizing Mandarin-Tibetan Bilingual Speech Synthesis by Speaker Adaptive Training

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 王海燕杨鸿武甘振业裴东

【Author】 WANG Haiyan,YANG Hongwu,GAN Zhenye,PEI Dong(College of Physics and Electronic Engineering,Northwest Normal University,Lanzhou 730070,China)

【机构】 西北师范大学物理与电子工程学院

【摘要】 根据藏语和汉语在发音上的相似性,提出了一种基于隐马尔科夫模型的汉藏双语语音合成方法。以声韵母为合成基元,采用多个普通话说话人和1个藏语说话人的语料库,利用说话人自适应训练,获得一个汉藏双语混合语言的平均音模型。通过说话人自适应变换,从混合语言的平均音模型获得普通话或藏语的说话人相关模型,从而合成出普通话或藏语语音。实验结果表明,在藏语训练语句较少的情况下,本文方法合成的藏语语音明显优于仅采用说话人相关模型合成的藏语语音。

【Abstract】 This paper proposes a method for realizing HMM-based Mandarin-Tibetan bilingual speech synthesis according to the similarity of pronunciation between Mandarin and Tibetan.We select the Initial and the Final as synthesis units and train a set of average mixed-lingual models from a large Mandarin multi-speaker-based corpus and a small Tibetan one-speaker-based corpus by using the speaker adaptive training.Then we apply the speaker adaptation transformation to the speaker dependent training data to obtain a set of speaker dependent Mandarin models or Tibetan models from the average mixed-lingual models.The Mandarin speech or Tibetan speech is synthesized from the speaker dependent Mandarin models or Tibetan models respectively.Experimental results show that the proposed method outperforms the method only using Tibetan SD models in the case of the small amount of training Tibetan utterances.

【基金】 国家自然科学基金项目(61263036,612620555);甘肃省杰出青年基金项目(1210RJDA007)
  • 【会议录名称】 第十二届全国人机语音通讯学术会议(NCMMSC2013)论文集
  • 【会议名称】第十二届全国人机语音通讯学术会议(NCMMSC’2013)
  • 【会议时间】2013-08-05
  • 【会议地点】中国贵州贵阳
  • 【分类号】TN912.33
  • 【主办单位】中国中文信息学会语音信息专业委员会、中国声学学会语言、听觉和音乐声学分会、中国语言学会语音学分会
节点文献中: