节点文献
基于说话人自适应训练的汉藏双语语音合成
Realizing Mandarin-Tibetan Bilingual Speech Synthesis by Speaker Adaptive Training
【Author】 WANG Haiyan,YANG Hongwu,GAN Zhenye,PEI Dong(College of Physics and Electronic Engineering,Northwest Normal University,Lanzhou 730070,China)
【机构】 西北师范大学物理与电子工程学院;
【摘要】 根据藏语和汉语在发音上的相似性,提出了一种基于隐马尔科夫模型的汉藏双语语音合成方法。以声韵母为合成基元,采用多个普通话说话人和1个藏语说话人的语料库,利用说话人自适应训练,获得一个汉藏双语混合语言的平均音模型。通过说话人自适应变换,从混合语言的平均音模型获得普通话或藏语的说话人相关模型,从而合成出普通话或藏语语音。实验结果表明,在藏语训练语句较少的情况下,本文方法合成的藏语语音明显优于仅采用说话人相关模型合成的藏语语音。
【Abstract】 This paper proposes a method for realizing HMM-based Mandarin-Tibetan bilingual speech synthesis according to the similarity of pronunciation between Mandarin and Tibetan.We select the Initial and the Final as synthesis units and train a set of average mixed-lingual models from a large Mandarin multi-speaker-based corpus and a small Tibetan one-speaker-based corpus by using the speaker adaptive training.Then we apply the speaker adaptation transformation to the speaker dependent training data to obtain a set of speaker dependent Mandarin models or Tibetan models from the average mixed-lingual models.The Mandarin speech or Tibetan speech is synthesized from the speaker dependent Mandarin models or Tibetan models respectively.Experimental results show that the proposed method outperforms the method only using Tibetan SD models in the case of the small amount of training Tibetan utterances.
【Key words】 HMM-based speech synthesis; speaker adaptive training; polyglot speech synthesis; Tibetan speech Synthesis; Mandarin-Tibetan bilingual speech synthesis;
- 【会议录名称】 第十二届全国人机语音通讯学术会议(NCMMSC2013)论文集
- 【会议名称】第十二届全国人机语音通讯学术会议(NCMMSC’2013)
- 【会议时间】2013-08-05
- 【会议地点】中国贵州贵阳
- 【分类号】TN912.33
- 【主办单位】中国中文信息学会语音信息专业委员会、中国声学学会语言、听觉和音乐声学分会、中国语言学会语音学分会