节点文献
兰州方言语音生成方法研究
Research on Lanzhou-Dialect Speech Generation
【作者】 甘振业;
【作者基本信息】 西北师范大学 , 电路与系统, 2007, 硕士
【摘要】 本文提出了利用语音转换实现兰州方言语音的生成方法。在采用Pitch Target估计模型为声调模型的基础上,提出了采用线性修改模型(LMM)生成兰州方言的方法和采用高斯混合模型(GMM)生成兰州方言的方法。论文还提出了在生成方言语音的基础上采用语音修改方法实现音色可变兰州方言的方法。论文的主要工作及贡献如下:1.提出了兰州方言的声调表示方法。在声调模型的选择上,论文讨论了现今主要的声调模型。根据兰州方言语音的特点选择语音学模型中的Pitch Target估计模型作为声调表示模型。2.提出了一种基于线性修改模型(LMM)的兰州方言生成方法。对于训练集中的普通话语音和兰州方言语音利用Pitch Target估计模型提取特征参数,分别用七维的矢量表示两种语音的声调曲线,然后利用线性回归的方法分别求得七个特征参数的转换函数。在生成语音时,首先提取待转换普通话的七个特征参数,然后利用转换函数计算出兰州方言对应的七个特征参数,生成基频F0曲线,最后利用Straight算法合成方言语音。3.提出了基于高斯混合模型(GMM)的兰州方言变换方法,使得能够在大语料库的基础上,基于统计学模型,实现普通话到兰州方言的变换。首先利用Pitch Target模型提取源语音和目标语音的特征参数,构建方言变换的训练集;然后构建普通话和兰州方言训练语音库,训练出GMM的转换参数。根据转换参数进行方言变换,得到兰州方言的F0曲线,最后利用Straight算法合成出兰州方言。实验结果表明,增加训练音库的规模,可以得到质量更好的合成语音。4.提出了音色可变兰州方言语音的生成方法。影响语音听感的参数,主要包括时域和频域参数:基频、时长、非周期指数和频谱。利用Straight语音修改算法修改方言语音的基频、时长等时域参数和共振峰等频域参数,可以得到音色可变兰州方言语音。实验结果表明,该方法能够得到较高质量的多音色兰州方言语音。
【Abstract】 This dissertation proposed Lanzhou dialectal speech generation methods based on speech conversion. The dissertation adopt the Pitch Target estimates model as intonation model , proposed the generation method of Lanzhou dialectal speech based on Linear Modification Model (LMM) and Gaussian Mixture Model (GMM), and proposed a speech modification method for generating variety speech timbre Lanzhou dialectal speech. The main contributions of this dissertation are listed as follows:First, the dissertation proposed the intonation modeling method of Lanzhou dialectal speech. The dissertation discusses the present main intonation model. According to the characteristic of Lanzhou dialectal speech, choose the Pitch Target estimate model as intonation model.Second, the dissertation proposed the generation method of Lanzhou dialect based on Linear Modification Model (LMM). In this method ,we predict model parameters on the Mandarin speech and Lanzhou dialectal speech in the testing set, and use a 7 dimensions parameters denoting two speech F0 contours. Then, using line regression method calculates conversion function of 7 dimensions parameters. At the stage of generating speech, predict model parameters of candidate Mandarin speech, calculate accordingly Lanzhou dialectal speech 7 dimensions parameters, and generate its F0 contours, and synthesize the Lanzhou dialectal speech using Straight algorithm.Third, the dissertation proposed the generation method of Lanzhou dialectal based on Gaussian Mixture Model (GMM). This method can work on a big corpus based on statistics model. Firstly, using Pitch Target model predict feature parameters of the Mandarin speech and Lanzhou dialect speech in the training set, and train GMM conversion parameters. According to GMM conversion parameters, we get converted F0 contours of Lanzhou dialect speech. Then synthesize the Lanzhou dialect speech using Straight algorithm. The result show that increases the scale of training speech set, we can get better synthesizes speech.Forth, the dissertation proposed the generation method of variety speech timbre Lanzhou dialectal speech. The parameters which influence listening sense are pitch, duration, aperiodic exponent and frequency spectrum. Modifying the pitch, duration, aperiodic exponent and frequency spectrum of dialect speech using Straight algorithm, we can get variety speech timbre Lanzhou dialect speech. The results show we can get high quality Lanzhou dialectal speech by this method.
【Key words】 Lanzhou dialect; Pitch Target estimates Model; GMM model; speech conversion; Straight algorithm;
- 【网络出版投稿人】 西北师范大学 【网络出版年期】2008年 07期
- 【分类号】TN912.3
- 【被引频次】9
- 【下载频次】400