节点文献

汉语基元音素独立分量谱分析对比及语音合成研究

【作者】 尉洪

【导师】 施心陵;

【作者基本信息】 云南大学 , 通信与信息系统, 2011, 博士

【摘要】 语音合成技术是实现人机语音交互通信的关键技术之一,它希望计算机具备像人一样的说话能力。能灵活调整合成单元的音段参数和超音段参数,同时确保合成语音的高自然度是目前面临的一个主要问题。独立分量分析方法区别于传统的DFT、小波变换等分析方法,论文利用独立分量分析方法的优势,提取基元独立分量,分析其声学特征并结合语音合成展开探索性研究。论文应用独立分量分析方法,研究汉语发音基元时域和频域独立分量信号可区分的声学特征,结合基元生物发声机理讨论各独立分量的含义;对比分析基元独立分量传统短时FFT谱包络与LPC声道谱包络、高阶Wigner-Ville谱包络声学特性,研究分析在基元合成实验中的合成效果;通过基频曲线调整合成基元调域,对独立分量谱包络按共振峰特性加窗处理和调整各独立分量间混合权重来控制合成基元音色。论文的主要工作如下:1、论文使用独立分量分析方法,从时域提取各发音基元独立分量。对比分析了基元各时域独立分量间相关性大小、基频FO线、共振峰特性、F1-F2,F2-F3声学音位图等声学特征,发现基元各时域独立分量间可区分的特征。结合各发音基元的生物发声机理,声带的振动频率与基频对应,元音发音时舌位的高低与第一共振峰频率F1对应,舌位的前后与第二共振峰频率F2对应等,将基元各时域独立分量进行鉴别区分,赋予各独立分量确切的含义,如高基频分量,高舌位分量,前舌位分量等。频域ICA分析中,获取了基元频谱包络的独立分量。对比分析了蕴含在基元各频谱独立分量中的共振峰特性和F1-F2, F2-F3声学音位图,找出基元各频谱独立分量间可区分特征,将基元各频谱独立分量分别区分为高舌位谱分量,前舌位谱分量等。2、在时域ICA分析中,对同一发音基元各时域独立分量,提取了其传统短时FFT谱包络与LPC声道谱包络、高阶Wigner-Ville谱包络,对比分析了蕴含在三种频谱包络中的共振峰特性和谐波结构,发现三种频谱包络间的声学特征差别;对比分析了传统短时FFT谱包络与LPC声道谱包络、高阶Wigner-Ville谱包络在基元合成实验中的效果。实验环节,应用STRAIGHT合成算法,基于各基元独立分量的基频和三种不同的频谱包络,完成了各发音基元时域独立分量合成和时域独立分量混合合成实验。基于各基元三种不同频谱包络的谱独立分量,完成了基于谱独立分量的基元合成和基于谱独立分量混合的基元合成实验。实验结果表明,三种频谱包络有各自不同的声学表现,基元LPC声道谱包络表现出了较平缓的声道传输特性,共振峰结构较钝化,而WV谱包络拥有更加丰富的谐波特性,更尖锐的共振峰结构和更高的频率分辨率,信号的一些快速时变特征在WV谱包络上也有体现。从基元合成效果来看,WV谱合成基元清晰度可懂度较优,传统FFT谱合成效果次之。3、论文针对各发音基元时域独立分量的谱包络按第一、二共振峰特性进行加窗处理,获取不同的音色表现。将不同特性的独立分量按不同的权值加权组合产生出音色可调控的合成语音,通过基频曲线调整合成基元的调域音高和情感特征。论文实验总结得到了音色调整的规则1、规则2和规则3,用来调控合成语音的基频和频谱包络中共振峰特性。实验结果显示,谱包络的加窗处理对音色的调整可控制在一个较满意的范围内,没有出现合成语音清晰度可懂度急剧下降的情况。经加权混合处理后的合成基元效果比音色相对单纯的各独立分量合成基元信号有更丰富的表现力,但音色的调整处理基于独立分量进行,对合成音质的影响会更细腻一些。合成基元清晰度可懂度经MOS评测,时域独立分量基元合成平均得分在4.5,时域独立分量谱加窗基元合成平均得分在4.53,时域独立分量加权混合基元合成平均得分在4.8左右。基于谱独立分量的基元合成平均得分在4.45,基于谱独立分量混合的基元合成平均得分在4.6左右。

【Abstract】 Speech synthesis is one of the most important technology of the man-machine interface, and the ultimate aim is to make the computer to have the human speech ability. At present, one of the dominant problems is that, we can adjust the segmental and suprasegmental acoustic parameters flexibly, and simultaneously may ensure the higher naturalness of the synthesized speech.The independent component analysis is different from the traditional other approach, such as DFT, wavelet transform etc. It takes advantage of the independent component analysis approach to extract the independent components of the mandarin phoneme, and to analysis their acoustic characteristics. Based on the above analysis results, the exploratory research on the speech synthesis is developed.Based on the independent component analysis technique, the discriminating acoustic characteristics of the independent components of mandarin phoneme are discussed in the time and frequency domain, and the meanings of every independent component are identified by the combination of the acoustic mechanism of the mandarin phoneme. The traditional FFT spectral envelope, the LPC vocal track spectral envelope, and the higher order Wigner-Ville spectral envelope of the phoneme independent components are compared and analysized. The effects of the three spectral envelopes in the phoneme synthesis experiments are presented. The timbre of synthesized phoneme is controlled by adjusting the fundamental frequency curve, by windowing the spectral envelope of independent component on formants position, and by adjusting the mixing weight among the independent components.The main research of the paper is as follow:1. Based on the independent component analysis approach, every independent component of the phoneme in time domain is extracted. The correlation, the fundamental frequency FO curve, the acoustic position in F1-F2 and F2-F3 space and the formants of independent components are compared and analysized. Further, the distinguishable characteristics among the phoneme temporal independent components are found. With the help of the acoustic mechanism of the mandarin phoneme, considering the relation between the fundamental frequency and the vibrating frequency of vocal cords, the relation between the first formant F1 and the tongue position (high or low) in vowel utterance, and the relation between the second formant F2 and the tongue position (front or back) in vowel utterance, every temporal independent component of the phoneme is identified, and given a certain meaning, such as high fundamental frequency component, high tongue position component, front tongue position component.In the process of independent component analysis in frequency domain, every spectral independent component of the phoneme is extracted. The acoustic position in F1-F2 and F2-F3 space and the formants traits of the spectral independent components are compared and analysized. Further, the distinguishable characteristics among the phoneme spectral independent components in frequency domain are found. Every spectral independent component of the phoneme is distinguished and given a certain meaning, such as high tongue position spectral component, front tongue position spectral component.2. In the process of independent component analysis in time domain, with the same one phoneme independent component, the traditional FFT spectral envelope, LPC vocal track spectral envelope, and the higher order Wigner-Ville spectral envelope are extracted. The formants and harmonic structure hidden in the three spectral envelopes are compared, and the discriminating acoustic characteristic among the above spectra are found. The effects of the FFT spectrum, LPC spectrum and Wigner-Ville spectrum in the phoneme synthesis experiments are presented. In the experiments, applying the STRAIGHT algorithm, based on the fundamental frequency and three different spectral envelope of every phoneme independent component, the phoneme synthesis experiments with temporal independent component and with mixing temporal independent components are implemented. Based on the spectral independent components from the three different spectral envelope of every phoneme, the phoneme synthesis experiments with spectral independent components and with mixing spectral independent components are implemented.The experimental results show that, the three spectral envelopes have their obviously different acoustic characteristic. The phoneme LPC spectral envelope has revealed the gentle transferring traits of vocal tract, and the blunt formants structure. The Wigner-Ville spectrum has the more abundant harmonic components, more sharp formants, and higher frequency absolution. Some quick time-variant characteristics are displayed in the WV spectral envelope. From the effects of phoneme synthesis, the articulation and intelligibility of the synthesized phoneme with WV spectral envelope is better than with FFT spectra, and the LPC spectra comes third.3. The spectral envelope of every phoneme temporal independent component is windowed on the first and the second formants, to acquire the different timbre. The different independent components are combined by different weight to generate the timbre-controlled synthesized phoneme. The pitch and emotional traits of the synthesized phoneme are adjusted by fundamental frequency curve. In the paper, the rule 1, rule 2 and rule 3 of adjusting timbre are summarized to control the pitch and the formants in the spectral envelope of the synthesized phoneme.The experimental results show that, the timbre adjusted by windowing the spectral envelope is controlled within the satisfied range, not emerging the case in which the articulation and intelligibility of the synthesized phoneme decline sharply. The phoneme synthesized by weighting the independent components has revealed more expressive effects than only by every pure independent component. Based on the adjustment of independent components, the more exquisite timbre effect of the synthesized phoneme is acquired. The articulation and intelligibility of the synthesized phoneme is evaluated by mean opinion score. The mean score of the phoneme synthesis with temporal independent components is 4.5, the mean score of the phoneme synthesis with windowing temporal independent components spectrum is 4.53, and the mean score of the phoneme synthesis with mixing temporal independent components is 4.8. The mean score of the phoneme synthesis with spectral independent components is 4.45, and the mean score of the phoneme synthesis with mixing spectral independent components is 4.6.

  • 【网络出版投稿人】 云南大学
  • 【网络出版年期】2012年 01期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络