节点文献

情感语音合成

Affective Speech Synthesis

【作者】 苏庄銮

【导师】 汪增福;

【作者基本信息】 中国科学技术大学 , 模式识别与智能系统, 2006, 博士

【摘要】 语音是最理想的人机交互方式之一,而语音合成技术则是实现语音人机交互的基础。从第一个电子语音合成器问世以来,随着各种新技术手段的应用,特别是近年来随着基于基音同步叠加、结合大规模自然语音库和数据挖掘等智能算法的语音合成方法的流行,语音合成技术在可懂度和自然度上达到了相当的水平,并且开始产业化,逐渐进入人们的日常生活。语音合成技术的推广应用,对语音合成的质量提出了更高的要求。如何进一步提高语音合成的表现力,特别是让合成语音能够模拟表达说话人的情感状态,是语音合成未来发展的趋势,也是语音合成研究领域所面临的一个难题。情感语音合成是一个跨学科的、具有很高理论价值和应用价值的研究课题;作为语音合成的一个新的研究方向,正受到众多研究者越来越多的关注。 本文以情感语音的基频特征为主要研究对象,以合成情感语音为主要研究目标,对基于基频特征的情感语音建模以及情感语调规则指导下的情感语音合成器设计等问题进行了较深入的研究。在此基础上,构建了一个语音合成系统,该系统除了可以验证本文提出的模型方法外,还可以作为语音处理相关研究的实验平台,为以后的研究工作创造良好的实验条件。 本文的主要创新点如下: (1)从情感语音处理的需求出发,通过对Fujisaki基频模型进行改进,提出了一种对情感语音基频进行建模的方法。该方法首先利用高通滤波器分离出基频曲线中的低频成分和高频成分,再分别从低频成分和高频成分中提取模型的短语命令参数和声调命令参数。提取命令参数时根据命令响应函数的特性,设计了从左往右、依次迭代提取的方法。该模型方法能够将基频曲线根据明确语音学含义进行参数化,并且模型参数分布与情感模式有一定的对应关系。同已有的同类研究相比较,本文所提出的基频模型能够反映语音的情感特征,所给出的模型参数提取方法简洁、有效,不需要任何手工标注。 (2)提出了一种数据驱动的语调模型方法,建立了特征语调的概念,并将相关的概念和方法用于分析普通话中情感模式对应的情感模式语调。在限定语料长度、结构以及说话人的前提下,采用主成分分析方法获取6个特征语调,借助这6个特征语调表示所有的语调。实验表明,本文提出的特征语调能够在可以接受的误差范围内拟合出所有语调,并且使用特征语调分析出的情感模式语调具有相当的情感表达能力。另外,本文还对采用特征语调表达混合情感模式的相关问题

【Abstract】 Speech is one of the perfect human-machine interfaces, and speech synthesis is a key technology for communication with speech between human and machine. Since the first speech synthesizer was born, with the application of new methods and techniques, especially the prevalence of combining massive raw speech database and intelligent algorithms such as data mining, TTS (Text To Speech) system based on pitch synchronous overlap adding has reached a high level on clarity and naturalness in recent years, began to be widely commercially used and will step into the people’s lives gradually. Being widely used, synthetic speech is required to be better. Improving the expressive ability for TTS system, especially letting the synthetic speech can express emotions like speaker, is accordant with the developing trend of speech synthesis. However, it is still a difficult problem lying ahead. As an interdisciplinary field, affective speech synthesis is a research topic with highly theoretical and applied value, and it has been a new direction of speech synthesis and has been focused on by more and more researchers.In order to synthesize the affective speech, this paper focuses on the fundamental frequency (F0) of affective speech and studies affective speech modeling based on F0 and affective speech synthesizer with the intonation-rules guidance, and some other related algorithms. Based on these studies, the paper has completed a speech synthesis system, which not only validated the modeling method proposed in the paper, but also can be an experimental platform for speech processing related research, and provide good experimental condition for the future research.The main innovative points of this paper are as follows:(1) In the paper an F0 modeling method for affective speech is proposed based on modified Fujisaki model, and a novel and effective approach is proposed to extract the parameters of the model automatically without any manual labels information. The approach separates the F0 contour into low frequency component (LFC) and high frequency component (HFC) with a high-pass filter, then estimates the phrase-command parameters of the model from LFC and the tone-command parameters from HFC. Because of the response characteristic of the command, a left-to-right iterative process is proposed to estimate the parameters in turn. The model can express the F0 contour with parameters which have explicit phonetic meanings and there is clear relationship between the distribution of the parameters and the emotion. Comparing with others, the F0 model proposed in this paper can express emotional features of affective speech. Furthermore, the method, which estimates the parameters of the model, is simple and effective, especially without any manual labels information.

节点文献中: 

本文链接的文献网络图示:

本文的引文网络