节点文献

基于语音信号的多模态情感识别关键技术研究

Research on Key Technologies of Multimodal Emotion Recognition Based on Speech Signals

【作者】 刘凯;

【导师】 朱维红;

【作者基本信息】 山东大学 , 电子信息(专业学位), 2023, 硕士

【摘要】 近年来,愈加丰富的计算机软件资源和不断提升的硬件设备极大地促进了自动语音识别、语音情感识别等语音领域相关技术的发展,智能语音交互作为人机交互的重要模块,已经成为人们日常生活中不可或缺的辅助工具。人们期待语音交互机器具有与人相似的观察、理解能力,并产生相应的情感状态,对所提问题做出更精确的应答,因此如何利用这些技术挖掘出语音数据中更深入的情感价值显得尤为重要。虽然语音能独立地进行情感理解和表达,但人与人之间是通过综合不同模态的信息进行相互交流,若仅通过单一的语音进行情感状态的判别,其准确度受限且片面。考虑到文本和语音两种模态数据之间存在一定的互补性,论文中首先利用语音识别技术获取语音中的文字信息,然后充分挖掘出文本和语音中的情感特征信息并将之融合,最后进行多模态情感识别。论文主要贡献总结如下:(1)针对语音识别任务中特征序列的长距离依赖问题,设计了一种卷积神经网络和多头注意力机制相结合的网络模型。该模型中卷积神经网络用于提取局部特征信息,多头注意力机制则获取全局信息并根据特征对输出的贡献大小进行加权,从而有效缓解长距离依赖问题,提高了模型的准确率。(2)针对多头注意力机制中头数(Attention Head)与子空间维度难以平衡的问题,设计了带有膨胀卷积和多头注意力机制的多分支融合网络模型。该模型分为多条并行的分支,每条分支中特征向量首先经过膨胀卷积网络,该网络可有效扩大卷积层的感受野。每条支路中膨胀卷积网络的膨胀率各不相同,这导致每条分支中多头注意力机制所获取的感受野也各不相同,多头注意力机制只会关注自身感受野范围内的信息而不过分关注全局信息,从而使得每条分支中的信息量大大减少,因此该模型可以使用更多的注意力头数获取更好的模型识别性能。(3)针对语音和文本这两种时间序列情感建模时,存在的不同模态情感特征难以对齐问题和跨模态长距离依赖问题,设计了一种基于交互注意力与自注意力机制的双流跨模态特征融合网络模型。该模型引入交互注意力机制,可以跨模态融合文本情感特征和语音情感特征,使得模型不再需要对齐两类模态的特征信息,并且两类情感特征充分互补提高了模型在有噪声环境下的鲁棒性。同时模型利用自注意力机制使每个模态能获取其全局上下文信息,缓解自身模态的长距离依赖问题。两种注意力机制的结合使得情感特征信息得到充分利用,显著提升了模型分类准确率。(4)针对多模态情感分类问题,设计了一种基于重加权BiGRU的多模态情感分类模型。该模型将BiGRU正反向隐藏层状态向量与输出向量相乘,获得每个时间步上的情感特征对输出的加权因子,并对输出向量重加权,充分利用正反向隐藏层中隐含的情感特征信息,抑制对输出影响较小的情感特征,从而有效提升了模型的分类效果。

【Abstract】 In recent years,with the enrichment of computer software resources and the continuous improvement of hardware equipment,the development of speech related technologies such as Automatic Speech Recognition(ASR)and Speech Emotion Recognition(SER)has been greatly promoted.As an important module of human-computer interaction,Intelligent Voice Interaction(IVI)has become an indispensable auxiliary tool in people’s daily life.People expect speech interaction machines to have similar observation and understanding abilities to humans,and generate their own corresponding emotional states to respond more accurately to questions raised by people.Therefore,how to use these technologies to mine the deeper emotional value of voice data is particularly important.Although speech can independently understand and express emotions,people communicate with each other through the integration of different modes of information.If a person’s emotional state is directly determined from only one expression method of speech,the results of the discrimination may be one-sided and limited in accuracy.Considering the complementarity between text and speech,the proposed algorithm first utilizes ASR to obtain text information from speech,then fully mining and fusing emotional feature in text and speech,and finally performing multimodal emotional recognition.The main contributions of this thesis are summarized as follows:(1)Aiming at the problem of long-distance dependence of feature sequences in ASR,a network model combining convolutional neural networks(CNN)and multi-head attention is designed.In this model,CNN is used to extract local feature information,while the multi-head attention obtains global information and weights the output based on the contribution of the feature,effectively alleviating the long-distance dependency problem and improving the accuracy of the model.(2)In view of the difficulty of balancing the attention head and subspace dimensions in multi-head attention,a multi branch fusion network with dilated convolutional networks(DCNN)and multi-head attention(DAMBFN)is designed.The model is divided into multiple parallel branches,and the features in each branch first pass through the DCNN,which can·effectively expand the receptive field of the convolutional layer.The dilated rate of the DCNN in each branch is different,which results in different sensory fields obtained by the multi-head attention in each branch.The multi-head attention only focuses on information within its own sensory field without excessively focusing on global information,which greatly reduces the amount of information in each branch.Therefore,the model can use a larger number of attention heads to obtain better model performance.(3)Aiming at the difficult alignment of different modal emotional features and longdistance dependence across modes in emotion modeling of speech and text time series,a dualflow cross modal feature fusion network(DCFFN)based on interactive attention and selfattention is designed.The model introduces an interactive attention that can fuse text emotional features and speech emotional features across modes,which eliminates the need for the model to align the feature information of the two types of modes,and the full complementarity of the two types of emotional features improves the robustness of the model in noisy environments.At the same time,the model utilizes a self-attention to enable each mode to obtain its own global context information.The combination of the two attention mechanisms makes full use of emotional feature information,significantly improving the accuracy of model classification.(4)For multimodal emotion classification,a multimodal emotion classification model based on Reweighted BiGRU(ReBiGRU)is designed.This model multiplies the state vector of the BiGRU forward and reverse hidden layers with the output vector to obtain the weighting factor of the emotional features on the output at each time step,and then reweights the output vector.This model fully utilizes the emotional feature information hidden in the forward and reverse hidden layers,effectively suppressing emotional features that have a small impact on the output,thereby improving the classification effect of the model.

  • 【网络出版投稿人】 山东大学
  • 【网络出版年期】2024年 04期
  • 【分类号】TN912.34
节点文献中: 

本文链接的文献网络图示:

本文的引文网络