节点文献

基于卷积神经网络和多特征融合的语音情感识别研究

Research on Speech Emotion Recognition Based on Convolutional Neural Network and Multi-feature Fusion

【作者】 韩梅;

【导师】 胡黄水; 孙翠玲;

【作者基本信息】 长春工业大学 , 电子信息(专业学位), 2024, 硕士

【摘要】 随着人工智能的发展以及智能化人机交互需求的增加,语音情感识别技术成为当前的研究热点。语音情感识别旨在从语音信号中捕获说话者的情感状态,能够帮助智能系统实现更加自然、个性化地交互。准确识别语音中的情感状态对于推动人机交互智能化发展具有重要意义。提取有效的情感特征是语音情感识别的关键。因此,本文基于卷积神经网络和多特征融合的语音情感识别算法展开研究,旨在从语音信号中获得更具辨别力的情感特征,以提高语音情感识别的准确率。具体研究工作如下:(1)针对传统卷积神经网络无法充分提取MFCC谱图中不同维度情感信息的问题,提出了基于多尺度扩张卷积和全局特征融合的语音情感识别算法。算法主要由浅层特征细化模块、多尺度扩张卷积网络模块和全局特征融合模块组成。浅层特征细化模块中使用并行卷积捕捉MFCC特征的时间、频率和时间-频率三个维度的浅层特征,并通过收缩块去除与情感不相关的特征,实现了对浅层多维特征的细化。多尺度扩张卷积网络中结合Res2net结构、扩张卷积和通道注意力构建了多尺度扩张卷积块,以进一步获取更加细粒度的局部多尺度信息,提高模型对局部细节的感知能力。使用并行的门控多层感知机设计了全局特征融合模块,并利用乘法门控控制信息流,将不同尺度的局部特征进行整合,得到全局的特征表示,以提高模型对整体语音信号的理解能力。算法在IEMOCAP数据集上进行了实验,结果表明在全部数据上的加权准确率和未加权准确率分别达到了75.78%和76.10%。(2)针对传统多特征融合中采用简单拼接而忽视不同特征之间交互信息的问题,提出了一种基于多特征交叉融合的语音情感识别算法。该算法将原始语音信号、MFCC特征和语谱图作为模型的输入,使用特征自适应交叉融合模块融合不同特征之间的情感信息。特征自适应交叉融合模块首先利用交叉注意力机制构建了交叉融合块捕获特征间的交互信息,同时,引入了注意力统计池化单元在聚合帧级情感表征时聚焦于更重要的特征,以增强模型提取话语级情感表征的能力,最后使用特征自适应加权策略根据不同特征对最终情感识别贡献度的不同,自适应的学习最优的融合系数。在IEMOCAP数据集和EMO-DB数据集上对算法进行性能测试,实验结果表明在IEMOCAP数据集上实现了72.66%的加权准确率和73.41%的未加权准确率,在EMO-DB数据集上实现了91.58%的加权准确率和91.89%的未加权准确率。(3)基于所提出的语音情感识别算法,利用Py Qt5设计并开发了语音情感识别系统。能够实现用户注册和登录、语音采集以及情感识别等功能。通过对系统进行综合功能测试,结果表明其人机交互友好,能够准确识别语音情感,具有一定的应用价值。

【Abstract】 With the development of artificial intelligence and the increased demand for intelligent human-computer interaction,speech emotion recognition technology has become a current research hotspot.Speech emotion recognition aims to capture the speaker’s emotional state from the speech signal,which can help intelligent systems achieve more natural and personalized interactions.Accurately recognizing emotional states in speech is crucial for advancing the development of human-computer interaction.Extracting effective emotional features is the key to speech emotion recognition.Therefore,in this paper,a research on speech emotion recognition algorithm based on convolutional neural network and multifeature fusion is carried out,aiming at obtaining more discriminative emotional features from speech signals in order to improve the accuracy of speech emotion recognition.The specific research work is as follows:(1)To address the problem that traditional convolutional neural networks cannot adequately extract different dimensions of emotion information in MFCC spectrograms,a speech emotion recognition algorithm based on multi-scale dilation convolution and global feature fusion is proposed.The algorithm comprises a shallow feature refinement module,a multi-scale expansion convolutional network module and a global feature fusion module.The shallow feature refinement module uses parallel convolution to capture the shallow features of MFCC features in three dimensions: time,frequency,and time-frequency.and removes the features that are not relevant to the sentiment by shrinking the block,which realizes the refinement of shallow multidimensional features.A Multi-Scale Dilated Convolution Block is constructed in the Multi-Scale Dilated Convolution Network by combining the Res2 net structure,dilated convolution,and channel attention to further acquire finer-grained local multiscale information and improve the model’s ability to perceive local details.A global feature fusion module is designed using parallel gated multilayer perceptron,and multiplicative gating is used to control the information flow and integrate local features at different scales to obtain a global feature representation to improve the model’s ability to understand the overall speech signal.The algorithm was experimented on the IEMOCAP dataset,and the results indicate that The algorithm was experimented on the IEMOCAP dataset and the results showed that the weighted accuracy and unweighted accuracy on the full data reached 75.78% and 76.10%,respectively.(2)To address the problem that traditional multi-feature fusion uses simple splicing while ignoring the interaction information between different features,a speech emotion recognition algorithm based on multi-feature cross-fusion is proposed.The algorithm takes raw speech signals,MFCC features and speech spectrograms as inputs to the model,and uses a feature adaptive cross-fusion module to fuse emotional information between different features.The feature adaptive cross-fusion module first constructs a cross-fusion block using the cross-attention mechanism to capture the interaction information between features.At the same time,introduces an attention statistics pooling unit to focus on more important features when aggregating frame-level emotion representations in order to enhance the model’s ability to extract discourse-level emotion representations.Finally,a feature adaptive weighting strategy is used to adaptively learn the optimal fusion coefficients based on the different contributions of different features to the final emotion recognition.The algorithms are tested for performance on the IEMOCAP and EMO-DB datasets,and the experimental results showed that a weighted accuracy of 72.66% and an unweighted accuracy of 73.41%were achieved on the IEMOCAP dataset,and a weighted accuracy of 91.58% and an unweighted accuracy of 91.89% were achieved on the EMO-DB dataset.(3)Based on the proposed speech emotion recognition algorithm,a speech emotion recognition system is designed and developed using Pyqt5.It can realize functions such as user registration and login,voice collection,and emotion recognition.Through the comprehensive function test of the system,the results show that its human-computer interaction is friendly,and it can accurately recognize the speech emotion,which has certain application value.

  • 【分类号】TP183;TN912.34
节点文献中: 

本文链接的文献网络图示:

本文的引文网络