节点文献

用语谱图融合小波变换进行特定人二字汉语词汇识别

Lexical semantic recognition for Chinese two-character words based on wavelet transform with fusion of spectrograms

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 魏莹王双维潘迪张玲许廷发梁士利

【Author】 WEI Ying;WANG Shuangwei;PAN Di;ZHANG Ling;XU Tingfa;LIANG Shili;School of Physics, Northeast Normal University;College of Science, Changchun University of Science and Technology;Key Laboratory of Photo Electronic Imaging Technology and System, Ministry of Education of China ( Beijing Institute of Technology);

【机构】 东北师范大学物理学院长春理工大学理学院光电成像技术与系统教育部重点实验室(北京理工大学)

【摘要】 传统的语音分析都是建立在短时平稳假定的基础上,采用固定窗傅立叶变换获取语音信号的时-频局部化信息,与非平稳的语音信号不完全吻合,为此提出一种新颖的鲁棒性的谱图融合的小波变换方法。在图像特征提取过程中,对二字汉语词汇语音的语谱图进行特征分析,首先采用二维离散db4小波基分别对宽窄带语谱图进行6层小波包分解,并计算出每层的水平细节能量值、垂直细节能量值和对角细节能量值。接着,将窄带语谱图提取出的水平细节能量值、垂直细节能量值和对角细节能量值,分别作为窄带语谱图的第1~3个特征集合。然后将宽带语谱图提取出的水平细节能量值作为第4个特征集合。上述4个特征集合作为识别的特征向量,以支持向量机为分类器对特定人二字汉语词汇整体识别。采用1000个语音样本进行仿真实验,结果表明,该算法是利用语谱图的整体特征逐字逐词进行语音识别,能够凸显语音信号的整体时频特性,正确识别率可达98%。利用语谱图的特性,针对汉语的自身特性,将每一条语音指令作为一幅图像进行词汇研究,保证了语句的整体性,同时有助于提高识别率,增强鲁棒性。

【Abstract】 The traditional speech analysis is based on the short time stationary assumption, and time frequency localization information of speech signal is obtained by using Fourier transform with fixed window. But it is not suitable for nonsmooth speech signal processing extraction framework was developed for speech recognition. The basic idea was to extract useful information from the projection matrix of spectrogram. A novel robust approach of wavelet transform was presented with the fusion of spectrograms. In the process of image feature extraction, the feature analysis of the spectrum of the Chinese twocharacter words was made. Firstly, two-dimensional discrete db4 wavelet was respectively used to decompose broadband and narrowband spectrograms, which was divided into 6 layers of wavelet package decomposition, and calculated the approximate energy values. Then, the extracted approximate energy values for the narrowband spectrogram were divided into level vertical detail energy and diagonal detail energy value, set respectively as the narrowband spectrogram of the first, the second and the third feature set. Meanwhile, the level of detail energy value for the broadband spectrogram was extracted as the fourth feature set. The above four feature sets were used as the feature vector of support vector machine as a classifier for the overall recognition of Chinese two-character words. The 1 000 voice samples were used in the simulation experiment. The results show that the algorithm is based on the whole feature of the speech spectrum, and word by word speech recognition, highlights the overall time-frequency characteristics of the speech signal, and its correct recognition rate can reach 98 percent. In this paper,the characteristics of the language spectrum were used to study the characteristics of Chinese. The vocabulary of each voice command was treated as an image, to ensure the integrity of the statement, while helping to improve the recognition rate robustness.

【基金】 国家自然科学基金资助项目(61471111)
  • 【文献出处】 计算机应用 ,Journal of Computer Applications , 编辑部邮箱 ,2017年S1期
  • 【分类号】TN912.34
  • 【被引频次】1
  • 【下载频次】60
节点文献中: 

本文链接的文献网络图示:

本文的引文网络