节点文献
基于P.563的话音质量客观评价
Objective Speech Quality Assessment Based on P.563
【作者】 杨波;
【导师】 别红霞;
【作者基本信息】 北京邮电大学 , 信号与信息处理, 2014, 硕士
【摘要】 随着各类多媒体业务的不断发展和普及,如何衡量多媒体业务的质量已经逐渐成为一个热点的问题。话音业务是多媒体通信的核心业务,它直接决定了人与人之间的沟通效率。话音信号的质量评价方法可以分为主观评价和客观评价两种,主观评价方式费时、费力,但是可靠性高,客观评价方式简单、快捷,但是准确度低。P.563算法是国际电信联盟确立的第一个话音客观质量单端评价标准,但是其算法设计存在一定的缺陷,使它的应用范围受到限制。论文从话音信号的预处理、特征参数提取、评价结果的映射模型三个方面对P.563算法加以改进,并基于RealV210开发平台实现了一套便携式话音质量客观评价系统。首先,在话音信号的预处理阶段基于频谱分布特点检测单频噪声,以提高P.563算法对含单频噪声话音的评价准确度。其次,提取能反映入耳感知特性的美尔倒谱系数,并建立未失真话音信号美尔倒谱系数的GMM模型。然后,利用GMM建立6种话音失真类型的参考模型,根据话音信号特征矢量与参考模型的距离判定话音失真类型。通过MARS技术建立从多维话音特征参数到客观评价结果的映射模型。最后,基于‘RealV210开发平台设计并实现了一套便携式的话音质量客观评价系统,他具有系统自检和实时话音质量评价两种工作模式。针对本文所构建的测试话音库,改进后算法的客观质量评价结果与主观评价结果之间的相关系数从0.44提高到了0.82,显示出良好的算法性能。
【Abstract】 With the continuous development and popularization of various multimedia services, to measure the quality of multimedia services has gradually become a hot issue. Voice service is of vital importance in the multimedia communication system, which determines the efficiency of communication between human beings directly. Methods for speech quality assessment can be mainly divided into two classes which are subjective evaluation and objective evaluation. Subjective evaluation of speech quality is time-consuming and expensive while it has higher reliability.In the contrast, objective quality assessment is quite simple and quick, but its accuracy of evaluation is much lower.The P.563is the first single-ended method for objective speech quality assessment established by the International Telecommunication Union. But the designing of algorithm in the P.563has some defects, which limit the scope of its application. In this paper, the P.563algorithm is improved from three aspects respectively, which are the pre-processing of speech, feature extraction of speech and the mapping model of evaluation result. Besides, a portable system for objective evaluation of speech quality is developed based on RealV210.Firstly, to detect the single-frequency noise in the speech based on the properties of its spectrum distribution during the pre-processing stage which aims to improve the predicting accuracy of the P.563for speech with single-frequency noise.Secondly, the Mel-Frequency Cepstral Coefficients (MFCCs) are extracted from undegraded speech signal which reflect the perception of human beings’ audio system, and build a Gaussian Mixture Model (GMM) for them.Thirdly, to build six kinds of reference models for6types of speech distortion using GMM and determine the distortion type of speech according to the distance between the feature vector and these reference models. Build the mapping relation between the multidimensional characteristic of speech and the result of objective evaluation.Finally, design and develop a set of portable system for objective assessment of speech quality based on RealV210, the system has two operating modes which are the self-inspection mode and the real-time speech quality evaluation mode.As to the speech database for testing built in this paper, the correlation between results of subjective evaluation and results of the algorithm in this paper is improved from0.44to0.82, which shows the favorable performance.
【Key words】 speech quality; objective assessment; MARSRealV210; single-frequency noise; Gaussian mixture model;