节点文献

基于深度学习的立体图像舒适度质量评价研究

Stereoscopic Image Quality Assessment Based on Deep Learning

【作者】 王凯

【导师】 周军;

【作者基本信息】 上海交通大学 , 信息与通信工程专业, 2017, 硕士

【摘要】 近年来,3D立体电影等大范围普及流行,为了提升观看3D电影等的舒适度,很多舒适度评价方法被提出,这些方法均可归类到主观质量评价和客观质量评价两大类中,后者在兼顾评价效率的同时降低了质量评价的成本,因而近年来被广泛研究。就大部分图像/视频的客观质量方法而言,均是在视差图、深度信息等基础上人工提取有效特征,易导致重要特征的遗漏或提取的特征存在瑕疵。另一方面,近些年深度学习快速发展,已经成功运用在语音识别、图像分类、文本分析等场景,其重要特色之一就是可以自动学习特征,因而相对人工提取更容易学习到全面细微的重要特征。本文在此思路的指导下,提出了用于自动提取3D立体图像特征的深度神经网络。在该深度神经网络中,左右两路卷积限制玻尔兹曼机(CRBM)被用来提取初始图像特征,顶层因子化的三阶玻尔兹曼集(FTO-RBM)对左右图像特征混合训练,经多层全连层神经网络连接后得到预估值并构建3D立体图像质量评价模型,而后通过后向传播算法微调整个深度神经网络。随后,本文分析并调试了深度神经网络的一些重要参数,并使用LIVE 3D Phase II和IEEE-SA公开库及基于单刺激、成对比较主观质量评价方法的3D立体图像库对该模型进行了测试,验证了该模型的性能。基于上述3D立体图像质量评价模型,本文分析了其可改进之处,并提出了基于池化的遍历优化算法。该算法主要基于平均池化和特征权重分布池化方法,并通过遍历的方式优化现有特征,更新特征图后再基于支撑向量回归算法构建了3D立体图像质量评价模型。随后,分析并调试了优化算法中的重要参数,研究了左右两路CRBM合并优化的效果。此后再次进行了测试,结果显示优化后的模型相较优化前有更佳的性能,且达到了现有的最佳3D立体图像质量评价水准。在对模型优化后,本文基于单刺激及成对比较方法分析了主观评价质量对于模型的影响。本文通过对单刺激及成对比较主观评价方法得到的平均意见值(MOS)添加不同方差的高斯噪声,分析了包含不同高斯噪声的MOS值对模型的性能影响。此外,为了能适应现在日益庞杂的数据量,本文对用于特征提取的深度神经网络做了GPU并行化加速处理。此处主要针对深度神经网络的离线训练和在线测试两部分做了并行优化。在离线训练阶段,考虑到基于Python的Theano库在自动求导、模块化程度高等方面的优势,使用该库对深度神经网络做并行优化,并对上述两大公开库进行了实验,就实验结果做了详细对比及分析。在线测试阶段,考虑到其对加速性能要求更高,此处使用CUDA对测试过程的各步骤进行并行化处理,给出了核函数及线程分配的分析,使得加速性能有了进一步提升。

【Abstract】 In recent years,3D stereoscopic films have been widely popularized.In order to enhance the comfort of watching 3D movies,many comfort evaluation methods have been proposed.All these methods can be classified into subjective image quality assessment methods and objective image quality evaluation methods.The latter has been widely studied in recent years because of the efficiency of evaluation and the lower cost of image quality assessment(IQA).Most of the objective image/video quality evaluation methods are based on manually extracted valid features,which come from the disparity map,depth information and so on.Such methods may result in missing important features easily,or make extracted features flawed.On the other hand,in recent years,deep learning has been rapidly developed and successfully applied in speech recognition,image classification,text analysis,etc..One of its important features is,compared to manually extract features,the ability to automatically learn the comprehensive and detailed features.Under the guidance of this idea,this paper proposes a deep neural network for the automatic extraction of 3D stereo image features.In this depth neural network,(CRBM)is used to extract the initial image features,the top-level factorized third-order Boltzmann sets(FTO-RBM)are applied to the left and right image feature mixture training,and the multi-layer full-Layer neural network to obtain the estimated value and construct the 3D stereoscopic image quality evaluation model,and then fine-tune the entire depth neural network by backward propagation algorithm.Afterwards,this paper analyzes and debugs some important parameters of the depth neural network,and tests the model using LIVE 3D Phase II and IEEE-SA open library and 3D stereo image database based on single stimulus and pairwise comparison subjective quality evaluation method,The performance of the model was verified.Based on the above three-dimensional image quality evaluation model,this paper analyzes the improvement of the 3D image quality evaluation model and puts forward a new optimization algorithm based on iteration.The algorithm is based on the average pooling and feature weight distribution pooling method.The existing feature is optimized by traversing,and the 3D image quality evaluation model is constructed based on the support vector regression algorithm after updating the feature map.Then,the important parameters in the optimization algorithm are analyzed and debugged,and tested again using the above two public libraries.The results show that the optimized model has better performance than the optimized one and achieves the best3 D Stereo image quality evaluation level.After optimization of the model,the influence of the subjective evaluation quality on the model was analyzed based on the single stimulus and pairwise comparison method.In this paper,the influence of MOS model with different Gaussian noise on the performance of the model is analyzed by adding the Gaussian noise with different variance to the average observation value(MOS)obtained by the subjective evaluation method of single stimulus and pairwise comparison.In addition,in order to be able to adapt to the increasingly complex data volume,this paper makes GPU parallelization accelerated processing for the deep neural network used for feature extraction.In this paper,we focus on the parallel optimization of two parts: offline training and online testing.In the offline training phase,taking into account the advantage of the Python-based Theano library in automatic derivation and high modularity,this library is used to parallelize the optimization of the deep neural network,and the two public libraries are experimented.Made a detailed comparison and analysis.In the online testing phase,considering the requirement of higher performance,we use CUDA to parallelize the steps of the testing process.The analysis of the kernel function and the thread allocation is given,which makes the acceleration performance further improved.

  • 【分类号】TP391.41;TP18
  • 【被引频次】1
  • 【下载频次】52
  • 攻读期成果
节点文献中: 

本文链接的文献网络图示:

本文的引文网络