节点文献
复杂背景下多视角人脸检测与识别
Multi-view Face Detection and Recognition in Complex Background
【作者】 田春娜;
【作者基本信息】 西安电子科技大学 , 智能信息处理, 2008, 博士
【摘要】 自动人脸检测与识别属于模式识别和计算机视觉的交叉学科,具有重要的科学意义。随着低价位摄像装置的出现与计算机技术的快速发展,以及商业和军事应用领域需求的不断提高,人脸检测与识别技术的研究得到了愈加广泛的关注和参与。自然人脸图像受多种因素的影响具有很强的模式多样性,这使得人脸检测与识别问题极富挑战性。复杂背景和多姿态是人脸检测与识别技术走向实用化需解决的关键问题。本文系统地分析了现有人脸检测与识别方法的优缺点。针对如下问题进行了深入的研究:基于可能近似正确学习模型的人脸检测方法中的样本选择和多姿态问题;识别任务中的人脸对齐问题;多因素影响下人脸图像线性子空间识别方法中的视角非线性变化问题。主要的创新性研究成果概括如下:(1)针对主动机器学习中训练示例的选择问题,提出了嵌入式Bootstrap主动样本选择算法。通过将Bootstrap和嵌入式Bootstrap的思想进行形式化描述,从理论和实验中均证明了新算法在保持和原Bootstrap算法训练时间相当的前提下,可得到更典型的训练示例集,从而解决了计算条件对训练集规模的限制,使训练所得的预测器具有更高的性能。并将Bootstrap和嵌入式Bootstrap应用于基于AdaBoost的正面人脸检测任务的负样本选择中,实验结果表明了理论分析的正确性。该算法适用于一大类主动学习中训练示例的选择问题。(2)当人脸姿态变化较大时,人脸图像可利用的共同特征明显减少,为解决多姿态人脸的类内差异,提出基于聚类有效性分析和FloatBoost的多姿态人脸检测算法。对多种姿态采取“分而治之”的策略,通过模糊c-均值聚类和基于修正的划分模糊度的聚类有效性分析对多姿态人脸图像进行分类,然后用FloatBoost学习得到树型的多姿态人脸检测器。与原树形检测器算法相比,新算法在训练阶段速度更快,在检测阶段更有效。(3)面部特征精确配准是实现鲁棒人脸识别的前提,为了更精确的描述人脸特征和提高分类的精度,提出一种新的人脸序列图像的对齐方法。文中借鉴人脸图像中的灰度分布平滑性,提出快速鲁棒的模糊连通度分割方法实现序列图像中的人脸分割,并根据人脸主要器官进行正面人脸对齐,实验结果表明该方法有助于识别性能的提高。此外,所提出的分割算法在多噪声和序列医学图像的分割上也取得较原基于尺度的分割方法更快的速度。(4)基于张量分解的线性子空间识别方法主要用于解决多因素影响下的复杂人脸识别,本文针对其中的视角非线性变化问题进行深入的研究,提出基于视角流形建模的多姿态人脸识别算法。首先采用张量分解将影响人脸图像的多种因素分离;再结合视角流形建模来解决人脸视角空间的非线性问题,最后给出身份和视角参数的自动求解方法。对测试视角与训练视角进行交叉验证的结果表明,与原基于张量脸的识别算法相比,新算法的识别率提高了12%左右。(5)针对人脸姿态低维流形与高维人脸数据之间的非线性映射关系,提出一种基于非线性张量分解的人脸生成模型。文中深入地研究了多视角人脸表达中的三种视角流形生成方法的有效性,并提出一种基于EM-like的模型参数求解方法。对测试视角与训练视角进行交叉验证的结果表明,与基于张量脸的识别算法相比,识别率提高了19%左右。此外,基于非线性张量分解的方法为混合线性-非线性因素影响下的精细目标建模提供了有效途径。上述研究成果分别从复杂背景条件下的多姿态人脸检测、对齐与识别等方面给出了具体的研究方案和实验结果,为人脸检测与识别的理论研究和应用推广提供了新思路。此外,文中提出的主动样本选择方法、图像分割方法、多姿态人脸检测和识别模型具有一定的通用性,为相关的模式识别问题提供了有效的理论依据和有效途径。
【Abstract】 Automatic face detection and recognition is emerging as an active research topic in the areas of pattern recognition and computer vision. At least two reasons account for this trend: one is the availability of low cost cameras and rapid progress of powerful personal computers; another is the wide range of commercial and law enforcement applications. It is a challenging topic because natural face images are formed by the interaction of multiple factors which broaden the diversity of face images. Complex background and multi-view are key problems that have to be solved in real enviorment based face detection and recognition tasks.A systemic review of existing face detection and recognition algorithms are given in this paper. Some crucial problems are discussed, for examples, the example selection and multi-view problems of probably approximately correct learning model based face detection, frontal face alignment algorithm, view manifolds and generative face models for multi-view face recognition. The main achivements of this paper are summarized as follows.(1) To handle the computation resource constraints to the size of training example set, an embedded Bootstrap example selection algorithm is proposed for active machine learning from example. Through formulation, theoretical comparison analysis indicates that the embedded Bootstrap algorithm, using almost the same training time with the traditional Bootstrap, selects more utility examples to represent a potentially overwhelming quantity of training samples. Thus a more effective predictor can be trained. Furthermore, both algorithms are applied to the negative example selection for AdaBoost based face detection system. And the experimental results show that the embedded Bootstrap strategy outperforms the traditional Bootstrap, which agrees with the theoretical analysis. This algorithm adapts to a wide range example selection tasks.(2) Confronting with big view variation, the common features of certain human’s face images are dramastically reduced. Therefor, a novel face detection tree based on cluster validity analysis and FloatBoost learning is proposed to accommodate the in-class variability of multi-view faces. The tree splitting procedure is realized through dividing face training examples into the optimal sub-clusters using the fuzzy c-means algorithm together with a new cluster validity function based on the modified partition fuzzy degree. Then each sub-cluster of face examples is conquered with the FloatBoost learning to construct branches in the node of the detection tree. During training, the proposed algorithm is much faster than the original detection tree. The experimental results illustrate that the proposed detection tree is more efficient than the original one while keeping its detection speed.(3) The face recognition performance is influnced by the facial feature alignment accuracy. In order to describe the facial feature accurately, a face alignment algorithm is proposed for sequential face images. According to the smoothness of face images, a fast fuzzy connected sequential image segmentation algorithm is proposed to segment the face region from complex background. By locating the important features on the face, face images are aligned automatically. Experimental results show that the recognition rates of the aligned faces are higher than that of the non-aligned ones. The proposed relative fuzzy connected interactive segmentation algorithm is robust to different noise models, and it can also improve the speed of segmentation than the original scale based one. The experimental results on medical images show satisfacting results.(4) Multi-view is one of the most challenging factors for face recognition because of the nonlinearity in view subspace. In this paper, tensor decomposition is applied to separate influential factors of multi-view face images for face modeling. Then to handle the nonlinearity in view subspace, a novel data-driven manifold is proposed to model view variation. In this way, a uniform multi-view face model is achieved to deal with the linearity in identity subspace as well as the nonlinearity in view subspace. Meanwhile, a parameter estimation method is developed to solve the view coordinate in the manifold and the identity coefficient automatically. The proposed model yields improved facial recognition rates around 12% against the traditional TensorFace.(5) A new multi-view face recognition method that extends a recently proposed nonlinear tensor decomposition technique is proposed. This technique provides a generative face model that can deal with both the linearity and nonlinearity in multi-view face images. Particularly, the effectiveness of three kinds of view manifold for multi-view face representation is studied in this paper. An EM-like algorithm is developed to estimate the identity and view factors iteratively. The new face generative model can successfully recognize face images captured under unseen views, and the experimental results provide the new method is superior to the traditional TensorFace based algorithm around 19%. This approach is also applicable to other hybrid linear-nonlinear multi-factor object modeling tasks.In this paper, the research scheme and experimental results of multi-view face detection, face alignment and multi-view face recognition in complex background are given. These research achivements enrich the theories and applications of face dection and recogniton tasks. In addition, the proposed example selection algorithm, fast and robust image segmentation algorithm, and multi-view face detection and recognition models are applicable to other related pattern recognition tasks.
【Key words】 Face detection; Multi-view face recognition; Face alignment; Learning from examples; Manifold learning; Tensor analysis; Nonlinear tensor decomposition;