节点文献

基于深度学习的双模态生物特征识别研究

Research on Bi-mode Biometrics Based on Deep Learning

【作者】 蒋浩

【导师】 邹采荣;

【作者基本信息】 东南大学 , 信息与通信工程, 2017, 硕士

【摘要】 鉴于生物特征具有优秀的独立区分特性,生物特征识别技术几乎涉及到所有区分人的相关领域。指纹、虹膜、人脸、声纹等生物特征已经广泛使用到公安部门破案侦查,移动设备解锁,目标跟踪等领域。随着电子设备使用的范围越来越广,使用的频率越来越高。具有优秀识别率的生物特征识别技术才能保证这些领域的长久发展。由于现实生活中使用生物特征识别的环境条件的差异。单一的生物特征识别因为其自身条件的局限性,导致无论哪一种生物特征有会有其应用的局限性。例如指纹识别使用的前提是设备与目标人物必须有肢体的接触,视频监控等远距离设备无法使用指纹识别技术;在光线不足或是摄像头没有正对目标人物脸部的情况下人脸识别技术的识别率将会急剧下降;同样,虹膜识别技术也要保证目标人物眼部靠近传感器,才能实现此生物特征的后续识别过程。多生物特征融合识别技术可以很好的解决这个问题。多模态生物特征融合识别技术可以根据特征的选取和融合方法提高识别设备的准确性,普适性和鲁棒性。本文的主要研究内容和实验成果如下:1.研究了基于深度卷积网络(Convolutional Neural Network,CNN)和的人脸识别,创新的提出了基于Vgg_Face改进模型的人脸识别算法。该方法结合深度卷积网络模型的核心思想,在Vgg_Face模型的最后一层添加了全连接层降低人脸特征维度。在Vgg_Face模型的基础上微调(fine-turning)新的模型。使用CASIA-Webface人脸数据库训练网络模型。2.首先研究了说话人的感知预测系数(Perceptual Linear Predictive,PLP),然后推导了说话人的I-Vector特征提取方法,I-Vector特征反映说话人的差异,并且具有优秀的跨信道性能,是目前说话人识别主流的识别特征。然后将I-Vector和PLP特征融合送入深度置信网络(Deep Belief Network,DBN)训练,生成说话人识别模型。3.创新的提出了人脸模型和说话人特征融合识别方法。结合TED-LIUM语音库和CASIA-WebFace人脸库,将不同的人脸和语音随机组合成新的人脸—说话人综合库。将该库提取得到的融合特征去训练DBN。最终得到的模型识别目标人物,其识别率比单独使用人脸特征或者说话人特征的识别率有一定得提高。

【Abstract】 In view of the fact that biological characteristics have excellent independent distinguishing characteristics,Biometric identification technology involves almost all the relevant areas of human distinction.Fingerprints,iris,face,voiceprint and other biological features have been widely used in the public security departments to detect detection,mobile equipment unlock,target tracking and other fields.With the use of electronic devices more and more widely and the frequency is getting higher and higher.Only the Biometric identification technology with excellent recognition rate can guarantee the long-term development of these fields.Due to differences in environmental conditions using biometrics recognition in real life.Single biometrics because of its own limitations,lead to any kind of biometrics there will be limitations of its application.For example,the premise of the use of fingerprint identification equipment is the target people must have physical contact,video surveillance and other remote devices can not use fingerprint recognition technology;The recognition rate of face recognition technology will drop sharply if there is not enough light or if the camera is not facing the target person’s face;Similarly,the iris recognition technology must ensure that the target person’s eye close to the sensor in order to achieve this biological characteristics of the follow-up process.Multi-biometric fusion identification technology can be a good solution to this problem.Multi-modal biometric fusion recognition technology can improve the accuracy,universality and robustness of the identification device according to the feature selection and fusion method.The main contents and experimental results of this paper are as follows:1.Based on the Convolutional Neural Network(CNN)and the face recognition,a new face recognition algorithm based on Vgg Face improved model is proposed.This method combines the core idea of the deep convolution network model.In the last layer of the Vgg_Face model,the full join layer is added to reduce the face feature dimension.Fine-turning a new model based on the Vgg_Face model.Use the CASIA-Webface face database to train the network model.2.Firstly,the Perceptual Linear Predictive(PLP)of the speaker is studied,and then the speaker’s I-Vector feature extraction method is deduced.The I-Vector feature reflects the speaker’s differences and has excellent cross-channel performance.It’s the mainstream identification characteristics of Speaker identify.And then sent the I-Vector and PLP features into the deep confidence network(Deep Belief Network,DBN)to training speaker recognition model.3 The fusion method of face model and speaker character fusion is put forward.Combined with TED-LIUM speech database and CASIA-WebFace face database,different face and voice randomly will be combined into a new face-speaker integrated library.The dataset is extracted from the fusion feature to train the DBN.The resulting model identifies the target person whose recognition rate is higher than the recognition rate of the individual face feature or speaker feature alone.

  • 【网络出版投稿人】 东南大学
  • 【网络出版年期】2018年 04期
  • 【分类号】TP391.41;TP181
  • 【被引频次】4
  • 【下载频次】255
  • 攻读期成果
节点文献中: 

本文链接的文献网络图示:

本文的引文网络