节点文献

三维人脸成像及重建技术综述

3D face imaging and reconstruction technology: a review

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 刘菲张堃博杨青周树波王云龙孙哲南

【Author】 Liu Fei;Zhang Kunbo;Yang Qing;Zhou Shubo;Wang Yunlong;Sun Zhenan;School of Management and Engineering, Capital University of Economics and Business;National Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences;College of Information Science and Technology, Donghua University;

【通讯作者】 孙哲南;

【机构】 首都经济贸易大学管理工程学院中国科学院自动化研究所模式识别实验室东华大学信息科学与技术学院

【摘要】 得益于新型三维视觉测量技术及深度学习模型的飞速发展,三维视觉成为人工智能、虚拟现实等领域的重要支撑技术,三维人脸成像及重建技术取得了突破性进展,不仅能够更好地应对光照、遮挡、表情和姿态等变化,同时增大了伪造攻击难度,大大推动了真实感“虚拟数字人”的重建与渲染,有效提升了人脸系统的安全性。本文对三维人脸成像技术和重建模型进行了全面综述,尤其对基于深度学习的三维人脸重建进行系统深入地分析。首先,对三维人脸成像设备及采集系统进行详细梳理及对比归纳,并介绍了基于新传感技术的人脸成像系统;然后,对基于深度学习的三维人脸重建模型进行系统分析,从输入数据源角度分为基于单目图像、基于多目图像、基于视频和基于语音的三维人脸重建算法4类。通过深入分析,总结三维人脸成像的研究现状及面临的难点与挑战,对未来发展方向及应用进行积极探讨与展望。本文涵盖了近5年经典的三维人脸成像及重建相关的技术与研究,为人脸研究、发展和应用提供了很好的参考。

【Abstract】 As the breakthrough technology of artificial intelligence(AI) in the big data era, deep learning(DL) has prompted the renewed upsurge of face technology. Powered by rapid developments of new technologies, such as threedimensional(3D) vision measurement, image processing chips, and DL models, 3D vision transformed into a key supporting technology in AI, visual reality, etc. The studies and applications of 3D facial imaging and reconstruction technologies have achieved important breakthroughs. 3D face data represent exact multidimensional facial attributes on account of rich visual information, such as texture, shape, space, etc. Moreover, 3D face data shows robust changes in large occlusions, expressions, and poses and increases the difficulty of forgery attack. Therefore, 3D face imaging and reconstruction effectively promote realistic “virtual digital human” reconstruction and rendering. In addition, these processes contribute to the improved security of the face system. In this paper, we comprehensively study the 3D face imaging technology and reconstruction models. The 3D face reconstruction methods based on DL are systematically and deeply analyzed. First, the development and innovation of 3D face imaging devices and capturing systems are discussed through a summary of public 3D face datasets. The devices and systems include consumer imaging devices(such as Kinect) and complex hybrid systems that fuse active and passive 3D imaging technologies to achieve precise geometry and appearance. Moreover, 3D face imaging based on new sensing technologies are introduced. Then, from the perspective of input resources, 3D face reconstruction methods based on DL are categorized into monocular, multiview, video and audio reconstruction methods. 3D face imaging technology introduces public classic 3D face datasets, popular 3D face imaging devices, and capturing systems. Most high-quality 3D face datasets, such as BU-3DFE, FaceScape and FaceVerse, are captured through a large imaging volume with a certain number of high-resolution cameras and controlled lighting conditions. They play key roles in applications of realistic rendering, driven animation, retargeting, etc. On the other hand, novel optical devices and imaging modules with small size and lightweight algorithm must be innovated for tiny AI as intelligent mobile devices. For 3D face reconstruction based on DL, monocular reconstruction has become the most popular technology. The state-of-the-art 3D face reconstruction method is generally self-supervised training on large-scale 2D face databases. The difficulties encountered in 3D face reconstruction include the lack of large-scale 3D face datasets, occlusions and poses of in-the-wild 2D face images, continuous expression deformations, etc. The DL network structure is categorized into general deep convolutional neural network(such as ResNet, U-Net, and Autoencoder), generative adversarial networks(GANs), implicit neural representation(INR)(such as neural radiance field(NeRF) and signed distance functions(SDF)), and Transformer. 3DMM and FLAME are widely used 3D face representation models. The StyleGAN model gives excellent performance in recovering high-quality face texture. INR has achieved remarkable results in 3D scene reconstruction, and the NeRF model plays an important role in the reconstruction of accurate head avatars. The combination of NeRF with GAN shows great potential in the reconstruction of high-fidelity 3D face geometry and realistic rendering appearances. Moreover, the Transformer model, which greatly improves the breakthrough of accuracy and speed, is mainly used in audio-driven 3D face reconstruction. Through in-depth analyses, the research difficulties accompanying 3D face are summarized, and future developments are actively being discussed and explored. Although recent research has made amazing progresses, challenges on how to improve the robustness and generalization to real-world lighting, extreme expressions/poses, and how to effectively disentangle facial attributes(such as identity, expression, albedo, and specular reflectance) and recover accurate detailed geometry of facial motions(such as wrinkles). In this study, we proposed a comprehensive and systematic review and covered classical technologies and studies on 3D face imaging and reconstruction in the last five years to provide a good reference for face studies, developments, and applications.

【基金】 国家自然科学基金项目(61806197,61803372,62071468,U23B2054,62276263)~~
  • 【文献出处】 中国图象图形学报 ,Journal of Image and Graphics , 编辑部邮箱 ,2024年09期
  • 【分类号】TP391.41
  • 【下载频次】130
节点文献中: