节点文献

三维视觉前沿进展

Recent progress in 3D vision

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 龙霄潇程新景朱昊张朋举刘浩敏李俊郑林涛胡庆拥刘浩曹汛杨睿刚吴毅红章国锋刘烨斌徐凯郭裕兰陈宝权

【Author】 Long Xiaoxiao;Cheng Xinjing;Zhu Hao;Zhang Pengju;Liu Haomin;Li Jun;Zheng Lintao;Hu Qingyong;Liu Hao;Cao Xun;Yang Ruigang;Wu Yihong;Zhang Guofeng;Liu Yebin;Xu Kai;Guo Yulan;Chen Baoquan;The University of Hong Kong;Jiluo Technology ( Shanghai) Co.,Ltd.;Nanjing University;Institute of Automation,Chinese Academy of Sciences;School of Artificial Intelligence,University of Chinese Academy of Sciences;SenseTime Research Institute;National University of Defense Technology;University of Oxford;Sun Yat-sen University;Zhejiang University;Tsinghua University;Peking University;

【通讯作者】 陈宝权;

【机构】 香港大学际络科技(上海)有限公司南京大学中国科学院自动化研究所中国科学院大学人工智能学院商汤研究院国防科技大学牛津大学中山大学浙江大学清华大学北京大学

【摘要】 在自动驾驶、机器人、数字城市以及虚拟/混合现实等应用的驱动下,三维视觉得到了广泛的关注。三维视觉研究主要围绕深度图像获取、视觉定位与制图、三维建模及三维理解等任务而展开。本文围绕上述三维视觉任务,对国内外研究进展进行了综合评述和对比分析。首先,针对深度图像获取任务,从非端到端立体匹配、端到端立体匹配及无监督立体匹配3个方面对立体匹配研究进展进行了回顾,从深度回归网络和深度补全网络两个方面对单目深度估计研究进展进行了回顾。其次,针对视觉定位与制图任务,从端到端视觉定位和非端到端视觉定位两个方面对大场景下的视觉定位研究进展进行了回顾,并从视觉同步定位与地图构建和融合其他传感器的同步定位与地图构建两个方面对同步定位与地图构建的研究进展进行了回顾。再次,针对三维建模任务,从深度三维表征学习、深度三维生成模型、结构化表征学习与生成模型以及基于深度学习的三维重建等4个方面对三维几何建模研究进展进行了回顾,并从多视RGB重建、单深度相机和多深度相机方法以及单视图RGB方法等3个方面对人体动态建模研究进展进行了回顾。最后,针对三维理解任务,从点云语义分割和点云实例分割两个方面对点云语义理解研究进展进行了回顾。在此基础上,给出了三维视觉研究的未来发展趋势,旨在为相关研究者提供参考。

【Abstract】 3 D vision has numerous applications in various areas, such as autonomous vehicles, robotics, digital city, virtual/mixed reality, human-machine interaction, entertainment, and sports. It covers a broad variety of research topics, ranging from 3 D data acquisition, 3 D modeling, shape analysis, rendering, to interaction. With the rapid development of 3 D acquisition sensors(such as low-cost LiDARs, depth cameras, and 3 D scanners), 3 D data become even more accessible and available. Moreover, the advances in deep learning techniques further boost the development of 3 D vision, with a large number of algorithms being proposed recently. We provide a comprehensive review on progress of 3 D vision algorithms in recent few years, mostly in the last year. This survey covers seven different topics, including stereo matching, monocular depth estimation, visual localization in large-scale scenes, simultaneous localization and mapping(SLAM), 3 D geometric modeling, dynamic human modeling, and point cloud understanding. Although several surveys are now available in the area of 3 D vision, this survey is different from few aspects. First, this study covers a wide range of topics in 3 D vision and can therefore benefit a broad research community. On the contrary, most existing works mainly focus on a specific topic, such as depth estimation or point cloud learning. Second, this study mainly focuses on the progress in very recent years. Therefore, it can provide the readers with up-to-date information. Third, this paper presents a direct comparison between the progresses in China and abroad. The recent progress in depth image acquisition, including stereo matching and monocular depth estimation, is initially reviewed. The stereo matching algorithms are divided into non-end-to-end stereo matching, end-to-end stereo matching, and unsupervised stereo matching algorithms. The monocular depth estimation algorithms are categorized into depth regression networks and depth completion networks. The depth regression networks are further divided into encoder-decoder networks and composite networks. Then, the recent progress in visual localization, including visual localization in large-scale scenes and SLAM is reviewed. The visual localization algorithms for large-scale scenes are divided into end-to-end and non-end-to-end algorithms, and these non-end-to-end algorithms are further categorized into deep learning-based feature description algorithms, 2 D image retrieval-based visual localization algorithms, 2 D-3 D matching-based visual localization algorithms, and visual localization algorithms based on the fusion of 2 D image retrieval and 2 D-3 D matching. SLAM algorithms are divided into visual SLAM algorithms and multisensor fusion based SLAM algorithms. The recent progress in 3 D modeling and understanding, including 3 D geometric modeling, dynamic human modeling, and point cloud understanding is further reviewed. 3 D geometric modeling algorithms consist of several components, including deep 3 D representation learning, deep 3 D generative models, structured representation learning and generative models, and deep learning-based 3 D modeling. Dynamic human modeling algorithms are divided into multiview RGB modeling algorithms, single-depth camera-based and multiple-depth camera-based algorithms, and single-view RGB modeling methods. Point cloud understanding algorithms are further categorized into semantic segmentation methods and instance segmentation methods for point clouds. The paper is organized as follows. In Section 1, we present the progress in 3 D vision outside China. In Section 2, we introduce the progress of 3 D vision in China. In Section 3, the 3 D vision techniques developed in China and abroad are compared and analyzed. In Section 4, we point out several future research directions in the area.

  • 【文献出处】 中国图象图形学报 ,Journal of Image and Graphics , 编辑部邮箱 ,2021年06期
  • 【分类号】TP391.41
  • 【被引频次】19
  • 【下载频次】3395
节点文献中: 

本文链接的文献网络图示:

本文的引文网络