节点文献

2024年度三维视觉前沿趋势与十大进展

Research trends and major developments in 3D vision in 2024

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 刘烨斌; 苏昊; 高林; 弋力; 王鹤; 廖依伊; 施柏鑫; 曹炎培; 洪方舟; 董豪; 张举勇; 王鑫涛; 许华哲; 杨蛟龙; 康炳易; 楚梦渝; 孙赫; 陈文拯; 马月昕; 张鸿文; 郭裕兰; 周晓巍; 章国锋; 韩晓光; 戴玉超; 陈宝权;

【Author】 Liu Yebin;Su Hao;Gao Lin;Yi Li;Wang He;Liao Yiyi;Shi Boxin;Cao Yanpei;Hong Fangzhou;Dong Hao;Zhang Juyong;Wang Xintao;Xu Huazhe;Yang Jiaolong;Kang Bingyi;Chu Mengyu;Sun He;Chen Wenzheng;Ma Yuexin;Zhang Hongwen;Guo Yulan;Zhou Xiaowei;Zhang Guofeng;Han Xiaoguang;Dai Yuchao;Chen Baoquan;Department of Automation, Tsinghua University;Department of Computer Science and Engineering, University of California;Institute of Computing Technology, Chinese Academy of Sciences;School of Intelligence Science and Technology, Peking University;College of Computer Science and Technology, Zhejiang University;VAST AI Research;College of Computing and Data Science, Nanyang Technological University;Department of Mathematics, University of Science and Technology of China;Kuaishou Technology;Microsoft Research Asia;Byte Dance Ltd.;School of Information Science and Technology, Shanghai Tech University;School of Artificial Intelligence, Beijing Normal University;School of Electronics and Communication Engineering, Sun Yat-sen University;School of Science and Engineering, The Chinese University of Hong Kong (Shenzhen);School of Electronics and Information, Northwestern Polytechnical University;

【通讯作者】 陈宝权;

【机构】 清华大学自动化系; 加州大学圣迭戈分校计算机科学与工程系; 中国科学院计算技术研究所; 北京大学智能学院; 浙江大学计算机科学与技术学院; VAST哇嘶嗒科技有限公司; 南洋理工大学计算机与数据科学学院; 中国科学技术大学数学科学学院; 快手科技有限公司; 微软亚洲研究院; 字节跳动科技有限公司; 上海科技大学信息科学与技术学院; 北京师范大学人工智能学院; 中山大学电子与通信工程学院; 香港中文大学(深圳)理工学院; 西北工业大学电子信息学院;

【摘要】 三维视觉作为计算机视觉、图形学、人工智能与光学成像的交叉学科,是构建具身通用智能与元宇宙的核心基石。2024年,以神经辐射场和高斯泼溅为代表的可微表征技术持续发展和完善并逐渐突破传统三维重建边界,无论从微观细胞组织到宏观物理天体,还是从静态场景到动态人体,均取得显著的精度提升;在生成式人工智能技术和大模型规模定律(scaling law)的推动下,三维视觉迎来从优化到可泛化前馈生成的范式跃迁,并在可控数字内容生成方向取得重要进展和突破;具身智能持续备受关注,研究者逐渐意识到三维虚拟仿真数据和三维人体运动数据的捕捉和生成,是训练具身智能的核心关键;随着世界模型和空间智能的概念成为科技界热议的焦点,对物理世界进行建模、对空间关系进行理解、对未来状态进行预测成为重要研究方向,而这些都离不开三维视觉技术的支撑;此外,计算成像技术的革新则通过非传统视觉传感器与新型重建算法,突破了传统三维重建的物理限制与性能瓶颈。这些技术突破正推动着三维视觉进入“感知—建模—生成—交互”全链路智能化、规模化学习的新阶段。为促进学术交流,本文分析总结三维视觉领域前沿趋势,并遴选年度十大研究进展,为学术界与产业界提供参考观点。

【Abstract】 As an interdisciplinary field that integrates computer vision, graphics, artificial intelligence(AI), and optical imaging, three-dimensional(3D) vision serves as the foundational cornerstone for building embodied general intelligence and the metaverse ecosystem. In 2024, differentiable representation technologies, exemplified by neural raddiance field(NeRF) and Gaussian splatting, continued to evolve and refine, gradually transcending the boundaries of traditional 3D reconstruction. From microscopic cellular structures to macroscopic physical celestial bodies, and from static scenes to dynamic human bodies, significant improvements in accuracy have been achieved across a spectrum. Propelled by advancements in generative AI technologies and the scaling laws of large models, the field of 3D vision has witnessed a paradigm shift from optimization to generalizable feedforward generation, marking important progress and breakthroughs in the direction of controllable digital content generation. Embodied intelligence remains a focal point of interest, with researchers increasingly recognizing the capture and generation of 3D virtual simulation data and 3D human motion data as central elements of training embodied intelligence. As the concepts of world models and spatial intelligence become hot topics among technology researchers, modeling the physical world, understanding spatial relationships, and predicting future states have emerged as crucial research directions, all of which rely on the support of 3D vision technologies. Furthermore, through nontraditional visual sensors and novel reconstruction algorithms, innovations in computational imaging technology have exceeded the physical limitations and performance bottlenecks of traditional 3D reconstruction. These technological breakthroughs are propelling 3D vision into a new era characterized by an intelligent, large-scale learning process that encompasses the processes of perception, modeling, generation, and interaction. Specifically, the development trends in 3D vision primarily manifested in the following aspects:1)Controllable and physics-aware generation of the visual elements of AI-generated content(AIGC). With the rapid development of AIGC technology, visual content generation has evolved from simple two-dimensional(2D) image creation toward more controllable and physics-aware approaches. This trend requires the combination of physical prior knowledge and multidimensional control parameters, such as 3D viewpoints, lighting conditions, and 3D character motion, to achieve higher-quality content generation. 3D vision technology plays a pivotal role in this process, providing essential spatial-temporal and physical constraints for AIGC. 2)4D spatial intelligence: bridging virtual and physical worlds. 4D(3D space + time dimension) spatial intelligence has emerged as a core technology connecting virtual worlds(e. g., the metaverse) and physical realities(e. g., embodied intelligent robots).This technology focuses on establishing a digital mapping of dynamic physical environments. By leveraging 3D vision and multimodal large model technologies, AI systems can construct 4D spatial models to understand spatial relationships, predict motion trajectories, and simulate future evolutions. In turn, intelligent agents can interactively learn within physical or virtual 4D environments to acquire intelligence. 3)Data-driven embodied intelligence: 3D virtual simulation and human motion capture. The advancement of embodied intelligence relies heavily on high-quality 3D virtual simulation data and the capture/generation of human 3D motion data. These datasets fuel the training of embodied intelligent robots to achieve sophisticated behavioral control. Through high-precision 3D vision technologies, robots can better understand and simulate human actions, resulting in their enhanced intelligence in complex tasks. 4)Differentiable 3D representation and integration with large model technologies. From microscopic cellular structures to indoor environments, human/animal modeling, autonomous driving/city modeling, and even astronomical black hole reconstruction, novel 3D representations, such as NeRF and 3D Gaussian splatting(3DGS), drive performance improvements in scene generation and reconstruction across scales. Their efficiency and flexibility have opened up new possibilities for 3D vision applications. Furthermore, by integrating large-scale 3D data with transformer-based architectures and advanced generative methods, such as diffusion models, fundamental 3D vision tasks are being unified into efficient end-to-end frameworks, resulting in the scaled-up learning of core 3D vision paradigms. The breakthroughs in 3D vision in 2024 have injected new momentum into technological development, with future trends focused on several key directions. Spatiotemporal-consistent world models that integrate 4D spacetime and physical laws provide dynamic prediction and interaction support for complex scenarios, such as autonomous driving and embodied intelligence. It is expected that the deep integration of generative AI and 3D content generation technologies will overcome data bottlenecks, enabling the automated creation of high-fidelity and controllable 3D content.Enhanced cross-modal generalization capabilities will also strengthen the fusion of vision, language, and motion modalities, thus improving the adaptability and robustness of robotic strategies. Physics-driven dynamic reconstruction, combined with physical engines, will achieve high-precision modeling and interactive editing of dynamic scenes, thus advancing digital twins and virtual reality. Meanwhile, 3D imaging technologies will accelerate scientific exploration in astrophysics and cell biology. Efficient real-time processing and lightweight solutions will also boost the reconstruction and rendering efficiency of large-scale dynamic scenes, promoting edge device applications. Furthermore, ethical and privacy protection will emerge as critical concerns, balancing innovation and security through encryption technologies and regulatory frameworks that govern 3D data acquisition and generation. In summary, this article reveals that, in the future, 3D vision will evolve toward an era of more intelligent and universal “spatial intelligence” in which technological breakthroughs will reshape human-computer interaction, scientific exploration, and industrial ecosystems. Therefore, to foster academic exchange, this article analyzed cutting-edge trends in 3D vision and highlights the top ten research breakthroughs of the year, thus offering insights and references for both academia and industry.

【基金】 国家自然科学基金项目(62125107,62306016,62376012,62322210,62271410,U24B20154,62441224,62441223,62136001,62371007)~~
  • 【文献出处】 中国图象图形学报 ,Journal of Image and Graphics , 编辑部邮箱 ,2025年06期
  • 【分类号】TP391.41
  • 【下载频次】116
节点文献中: 

本文链接的文献网络图示:

本文的引文网络