节点文献

基于双视点采集的大视角三维光场重建算法

Large-Angle Three-Dimensional Light Field Reconstruction Algorithm Based on Dual-Viewpoint Acquisition

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 郭蓉璇于迅博高鑫李宁驰王姝瑶颜玢玢桑新柱

【Author】 Guo Rongxuan;Yu Xunbo;Gao Xin;Li Ningchi;Wang Shuyao;Yan Binbin;Sang Xinzhu;School of Electronic Engineering, Beijing University of Posts and Telecommunications;

【通讯作者】 桑新柱;

【机构】 北京邮电大学电子工程学院

【摘要】 为解决基于稀疏视点的光场重建面临的遮挡信息缺失问题,提出一种基于双视点采集的大视角密集视点三维(3D)光场重建算法,利用双视点图像间的空间关联性,引入扩散模型对遮挡区域进行补全,结合基于3D高斯表示的体渲染方法实现具有几何一致性的多视角图像生成。与传统插值或几何重建方法相比,本文方法利用人工智能模型实现了对遮挡区域从无到有的3D内容生成,不依赖密集视角信息即可对缺失结构进行推理,为基于稀疏输入的3D光场内容获取提供了一种新的技术路径。实验结果表明,在仅输入两张视角图像的条件下,本文方法可有效重建覆盖水平视角60°、共计96个视点的高保真3D光场图像,本文方法得到的重建图像的峰值信噪比(PSNR)和结构相似性(SSIM)优于现有方法,具备良好的实用性和推广价值。

【Abstract】 Objective Three-dimensional(3D) light field display technology has emerged as a promising solution for next-generation immersive visual experiences. However, conventional acquisition methods require a large number of densely sampled viewpoints to capture highfidelity light field data, which imposes significant hardware costs and limits practical applicability. In particular, in scenarios involving large baseline configurations or wide-angle displays, occlusion between foreground and background objects often leads to severe information loss, posing a major challenge for accurate 3D scene reconstruction. To overcome these limitations, this study proposes a novel 3D light field reconstruction algorithm based on dual-viewpoint acquisition, which aims to synthesize dense-view images across a wide angular range using only two input views. This method targets the occlusion completion problem as its core innovation, leveraging recent advances in diffusion models to effectively restore missing scene details while maintaining spatial consistency.Methods The proposed framework comprises two key stages: a stereo diffusion-based occlusion completion module and a 3D Gaussian-based volumetric rendering module. In the first stage, a pair of sparsely sampled images(with angular separations up to 60°) is used to generate an initial point cloud through uncalibrated dense stereo reconstruction. A latent diffusion model is then employed to complete missing regions in the point cloud rendering, particularly targeting occluded areas where traditional stereo matching fails due to lack of correspondence. These occluded regions refer to spatial locations in the 3D scene where the reconstructed point cloud becomes sparse, discontinuous, or missing due to visibility constraints between the two input views. Such regions typically lack reliable depth and geometry information, rendering them difficult to recover using conventional stereo methods.To address this, the proposed diffusion model—designed based on an improved latent denoising diffusion probabilistic model(DDPM) —learns to restore these geometrically incomplete regions by modeling the residual mapping between the rendered point cloud images and the corresponding ground truth images. The process is conducted in a low-dimensional latent space, where the model progressively refines noisy latent representations toward photo-realistic outputs, ensuring consistency with surrounding context and structural coherence. Notably, this training process does not require explicit annotations of occluded areas; instead, the model automatically learns to infer and complete missing content through its denoising objective. The input to the diffusion model consists of 24 point cloud rendered images(each with an 8-channel feature map) along with a reference image encoded as conditional guidance.In the second stage, the completed intermediate views are used to construct a 3D Gaussian splatting(3D-GS) representation, enabling volumetric rendering across a continuous viewpoint space. This combination allows the proposed method to synthesize geometrically consistent dense-viewpoint images suitable for high-fidelity 3D display applications.Results and Discussions To evaluate the effectiveness of the proposed method, we conducted extensive experiments on a largeangle light field dataset primarily composed of human portrait images. Each experimental group involved reconstructing 96 dense viewpoints using only two input images captured at angular separations ranging from 10° to 60°. The performance was compared against three representative baselines: 1) an optical flow-based interpolation method, 2) a 3D-GS method with 9 evenly spaced input views, and 3) the standalone output of the diffusion module without subsequent Gaussian rendering.Quantitative evaluation using peak signal-to-noise ratio(PSNR) and structural similarity index(SSIM) was conducted to assess the visual fidelity of the synthesized views. Under the most challenging setting with a 60-degree angular span, the proposed method achieved a PSNR of 32.02 dB and an SSIM of 0.9482, significantly outperforming the optical flow interpolation method(22.85 dB, 0.8636) and the 3D Gaussian splatting baseline(21.98 dB, 0.8414). Even when only the stereo diffusion completion module was used without volumetric rendering, the method maintained a strong performance of 26.92 dB PSNR and 0.9075 SSIM. Similar performance advantages were observed across other angular spans, demonstrating the robustness and generalizability of the proposed framework over wide viewing ranges.To further validate the effectiveness of our method in handling occlusion, we introduced an occlusion-aware local quality analysis. Occlusion masks were computed using bidirectional optical flow from the RAFT model based on forward-backward consistency. Regions with large flow inconsistency were marked as occluded. These masks were overlaid on the input images for visualization and used to isolate foreground occluded regions for targeted evaluation. The results show that our method achieves superior reconstruction fidelity within occluded areas, with clear advantages in preserving structural continuity and recovering texture details, especially in complex scenes where other methods suffer from artifacts or geometry distortions.Finally, the reconstructed light field images were loaded onto a 65-inch 3D light field display for visualization. The synthesized results exhibit smooth disparity transitions and correct spatial perception, confirming the method’s effectiveness in large-angle 3D light field display scenarios.Conclusions This study presents a dual-viewpoint-based 3D light field reconstruction algorithm capable of generating dense, highquality light field images over a wide angular span with minimal input data. By integrating a stereo diffusion model for occlusion completion and a 3D Gaussian rendering framework for spatially consistent synthesis, the proposed method effectively addresses the challenges of sparse-viewpoint acquisition and occlusion-induced data loss. Experimental results demonstrate that our method not only surpasses existing techniques in global reconstruction quality but also exhibits notable advantages in accurately recovering occluded regions. While the current implementation has not yet achieved real-time performance(with a generation time of approximately 2-3 minutes for 96 views), future work will explore network optimization and acceleration strategies to enhance inference speed. The method’s high fidelity, low input requirement, and ability to handle large viewing angles make it a promising solution for practical 3D light field display systems.

【基金】 国家重点研发计划(2023YFB2806901);信息光子学与光通信全国重点实验室(北京邮电大学)基金(2024ZZ02);国家自然科学基金(62175015)
  • 【文献出处】 光学学报 ,Acta Optica Sinica , 编辑部邮箱 ,2026年05期
  • 【分类号】TP391.41;O439
  • 【下载频次】13
节点文献中: 

本文链接的文献网络图示:

本文的引文网络