节点文献
基于几何先验知识约束的双目视觉深度估计方法
Binocular Vision Depth Estimation Method Based on Geometric Prior Knowledge Constraints
【摘要】 近年来,随着自动驾驶、机器人导航及三维重建等领域的迅速发展,深度估计技术作为感知环境三维结构的关键手段,受到广泛关注。然而,现有基于监督学习的深度估计方法虽然在特定数据集上表现优异,但其泛化能力较弱,且依赖大规模、高质量的标注数据,这严重限制了其在真实工业场景中的应用。因此,本研究提出一种基于几何先验知识约束的双目视觉深度估计方法。首先,组合残差卷积与上下文编码器,从图像数据中提取多尺度特征。接下来,利用特征金字塔结构捕捉不同尺度匹配信息,并保留图像边缘结构细节。然后,设计多级门控制循环(Gated Recurrent Unit,GRU)单元结合不同尺度特征信息对特征匹配参数进行更新,优化视差匹配结果,实现双目视觉深度估计。特别地,本文构建了一种结合监督信号与物理先验的混合损失函数。该函数在传统监督损失的基础上,引入了源自自监督学习范式的几何约束作为正则化项,具体包括左右视差一致性约束和视差结构一致性约束。其中,左右一致性约束通过强制左右视图预测视差满足几何对应关系,以增强模型的几何理解并缓解遮挡区域的误匹配,而结构一致性约束则通过引导视差图在纹理平坦区域保持平滑、在物体边缘处保持清晰,进而提升深度图的结构完整性与视觉质量,以实现增强双目视觉深度估计模型的泛化能力。为验证所提方法的有效性,本文在KITTI 2015和Middlebury等公开数据集上进行训练与评估,并利用SceneFlow数据集进行跨数据集泛化性能测试。实验结果表明,引入几何先验约束后,基线模型的性能得到稳定提升,在KITTI数据集上,端点误差(End-Point Error,EPE)降低了3%~5%,综合误匹配率(D1-all)降低了5%~8%。同时,在Middlebury数据集上的结果进一步证实了该方法在不同场景下的良好泛化性与鲁棒性。消融实验验证了各模块的贡献,超参数敏感性实验确定了损失函数权重的最优配置。此外,迁移实验表明,本文提出的几何先验约束机制具有良好的可移植性,能够适配于多种主流深度估计网络架构,并普遍带来性能增益。
【Abstract】 In recent years, with the rapid development of fields such as autonomous driving, robot navigation, and 3D reconstruction, depth estimation technology, as a key means of perceiving the three-dimensional structure of the environment, has garnered widespread attention. However, although the existing deep estimation methods based on supervised learning perform well on specific datasets, their generalization ability is weak and they rely on large-scale, high-quality labeled data, which severely limits their application in real industrial scenarios. Hence, this study proposes a binocular vision depth estimation method based on geometric prior knowledge constraints. First, this study combines residual convolution with the context encoder to extract multi-scale features from image data, and utilizes the feature pyramid structure to capture matching information at different scales for retaining the edge structure details of the image. Then, a multi-level gated recurrent unit(GRU) unit is designed to update the feature matching parameters in combination with feature information of different scales, optimize the disparity matching results, and achieve binocular vision depth estimation. Notably, this paper constructs a hybrid loss function that combines supervised signals with physical priors. Based on the traditional supervised loss, this function introduces geometric constraints derived from the self-supervised learning paradigm as regularization terms, specifically including the left-right disparity consistency constraint and the disparity structure consistency constraint. The left-right consistency constraint enforces geometric correspondence between the predicted disparities of the left and right views, enhancing the model geometric understanding and mitigating mismatches in occluded areas. The structural consistency constraint guides the disparity map to remain smooth in texture-flat regions and sharp at object edges, thereby improving the structural integrity and visual quality of the depth map, ultimately enhancing the generalization capability of the binocular vision depth estimation model. To verify the effectiveness of the proposed method, this paper conducts training and evaluation on public datasets such as KITTI 2015 and Middlebury, and uses the SceneFlow dataset for cross-dataset generalization performance. Experimental results show that after introducing geometric prior constraints, the baseline model’s performance is consistently improved: on the KITTI dataset, the endpoint error(EPE) is reduced by 3% to 5%, and the overall mismatch rate(D1-all) is reduced by 5% to 8%. Simultaneously, results on the Middlebury dataset further confirm the method’s good generalization and robustness across different scenarios. Ablation experiments verify the contributions of each module, while hyperparameter sensitivity experiments determine the optimal configuration for the loss function weights. Additionally, transfer experiments demonstrate that the proposed geometric prior constraint mechanism exhibits good portability, adapting to various mainstream depth estimation network architectures and generally providing performance gains.
【Key words】 depth estimation; stereo matching; prior knowledge; deep learning; geometric constraints; hybrid supervised learning;
- 【文献出处】 电子学报 ,Acta Electronica Sinica , 编辑部邮箱 ,2026年01期
- 【分类号】TP391.41;TP18
- 【下载频次】12