节点文献

神经辐射场稠密点云SLAM技术研究

Research on SLAM Technology for Dense Point Cloud in Neural Radiation Field

【作者】 李伟;

【导师】 何元烈; 邓超;

【作者基本信息】 广东工业大学 , 软件工程(专业学位), 2025, 硕士

【摘要】 视觉同步定位与建图(Visual Simultaneous Localization and Mapping,VSLAM)因为丰富的视觉信息,在自主机器人领域被广大研究学者和工程技术人员所关注。近年来,神经辐射场(Neural Radiance Field,Ne RF)作为一种隐式场景表示的新视觉合成技术,在计算机视觉和图形领域取得了较为显著的进展。在结合神经辐射场和视觉SLAM上,采用RGB-D相机作为图像的输入,由于物体和相机的自身特性,会存在像素深度值为0的深度缺失情况,同时也会存在图像噪声。深度噪声和深度缺失会增加SLAM系统额外负担,并且会影响到相机位姿的估计;另外神经辐射场依赖于相机的初始位姿,而当前的模型在初始位姿估计上采用的恒速模型,在室内场景出现的非匀速直线运动上,位姿估计存在较大误差。针对上述问题,本文主要研究工作和贡献如下:(1)针对输入的深度图像大面积缺失,从而引起神经辐射场采样失效的问题,本文提出了采用图像滤波器进行补全模型:根据RGB图像引导深度图进行保护边缘的的滤波,再通过深度缺失部分掩膜的面积,进行不同滤波器方案的选择,面积大于阈值的子图进行降采样,降低计算开销,然后进行多方向的滤波方案;面积低于阈值的,通过阿尔法均值滤波。(2)滤波器在高纹理区域,因不能保留物体边缘,故采用训练好的3D U-net网络模型进行深度缺失部分的补全。针对已训练好的网络模型泛化性较差的问题,本文提出了通过输入数据修复增强、输入数据动态归一化和测试时自适应的方法来提高模型的泛化性。再通过融合滤波器预测的深度图补全,最后实现完整的深度图缺失补全。(3)当前模型在估计每帧图像的相机初始位姿时,所采用的是恒速模型。在室内场景里,物体的运动呈现非匀速直线运动,导致该模型存在明显的局限性。本文提出了采用传统视觉SLAM上通过特征点匹配进行帧间相对位姿的估计,并降低深度图像的损失权重,提高彩色图像的损失权重,通过MLP网络对位姿进行反向优化,提高相机位姿的估计精度。

【Abstract】 Visual synchronous localization and mapping(VSLAM)is concerned by many re-searchers and engineers in the field of autonomous robots because of its rich visual information.In recent years,as a new visual synthesis technology for implicit scene representation,neural radiation field(nerf)has made significant progress in the field of computer vision and graphics.In the combination of neural radiation field and visual slam,the rgb-d camera is used as the input of the image.Due to the characteristics of the object and camera,there will be depth loss with pixel depth value of 0,and there will also be image noise.Depth noise and depth missing will increase the additional burden of slam system,and will affect the estimation of camera pose;In addition,the neural radiation field depends on the initial pose of the camera,while the current model uses the constant speed model in the initial pose estimation,which has a large error in the non-uniform linear motion of the indoor scene.In view of the above problems,the main research work and contributions of this thesis are as follows:(1)To solve the problem that the input depth image is missing in a large area,which causes the sampling failure of neural radiation field,this thesis proposes an image filter to complete the model:filter the edge protection according to the RGB image guided depth map,and then select different filter schemes through the area of the mask of the depth missing part.Subsam-ples whose area is greater than the threshold are downsampled to reduce the computational overhead,and then carry out multi-directional filtering schemes;If the area is lower than the threshold,it shall be filtered by alpha mean.(2)The filter is in the high texture area,so the trained 3D U-net network model is used to complete the depth missing part because the edge of the object cannot be retained.To solve the problem of poor generalization of the trained network model,this thesis proposes methods to improve the generalization of the model through input data repair and enhancement,dynamic normalization of input data,and adaptive testing.Then complete the missing depth map by fusing the predicted depth map of the filter.(3)The current model uses a constant speed model to estimate the initial camera pose of each image.In indoor scenes,the motion of objects presents non-uniform linear motion,which leads to the obvious limitations of this model.This thesis proposes to estimate the relative pose between frames through feature point matching on the traditional visual SLAM,reduce the loss weight of the depth image,improve the loss weight of the color image,and reverse optimize the pose through the MLP network to improve the accuracy of camera pose estimation.

  • 【分类号】TP391.41;TP242
节点文献中: 

本文链接的文献网络图示:

本文的引文网络