节点文献

基于语义信息的无人车激光-惯性同步定位与地图构建方法

LiDAR-Inertial Simultaneous Localization and Mapping Method for Unmanned Vehicles Based on Semantic Information

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 罗娟赵瑞祺张传伟杨佳佳

【Author】 Luo Juan;Zhao Ruiqi;Zhang Chuanwei;Yang Jiajia;College of Railway Transport, Shaanxi Communications Vocational and Technical College;College of Mechanical Engineering, Xi’an University of Science and Technology;

【通讯作者】 张传伟;

【机构】 陕西交通职业技术学院铁道运输学院西安科技大学机械工程学院

【摘要】 针对无人驾驶场景中动态物体干扰引发定位精度下降的问题,提出一种融合激光雷达、惯性测量单元(IMU)与相机的语义同步定位与地图构建框架(SeLI-SLAM)。在前端,利用IMU预积分进行运动估计以消除点云运动畸变,并设计了一种改进的DeepLabv3+语义分割网络。该网络采用轻量化主干网络、引入注意力机制并优化解码器结构,从而提升了分割精度与效率,为点云赋予了更可靠的语义标签。结合对应实例簇的相似度度量,有效检测并剔除动态物体。在后端,提出基于语义约束的回环检测与图优化机制:首先利用几何阈值筛选候选帧,再结合关键语义特征匹配与点云重匹配进行验证,最后通过图优化降低了全局累积误差,提升了轨迹一致性。在公开KITTI数据集及自建HT数据集上进行的实验结果表明,SeLI-SLAM在大规模城市道路环境中能够实现更高的定位精度和更优的建图效果,验证了该方法的有效性与鲁棒性。

【Abstract】 Objective Simultaneous localization and mapping(SLAM) is a core technology for autonomous vehicle navigation. It provides reliable perception of the surrounding environment and self-location information,serving as a critical foundation for perception and decision-making in autonomous driving systems. Traditional SLAM techniques perform well in static environments. However,in real-world complex traffic scenes,dynamic objects such as pedestrians and vehicles are prevalent. These objects often cause point cloud smearing,which degrades positioning accuracy and introduces errors into the map. Furthermore,to achieve semantic-based navigation tasks like “park the car next to a tree”,maps built solely by traditional geometric SLAM are insufficient. Therefore,it is necessary to incorporate semantic information into spatial representations to construct higher-quality semantic maps.Methods This paper presents a semantic SLAM framework named SeLI-SLAM,which tightly couples LiDAR,inertial measurement unit(IMU),and camera data. At the front end,the system utilizes IMU pre-integration for motion estimation to eliminate point cloud motion distortion. An improved DeepLabv3+ semantic segmentation network is designed,which employs a lightweight backbone network,incorporates an attention mechanism,and optimizes the decoder structure. This improves segmentation accuracy and efficiency,thereby assigning more reliable semantic labels to point clouds. Subsequently,dynamic objects are effectively detected and removed by combining a similarity measure for corresponding instance clusters. At the back end,the framework proposes a loop closure detection and graph optimization mechanism based on semantic constraints. Candidate frames are first screened using geometric thresholds and then verified through key semantic feature matching and point cloud re-matching.Finally,graph optimization is applied to reduce global cumulative error and improve trajectory consistency.Results and Discussions The improved DeepLabv3+ semantic segmentation model proposed in this paper demonstrates significant improvements in both accuracy and efficiency compared to the original DeepLabv3+. As shown in Table 2,the mean intersection over union(mIoU) and mean pixel accuracy(mPA) increased to 82.12% and 85.38%,representing gains of 1.96 and 1.93 percentage points,respectively. The inference speed also improved substantially,reaching 23.81 frames per second. Visual comparisons in Fig. 14 indicate superior performance in traffic sign recognition and clearer boundaries for vehicle and pedestrian segmentation. The proposed semantic-based loop closure detection method was evaluated using precision and recall metrics. As detailed in Table 5,the method achieved 100% precision across all test sequences(KITTI-00,HT-01 to HT-04) and high recall rates(92%-100%),demonstrating its ability to accurately identify loops and capture most true positives. The overall SeLI-SLAM framework was compared with four classical SLAM methods(A-LOAM,LeGO-LOAM,LIO-SAM,and SUMA++) on KITTI datasets(sequences 00 and 05) and the self-collected HT dataset. The results show that the proposed method achieved the lowest root mean square error(RMSE) across all six sequences,with an average RMSE of 3.635 m. Trajectory comparison plots visually confirm the higher conformity of our method’s trajectory to the ground truth,especially the well-closed loop in the HT-01 sequence. The constructed semantic map for HT-01 clearly depicts environmental elements such as parked cars,trees,and traffic signs,validating the framework’s capability to generate high-quality 3D semantic maps.Conclusions This paper proposes a reliable and high-precision semantic SLAM framework suitable for urban road environments. The system integrates camera,LiDAR,and IMU data to construct a multi-sensor fusion odometer enriched with semantic information. An improved image semantic segmentation algorithm with higher accuracy is employed to achieve point cloud semantic segmentation and effectively eliminate dynamic obstacles. Building upon this,the system performs robust loop closure detection using geometric constraints and key semantic features,recovers a globally consistent trajectory via graph optimization,and ultimately completes the construction of a 3D semantic map. To validate the practical performance of the proposed method,a systematic evaluation was conducted on the KITTI dataset and self-collected urban road scene data,with comparisons made against current mainstream SLAM methods(A-LOAM,LeGO-LOAM,LIO-SAM,and SUMA++). The experimental results demonstrate that the proposed method exhibits higher positioning accuracy and stronger environmental adaptability in large-scale urban road environments,effectively proving the feasibility and practicality of the proposed semantic SLAM framework.

【基金】 国家自然科学基金面上项目(51974229);陕西省科技创新团队(2021TD-27)
  • 【文献出处】 激光与光电子学进展 ,Laser & Optoelectronics Progress , 编辑部邮箱 ,2026年06期
  • 【分类号】TP391.41;TN958.98;U463.6
  • 【下载频次】69
节点文献中: 

本文链接的文献网络图示:

本文的引文网络