节点文献
多视图下激光雷达点云目标检测的研究
Research of LiDAR Object Detection Based on Multi-view Data Representation
【作者】 赵菁菁;
【导师】 马利庄;
【作者基本信息】 华东师范大学 , 计算机科学与技术, 2022, 硕士
【摘要】 目标检测是计算机视觉中的一个重要的基础任务。激光雷达传感器能快速、准确地主动测量场景中目标的三维坐标,被大量运用于自动驾驶、机器人、增强/虚拟现实等场景。基于特征表达方式,目前主流的点云目标检测方法可以分成四类:一是基于点特征的方法,二是基于鸟瞰图的方法,三是基于前视图的方法,还有它们的混合方法。点级别的算子能有效表示特征,但计算消耗较大,很难应用于大规模点云数据集。基于鸟瞰图的方法精度较高,但将空间分割成体素时会产生量化误差。前视图的方法计算高效,然而对于目标的深度估计不准。本文设计了新视图下的目标检测方法,并进一步探索如何利用多视图融合兼顾模型的效果和效率,还建模了跨传感器的点云仿真器用于缩小模型性能差距。首先,本文提出了一种新的基于柱状视图的三维目标检测方法。该方法将性能卓越的柱状分割器作为编码器,将上下文结构感知的非对称卷积和扩张稀疏卷积应用于骨干网络和区域提取网络,有效提取外观信息和位置信息,解决柱面视图下物体形变严重和尺度不一致的问题。为了将三维目标框适配到柱状坐标系,该方法进一步设计了有距离感知的柱状中心检测头,实现了定位的有效提升。通过实验表明,该方法对于小目标和远距离的目标有大幅的精度提升。其次,本文提出了一种基于多视图融合的实时三维目标检测方法,该方法简单、高效、且精确度高。该方法结合了前视图和鸟瞰图的优势,在前视图上做分割,并将分割后的前景点和特征投到三维点云进行检测。它是一个全稀疏网络,因此能够高效地聚合多帧数据,学习时序特征。相比于同样精度的网络,该方法的耗时降低了 5倍。最后,为了解决目标检测方法对不同线束数量的机械激光雷达传感器难以迁移的问题,本文提出一种模式感知的激光雷达模拟器,用于简化传感器光线追踪的计算并加速数据生成。本文用这种模拟器生成了一个跨传感器的激光雷达点云目标检测数据集,其中包含从虚拟现实激光雷达模拟器捕获的6组不同传感器但具有相同对应场景的大规模注释激光雷达点云。并根据域差异的特性,提出一种简单有效的最近邻下采样NNDS方法,能减少不同传感器数据的模型性能域差异。
【Abstract】 Object detection is an important task in computer vision.Lidar sensors can quickly and accurately measure the three-dimensional coordinates of targets in the scene actively,and are widely used in autonomous driving,robotics,augmented/virtual reality and other scenarios.Based on the feature expression,the current mainstream point cloud target detection methods can be divided into four categories:one kind is methods based on point features,the other kind is methods based on the bird’s eye view,the third is methods based on the front view,and their hybrid methods.Point-level operators can effectively represent features,but they are computationally expensive and difficult to apply to large-scale point cloud datasets.Bird’s-eye-view-based methods are more accurate,but incur quantization errors when dividing the space into voxels.The front-view methods is computationally efficient,but depth estimation of targets is inaccurate.This paper designs a target detection method under a new view,and further explores how to use multi-view fusion to take into account the effect and efficiency of the model.A cross-sensor point cloud simulator is proposed to narrow the model performance gap.First,this paper proposes a new 3D object detection method based on cylindrical view.The method uses a cylinder feature extraction network with outstanding performance as an encoder,and applies context-structure-aware asymmetric convolution and dilated sparse convolution to the backbone network and region extraction network to effectively extract appearance information and position information,and solve objects under cylindrical view.Serious deformation and scale inconsistency.In order to adapt the threedimensional target frame to the cylindrical coordinate system,the method further designs a cylindrical center detection head with distance perception,which can effectively improve the positioning.Experiments show that this method can greatly improve the accuracy of small targets and long-distance targets.Secondly,this paper proposes a real-time multi-view fusion 3D object detection method,which is simple,efficient and accurate.This method combines the advantages of front view and bird’s-eye view,performs segmentation on the front view,and casts the segmented foreground points and features into a 3D point cloud for detection.It is a fully sparse network,so it can efficiently aggregate multiple frames of data and learn temporal features.Compared with the same precision network,the time-consuming of this method is reduced by 5 times.Finally,to solve the problem that detection models on source domain performs bad on target domain where mechanical lidar sensors are of different number of beams,this paper proposes a mode-aware lidar simulator to simplify the calculation of sensor ray tracing and accelerate data generation.This simulator is used to generate a cross-sensor lidar point cloud object detection dataset which contains large-scale annotated lidar point clouds captured from a mixed reality lidar simulator for datasets of 6 different sensors but with the same scenes and consistent corresponding annotations.
【Key words】 point cloud; 3d object detection; multi-view; real-time; cross-sensor;