节点文献
面向智能驾驶的交通场景理解方法研究
Research on the Traffic Scene Understanding Method for Intelligent Driving
【作者】 李威;
【导师】 曲昭伟;
【作者基本信息】 吉林大学 , 交通信息工程及控制, 2021, 博士
【摘要】 智能驾驶作为战略性新兴产业的重要组成部分,是互联网时代向人工智能时代发展的过程中,世界新一轮经济与科技发展的战略制高点之一,是未来解决交通拥堵的重要枝术,能够大大提升生产效率和交通效率。发展智能驾驶,对于促进国家科技、经济、社会、生活、安全及综合国力有着重大的意义。本文以面向智能驾驶的交通场景理解方法为研究对象,围绕交通标志识别、行人检测和交通场景语义理解在实际应用中存在的问题进行了深入研究,主要分为以下四部分内容。(1)针对交通标志检测效果易受光线、天气和运动模糊影响的问题,提出一种基于颜色概率模型的交通标志检测方法。该方法根据交通标志中特有的颜色建立颜色概率模型,提高了交通标志与其背景的对比度,通过最大稳定极值区域方法对交通标志感兴趣区域进行提取,使用图像分类方法对交通标志感兴趣区域进行判定,确定交通标志所在区域。该方法在德国GTSDB通用数据集上进行了实验验证,因减少了交通标志感兴趣区域数量,缩减了搜索空间,从而能够提高检测效率和性能。(2)针对真实场景中提取的交通标志图像清晰度和拍摄角度存在多样性的特点,提出一种基于融合特征的交通标志分类方法,将基于注意力的Fish Net深度特征与改进的颜色直方图特征相融合,充分利用了交通标志颜色鲜明的特征,以及Fish Net网络特征泛化能力强的特点,采用自编码网络对交通标志进行分类。该方法在德国GTSRB通用数据集上进行了实验验证,并与经典方法进行了比较,本文方法的分类性能有明显提升。(3)针对行人因为距离观察点远近不同导致的尺度差异,以及行人在直立行走状态下宽高比具有特殊性的特点,提出了基于YOLOv3的多尺度行人检测模型YV3-PD,该模型对三种尺度行人目标分别进行检测,通过特征重组实现图像的长方形网格划分,从而改变图像网格的宽高比,进而提高了行人检测的性能。YV3-PD方法分别在INRIA Person数据集和Caltech Pedestrian数据集上进行了实验验证,实验结果表明,本文方法在不提高漏检率的前提下,检测性能明显提高。(4)针对现有交通领域目标检测方法所能检测的目标类别单一的问题,本文提出了基于图像描述生成技术的交通场景语义理解方法。该方法首次将图像描述生成技术用于交通场景理解中,创建了基于La RA数据集的交通场景描述数据集,提出了基于网络特征组合优化的语义理解编码器-解码器模型,不仅能够描述场景中存在的交通目标,还能够给出直接的驾驶决策和建议。在标准数据集Flickr30k和MSCOCO上的实验定量比较了该方法与其他经典方法的性能。在自建的交通场景描述数据集上的实验定性对比了本文方法与传统方法的区别,展示了本文方法的优势。
【Abstract】 As an important part of the emerging industries,intelligent driving has become one of the strategic commanding heights of the new economy and technology in the era of artificial intelligence.In the future,It will be a vital technology to solve traffic congestion and improve the production and traffic efficiency greatly.Intelligent driving is of great significance to promote science and technology,economy,society,life,safety & security and comprehensive national strength.In this thesis,we focus on the research of traffic scene understanding method for intelligent driving.Aiming at traffic sign recognition,pedestrian detection and traffic scene semantic understanding,we investigate their problems in practical application.The thesis mainly composed of the following four parts.(1)In view of the problem that the detection effect of traffic signs is easily affected by light,weather and motion blur,a traffic sign detection method based on color probability model is proposed.In this method,the color probability model is established according to the unique color of traffic signs.The ROI(Region of Interest)of traffic signs is extracted from the color probability graph by the MSER(Maximally Stable Extremal Regions),which is determined by image classification method,thus the area of traffic signs is located accurately.This method is verified on the general GTSDB dataset.Because the number of traffic signs is significantly decreased by the ROI extraction method and thus the search space is limited,the detection performance and efficiency are improved.(2)In view of the diversity of clarity and scale of traffic sign images extracted from real scenes,our traffic sign classification method is proposed.It combines the attention-based Fish Net depth feature with the improved color histogram,makes full use of the distinctive color characteristics of traffic signs and the advantages of Fish Net network in general feature extraction,and classifies traffic signs by self-coding network after the fusion of the extracted features.The proposed method is verified on the GTSRB dataset.Compared with other similar methods,the performance of our proposed method is improved obviously in classification.(3)In view of the scale difference caused by the distance between pedestrians and observation points,and the bounding box of pedestrians is rectangle when the pedestrians are walking upright,we propose a multi-scale pedestrian detection model YV3-PD based on YOLOv3.Firstly,the input image data is normalized by scale,and then the up-sampling of the high-level network is spliced twice with the features of the low-level network.After feature reorganization,pedestrian targets of different scales are detected.The combination of high-level network and low-level network makes the low-level feature receptive field unchanged with stronger feature description ability.Feature reorganization changes the aspect ratio of image grid division,so the accuracy of pedestrian detection is improved.The YV3-PD method is tested on both INRIA Person dataset and Caltech Pedestrian dataset.The experimental results show that the detection performance of our method is higher and the miss detection rate is lower.(4)In view of determining the target category in the existing traffic target detection,a traffic scene semantic understanding method based on image description generation technology is proposed.It is the first time to apply image description generation technology to the field of transportation.We created a traffic scene description dataset based on Lara,and introduced a semantic understanding encoder-decoder model based on network feature combination optimization to generate rich semantic information about traffic targets.It can not only describe the traffic targets,but also determine their orientation,which is beneficial for the driving decisions.We carried out the experiments on Flickr30 k and MSCOCO to compare the performance of the proposed method with other classical ones.The experiments on the self-built traffic scene description dataset qualitatively compared the difference between our method and traditional methods,and demonstrated our merits.
【Key words】 Intelligent driving; traffic scene understanding; traffic sign recognition; pedestrian detection; scene semantic understanding;