节点文献

基于深度强化学习的无人驾驶决策方法研究

Deep Reinforcement Learning Based Autonomous Driving Decision-Making Methods

【作者】 王浩;

【导师】 王忠立;

【作者基本信息】 北京交通大学 , 交通信息工程及控制, 2021, 硕士

【摘要】 在无人驾驶领域,驾驶决策是当前研究的热点和难点问题。深度强化学习(Deep Reinforcement Learning,DRL)算法寻求以端到端的方式解决问题,但一般需要大量的样本数据,同时面临输入数据复杂性高、模型复杂的问题,导致驾驶策略学习算法收敛速度慢,无法快速学习到有效策略。驾驶策略与多种因素相关,目前采用DRL的方法大多采用简单的约束奖励函数,仅能适应简单交通场景。由于实际交通场景复杂多变,导致现有的这类算法适应性较差。针对这些问题,本文提出一种基于多输入多因素约束奖励函数的驾驶决策算法框架。该方法的输入包括相机前视图、激光雷达数据及由感知结果生成的鸟瞰图。考虑到输入信息的维度高,本文研究了两种决策学习方法,并对方法的性能进行了评估。在两类策略学习方法中,奖励函数的设计均综合考虑了纵向误差、航向、驾驶平稳性、速度等多因素约束,可有效提高方法对场景的适应能力,加快策略学习的收敛速度。论文工作主要包括以下几个方面:(1)提出了一种多输入多因素约束奖励(Multi-Sensing input and Multi-factor Constraints,MSMC)的驾驶策略学习方法。分析了驾驶决策的输入信息特点,利用环境感知的结果,生成包含多类信息的鸟瞰图,选择三类数据作为学习方法的输入。设计了变分自动编码器(Variational Auto-Encoder,VAE)与软演员评论算法(Soft Actor Critic,SAC)相结合的决策学习框架。为了高效的利用多传感器输入的观测数据,利用VAE编码器网络提取低维潜在特征,作为决策学习算法的输入,从而加快策略训练过程。分析驾驶决策过程中的影响因素,基于矢量场制导定义了横向误差和方位角误差,在此基础上设计了多因素约束奖励函数,实现了可适应多场景的驾驶策略。(2)基于多输入多因素约束奖励的决策学习框架,采用随机潜在演员评论家算法(Stochastic Latent Actor Critic,SLAC)替代基于VAE的表示学习与SAC任务学习相互独立的策略学习方法,提出了多输入多因素约束奖励的SLAC方法(MSMC-SLAC)。利用部分可观测马尔科夫决策过程(Partially Observable Markov Decision Process,POMDP)与概率图模型,将特征表示学习与策略学习联合建模,进一步提高了驾驶决策学习的效率,并扩展了方法对场景的适应性。(3)对所提出的算法进行了仿真验证。采用CARLA仿真器模拟了不同交通场景,利用Pygame软件包进行可视化设计。评估算法在不同输入数据组合下的性能表现;对比了本文所提的方法与其他DRL算法的性能;对多因素约束奖励项的重要性进行实验分析;对本文提出的两种方法进行了对比实验。在仿真环境下采用不同的地图验证了MSMC-SLAC算法的对多种交通场景的适应性。图61幅,表15个,参考文献64篇。

【Abstract】 In the field of autonomous driving,decision-making is active and challenging topic currently.Deep reinforcement learning seeks to solve the problem in an end-to-end manner,but generally requires a large amount of sample data and confronted with high dimensionality of input data and complex models,which lead to slow convergence and can not learn effective strategies quickly.Driving strategies are related to a variety of factors,and most deep reinforcement learning based methods use simple constraint reward functions,can only adapt to simple traffic scenarios.Due to the complexity and variability of traffic scenarios,the existing algorithms are less adaptable.To solve these problems,this paper proposes a framework for a driving decision method based on a multi-input multi-factor constrained reward function.The inputs of the method include camera front view,Li DAR data and bird’s-eye view generated from the perception results.Considering the high dimensionality of the input information,two strategies learning methods are designed,and the performance of the methods is evaluated.In the two types of strategies learning methods,the reward functions are all designed with multi-factor constraints such as lateral error,heading,driving smoothness and speed,which can effectively improve the adaptability of the method to the scene and accelerate the convergence speed of strategies learning.The work of the paper mainly includes the following aspects:(1)A driving strategies learning method with multi-sensing input multi-factor constraint(MSMC)reward is proposed.Through analysis the input information characteristics of decision-making,and using the results of environment perception to generate bird’s-eye view,and three types of data are selected as the input of the strategies learning.A strategies learning framework combining Variational Auto Encoder(VAE)and Soft Actor Critic(SAC)algorithm is designed.In order to efficiently use the observed data from multi-sensor inputs,low-dimensional latent features are extract by VAE encoder network and then input to the algorithms to improve the training process speed.The various influencing factors in the decision-making are analyzed,lateral error and azimuth error are defined based on vector field guidance method,and then a multi-factor constrained reward function is designed,which ensure that the strategies can adapt to multiple scenarios.(2)Based on the multi-sensing input multi-factor constraint reward strategies learning framework,the Stochastic Latent Actor Critic(SLAC)algorithm is used to replace the VAE-based representation learning and SAC-based task learning with independent strategies learning methods,and the multi-sensing input multi-factor constraint reward SLAC(MSMC-SLAC)method is proposed.The efficiency of strategies learning is improved by jointly modeling feature representation learning and task learning,which use partially observable markov decision process(POMDP)with probabilistic graphical model and expand ability to adapt to scenarios.(3)The proposed algorithm is simulated and validated.Different traffic scenarios are simulated using the CARLA simulator and visualized using the Pygame software package.The performance of the algorithms is evaluated under different combo of input data;the performance of the proposed method compares with other deep reinforcement learning algorithms;The importance of the multi-factor constraint reward term is experimentally analyzed;Comparison experiments are conducted on the two methods proposed in this paper.The adaptability of MSMC-SLAC algorithm to a variety of traffic scenarios is verified under different simulation maps.Figure 61,Table 15,and 64 references.

  • 【分类号】U463.6;TP18
  • 【被引频次】1
  • 【下载频次】568
  • 攻读期成果
节点文献中: 

本文链接的文献网络图示:

本文的引文网络