节点文献

学习型自动驾驶决策算法闭环学习方法

A Closed-Loop Learning Method for Learning-Based Autonomous Driving Algorithms

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 黄岩军陈诗阳韦登伟李欣城陈虹

【Author】 Huang Yanjun;Chen Shiyang;Wei Dengwei;Li Xincheng;Chen Hong;School of Automotive Studies,Tongji University;College of Electronic and Information Engineering,Tongji University;

【通讯作者】 黄岩军;

【机构】 同济大学汽车学院同济大学电子与信息工程学院

【摘要】 为构建开放环境下高安全可信度的自动驾驶系统,本文针对自动驾驶场景的长尾分布问题,提出一种自动驾驶决策算法闭环学习方法,该方法通过安全关键场景生成与持续学习实现算法闭环。首先,对于常见驾驶场景下表现良好的基础算法,生成具有威胁性的安全关键场景,以此挖掘算法缺陷;其次,采用融合弹性权重巩固与线性多策略头的持续学习方法,在安全关键场景中进一步训练自车算法,避免灾难性遗忘问题;最后,通过多次闭环迭代提高算法场景的适应能力。本文以软演员-评论家算法为基础算法,验证所提闭环学习方法的有效性。经两轮环境差异较大、难度持续提升的闭环迭代测试,未采用持续学习策略和仅采用经验回放策略的两种基线方法与本文方法的碰撞率分别为25.40%、25.33%和14.43%。对比结果表明,本文方法抵御灾难性遗忘与探索学习新任务的综合能力更强,因此所提出的闭环学习方法可有效提高学习型自动驾驶决策算法的场景适应性,实现算法迭代优化。

【Abstract】 To build a highly secure and trustworthy autonomous driving system in an open environment, this paper proposes a closed-loop learning method for autonomous driving decision-making algorithms in response to the long-tail distribution problem in autonomous driving scenarios. This method achieves algorithmic closed-loop through the generation of safety-critical scenarios and continuous learning. Firstly, for the basic algorithm that performs well in common driving scenarios, safety-critical scenarios with threat are generated to identify algorithmic flaws. Secondly, a continuous learning method that combines elastic weight consolidation and linear multi-strategy heads is adopted to further train the self-vehicle algorithm in safety-critical scenarios, avoiding the problem of catastrophic forgetting. Finally, the algorithm′s adaptability to scenarios is enhanced through multiple closed-loop iteration. This paper takes the soft actor-critic algorithm as the basic algorithm to verify the effectiveness of the proposed closed-loop learning method. After two rounds of closed-loop iterative tests with significant environmental differences and continuously increasing difficulty, the collision rates of the two baseline methods without continuous learning strategy and only using experience replay strategy, and the method proposed in this paper are 25.40%, 25.33%, and 14.43% respectively. The comparison results show that the method proposed in this paper has a stronger comprehensive ability to resist catastrophic forgetting and explore new tasks. Therefore, the proposed closed-loop learning method can effectively improve the scene adaptability of learning-based autonomous driving decision-making algorithms and achieve iterative optimization of the algorithms.

【基金】 国家自然科学基金委企业创新发展联合基金重点项目(U23B2061);小米青年学者项目资助
  • 【文献出处】 汽车工程 ,Automotive Engineering , 编辑部邮箱 ,2026年03期
  • 【分类号】U463.6
  • 【下载频次】29
节点文献中: 

本文链接的文献网络图示:

本文的引文网络