节点文献

关键驱动经验回放机制在自动驾驶模型中的应用

Application of Key-Driven Experience Replay Mechanism in Autonomous Driving Models

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 贾超繁; 郭旭; 包长春; 孙志诚;

【Author】 Jia Chaofan;Guo Xu;Bao Changchun;Sun Zhicheng;School of Aviation,Inner Mongolia University of Technology;College of Energy and Power Engineering,Inner Mongolia University of Technology;

【通讯作者】 包长春;

【机构】 内蒙古工业大学航空学院; 内蒙古工业大学能源与动力工程学院;

【摘要】 针对传统深度强化学习中经验回放机制存在的局限性,特别是随机抽取经验可能导致训练效率低下和模型不稳定的问题,提出一种改进的经验回放方法。设计了一种关键驱动经验回放系统,通过智能体每次训练时优先抽取经验池中最大奖励、最小奖励和中值奖励的经验,以确保关键经验的有效利用。通过从经验池中随机抽取其他经验,保持样本的多样性。此方法不仅能够有效提高样本利用率,还能显著提升模型的训练速度和性能。实验在Highway-env自动驾驶仿真环境中进行模拟。实验结果表明,改进后的经验回放机制相较于传统的随机抽取方法在成功率和平均奖励等指标上均有显著提升,收敛速度明显加快,这表明关键经验的引入在复杂任务中具有重要的影响。通过实验可得,优化经验回放机制能够有效提高深度强化学习模型在自动驾驶等复杂任务中的表现。

【Abstract】 To address the limitations of the experience replay mechanism in traditional deep reinforcement learning,particularly the inefficiency and instability caused by random experience sampling,the study proposes an improved experience replay method. The study has designs a key-driven experience replay system,where the agent prioritizes sampling experiences with the highest,lowest,and median rewards from the experience buffer during each training session to ensure effective utilization of key experiences. Additionally,random sampling of other experiences from the buffer is used to maintain sample diversity. This method not only improves sample utilization,but also significantly enhances the training speed and performance of the model. The experiment is simulated in the Highway-env autonomous driving simulation environment. Experimental results indicate that optimizing the experience replay mechanism can effectively improve the performance of deep reinforcement learning models in complex tasks such as autonomous driving.

【基金】 内蒙古自治区科技计划项目(2022YFSJ0040)
  • 【文献出处】 黑龙江科学 ,Heilongjiang Science , 编辑部邮箱 ,2025年20期
  • 【分类号】U463.6;TP18
  • 【下载频次】22
节点文献中: