节点文献

基于规则与改进wall-following的多智能体协同围捕策略

Multi-Agent Cooperative Hunting Strategy Based on Rules and Improved Wall-Following

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 王佳旭冀承慧胡创业丁男

【Author】 WANG Jia-xu;JI Cheng-hui;HU Chuang-ye;DING Nan;College of Computer Science and Technology,Xinjiang Normal University;Key Laboratory of Intelligent Control and Optimization for Industrial Equipment,Dalian University of Technology;

【机构】 新疆师范大学计算机科学技术学院大连理工大学工业装备智能控制与优化教育部重点实验室

【摘要】 针对多智能体协同执行围捕任务面临的动态性,提出了一种基于规则与改进wall-following的深度强化学习算法(Rule-Based Deep Reinforcement Learning,RBDRL)。首先,RBDRL算法根据目标和障碍物在历史行为区间内动作选择的统计,进行连续执行多步动作的状态预测,并利用基于wall-following规则设计的Upward-Downward规则在四边形网格环境中生成闭环轨迹;其次,针对闭环轨迹中的冗余路径,采用缩减规则对轨迹进行优化;再次,将这些规则集成到深度强化学习框架中,并设计了综合型奖励机制,尤其在团队奖励中,特别纳入了对时间成本的考量;最后,将RBDRL算法分别与基于计数的深度强化学习算法和无规则的深度强化学习算法在包含不同规模和数量的静态与动态障碍物场景中进行对比实验。实验结果表明,所提方法在解决多智能体在动态环境中协同执行围捕任务的问题时,具有可行性与有效性。

【Abstract】 In response to the dynamics faced by multi-agent collaboration in performing encirclement tasks,a deep reinforcement learning algorithm based on rules and improved wall-following(Rule-Based Deep Reinforcement Learning,RBDRL) is proposed. Firstly,the RBDRL algorithm performs the state prediction of multi-step actions according to the statistics of the action choices of targets and obstacles in the historical behavior interval and uses the Upward-Downward rule designed based on the wall-following rule to generate a closed-loop trajectory in the quadrilateral-grid environment. Secondly,the trajectory of the redundant paths in the closed-loop trajectory is optimized using the reduction rule. Thirdly,these rules are integrated into the deep reinforcement learning framework,and a comprehensive reward mechanism is designed,especially in the team reward,especially the consideration of time cost. Finally,the RBDRL algorithm is compared with the counting-based deep reinforcement learning algorithm and the rulefree deep reinforcement learning algorithm for experiments in static and dynamic obstacle scenarios containing different sizes and numbers of obstacles,respectively. The experimental results show that the proposed method is feasible and effective in solving the problem of multi-agent cooperative hunting tasks in a dynamic environment

【基金】 国家自然科学基金资助项目(62072071,62262066)
  • 【文献出处】 计算机仿真 ,Computer Simulation , 编辑部邮箱 ,2026年01期
  • 【分类号】TP18
  • 【下载频次】21
节点文献中: 

本文链接的文献网络图示:

本文的引文网络