节点文献

基于动态图卷积强化学习的多智能体协同决策

Multi-agents Collaborative Decision Making Based on Dynamic Graph Convolutional Reinforcement Learning

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 刘瑞航刘海颖李志豪刘宇辰

【Author】 LIU Rui-hang;LIU Hai-ying;LI Zhi-hao;LIU Yu-chen;College of Astronautics, Nanjing University of Aeronautics and Astronautics;Nanjing Center for Applied Mathematics;

【机构】 南京航空航天大学航天学院南京应用数学中心

【摘要】 针对部分可观测环境下多智能体协同决策中存在的不确定性和高维度等问题,提出了一种动态图卷积强化学习算法——DGC-RL。算法由观测编码器、卷积层和Q网络三类模块组成,利用动态图结构动态地表示智能体之间的关系,并通过多头注意力卷积核从局部到全局地提取智能体之间的潜在特征,引入时间关系正则化项对损失函数进行修正以增强智能体之间的合作一致性。在典型对抗场景3v3,5v5,25v25和corridor下进行实验,结果表明上述算法相比于IQL、CommNet算法,在平均奖励和平均胜率上有所提升。所提算法为部分可观测环境下多智能体协同决策提供了一种新颖有效的解决方案,为探索动态图卷积强化学习在智能化作战场景中的进一步应用奠定了基础。

【Abstract】 To address the problems of uncertainty and high dimensionality in collaborative decision making of large-scale multi-agents in partially observable environments, a dynamic graph convolutional reinforcement learning algorithm is proposed. The algorithm consists of three types of modules: observation encoder, convolutional layer and Q-network. It utilizes the dynamic graph structure to dynamically represent the relationship between agents, and extracts the potential features between agents from local to global via a multi-head attention convolution kernel, and introduces a temporal relation regularization term to correct the loss function to enhance the cooperative consistency between agents. Experiments are conducted under typical confrontation scenarios 3v3,5v5,25v25 and corridor, and the results show that the algorithm has a considerable improvement in average reward and average win rate compared with IQL and CommNet algorithms. This algorithm provides a novel and effective solution for large-scale multi-intelligence collaborative decision-making in partially observable environments, laying the foundation for exploring further applications of dynamic graph convolutional reinforcement learning in intelligent combat scenarios.

【基金】 装备预先研究基金(50912020401)
  • 【文献出处】 计算机仿真 ,Computer Simulation , 编辑部邮箱 ,2025年11期
  • 【分类号】TP18
  • 【下载频次】32
节点文献中: 

本文链接的文献网络图示:

本文的引文网络