节点文献
基于强化学习和群集智能方法的多机器人协作协调研究
Research on Cooperation and Coordination of Multi-robot System Based on Reinforcement Learning and Swarm Intelligence Method
【作者】 王醒策;
【导师】 顾国昌;
【作者基本信息】 哈尔滨工程大学 , 计算机应用技术, 2005, 博士
【摘要】 随着多机器人系统的迅速发展,21世纪伊始就对其提出了分布、智能和协同化的要求。合理体系结构和高效协作协调算法的研究重要性日益突出,本文对这两方面进行了全面深入地研究,内容分为三部分:多机器人系统体系结构研究;多机器人系统强化学习算法研究和多机器人系统群集智能算法研究,以满足多机器人系统的低通讯量,变化性,分布性,分散性和动态性的要求。 体系结构是多机器人系统的研究基础,直接决定了机器人间的相互关系和功能的分配。本文面向多机器人系统强化学习算法,设计了多机器人分层体系结构。给出了势场栅格算法,研究了模糊控制算法和黑板式通讯。这种结构的并发性好,实时功能强,能够加强机器人对变化环境的应变能力。面向多机器人系统群集智能算法,提出了多机器人意图-行为结构,对这种结构,对各机器人的行为能力和群体交互方式进行了研究。探讨了基于对策论的无通讯协调,给出了愿望竞争算法、抑制疲劳算法,研究了机器人行为设定机制和基于信息素的通讯机制,得出该结构具有分布式控制和分散的数据量的特点的结论,这适合于相似的分布式控制系统。 强化学习理论由于其自学习性和自适应性的优点而得到了广泛地关注。但此理论在应用中还存在着状态空间压缩,结构信度分配等问题。本文面对状态空间压缩问题,提出自组织动态压缩空间算法;关于结构信度分配问题,提出兼顾系统整体利益和个体利益的内外强化信号算法,对传统强化学习算法进行了重大改进。这种状态空间压缩方法加快了算法对空间的遍历,提高了算法的学习速度;合理分配信
【Abstract】 In the 21st Century , the distribution, intelligence and the coordination requests are brought forward for the multirobot system .Reasonable architecture and effective cooperation algorithm are paid more and more attention. These two fields are studied in three parts in the dissertation, which are the study of multirobot architecture , the Reinfocement Learning algorithm of multirobot system and the Swarm Intelligence algorithm of multirobot system. All these researches can satisfy the low communication, variety, distribution and decentralization needs of the system.The architecture which determines the relationship of the robots and the task assigned is the basis of the multirobot study. Facing the Reinforcement Learning algorithm, the level architecture of the multirobot system is brought forward .At the same time, the potential grid method, action fuzzy control and the blackboard communication are studied. This architecture has the concurrent , real-time and flexible ability. Pointing to the Swarm Intelligent algorithm, the intention-behavior architecture is brought forward. And the group structure, robot ability and the communication are studied. The cooperation based Markov Game without communication is discussed. At one time, the intention competing method and the behavior restrained-exhausted method are investigated. Moreover, the robot behavior assigned mode and the communication based on pheromone spreading method are researched. This architecture has the advantages of distributed control and dispersible data. These two architectures can be widely used in the similar system.The Reinforcement Learning theory is attached importance for itsself-learning and the self-adaption. But the problems of state compressed, structure credit assigned and the task partition still prevent the theory extending .In the dissertation the self organization method is brought forward to solve the state compressed problem and the self and the group signal method is put forward to solved the structure credit assigned problem. Compressing state speeds up the ergodic so that the learning speed can be increased. Assigning the credit maps the state to behavior reasonably so that the system’s effect swing can be avoided. After the algorithm ameliorated , the self-adaption and the robust ability of the system increased rapidly.The swarm intelligent theory offers the thought that the soul of the intelligence exists more in the group than in the single part. The intelligence of the system can also be increased even though the intelligence of the single part is low. But the thought still exists the exploring fields such as lacking the transplant ability and the practical field. The group behaviors algorithms are designed based on the swarm intelligence in the dissertation. The simple interaction rules are set down to realize the complex task. The pheromone spreading method is constituted to realize the communication, which can decrease the information flowing and strengthen the stability of the system. The stability of the system is discussed and the conclusion is when the functions of the robot are chosen reasonably , the stability of the team formation can be guaranteed.At last, the compare of these two algorithms is presented and the capabilities are also analyzed. These conclusions gives the guidance of corresponding algorithms’ application in the different environment.Facing the multirobots’ team formation task, the application models of the two algorithms are discussed . In the simulation of the experiment , the feasibility of these technologies is verified further. The expands ofthe methods are strong and can be used in the similar system.
【Key words】 architecture; reinforcement learning; swarm intelligence; multirobot system; team formation;