节点文献

基于强化学习的机器人觅食问题研究

Reinforcement Learning Based Research on Robot Foraging Problem

【作者】 李建军

【导师】 李衍杰;

【作者基本信息】 哈尔滨工业大学 , 控制科学与工程, 2013, 硕士

【摘要】 机器人对环境的适应程度决定了其智能的高低,在复杂非结构环境中,机器人的适应能力受到严重挑战,这要求机器人必须具备学习能力,强化学习机制以其优秀的自适应性和对学习条件的弱要求,为机器人的行为学习提供了途径。针对觅食行为具有广泛的代表性和较高的实际应用价值,本文利用强化学习机制对机器人的觅食行为学习作出了一系列研究。强化学习算法的关键问题是算法的收敛性以及收敛速度,这决定了机器人觅食学习的成败和学习速度的快慢。文中提出将觅食行为分解成普通的行为集成块,极大地减小学习空间,建立标准的马尔可夫决策过程(MDPs),同时加入一定的先验知识加速学习过程。利用Q学习进行单个机器人觅食学习的仿真实验结果表明,分解任务和加入先验知识的措施对在线学习速度的提升效果明显。针对多机器人系统相比单个机器人具有的并行性、鲁棒性等特点,文中提出利用平均报酬的强化学习算法诱导多机器人产生协作觅食行为,并提出一种基于Schweitzer变换的相对值迭代(RVI)强化学习(RL)算法。和单个机器人觅食学习的情况类似,建立多机器人系统觅食的MDPs模型,将新算法应用于多机器人觅食学习。和Q学习对比的仿真实验结果表明,改进的RVI算法有效且具有较高的可靠性。

【Abstract】 The ability of a robot to adapt to the environment determines its intelligentdegree. Especially in the complex unstructured environment, the adaptability ofrobot is a critical issue, which requires the robot to have the ability to learn fromenvironment. With few pre-conditions, reinforcement learning can provide robotswith approaches to learn behaviors and help them have excellent adaptability. As weknow, foraging behavior is broadly representative and has important applicationvalue. Therefore, we conduct a series of research on foraging behavior by usingreinforcement learning in this thesis.The key problems of reinforcement learning are convergences and convergentrates, which determine whether the robot can learn to forage sucessfully and itslearning speed. This thesis first proposes a method that foraging behavior isdecomposed into blocks, each of which is integrated by elementary behaviors, so asto reduce learning space greatly. Then standard Markov Decision Processes (MDPs)model is established. On this basis, we add a little priori artificial knowledgerationally to accelerate the learning process. In addition, a simulation experimentthat single robot learns foraging with Q-learning algorithm is conducted, and theresult shows that method of decomposing task and priori knowledge improve onlinelearning speed obviously.Moreover, compared with the single robot system, multi-robot system hasconcurrency, robustness and other merits. Thus, this thesis discusses the idea thatmulti-robots forage cooperatively by using average-reward reinforcement learningalgorithm. We put forward a relative value iteration (RVI) reinforcement learningbased on Schweitzer’s transformation. Then, the MDPs model of multi-robotforaging is established similarly to single robot, and new RVI algorithm is applied tomulti-robot foraging. Finally, we compare the new algorithm with Q-learning in asimulation experiment and find that new RVI algorithm is effective and has highreliability.

  • 【分类号】TP181;TP242
  • 【被引频次】3
  • 【下载频次】288
  • 攻读期成果
节点文献中: