节点文献

基于强化学习的值迭代算法

Value Iteration Algorithm Based on Reinforcement Learning

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 崔军晓; 朱蒙婷; 王海燕; 章鹏; 王辉;

【Author】 CUI Jun-xiao;ZHU Meng-ting;WANG Hai-yan;ZHANG Peng;WANG Hui;Soochow University College of Computer Science and Technology;

【机构】 苏州大学计算机科学与技术学院;

【摘要】 强化学习(Reinforcement Learning)是学习环境状态到动作的一种映射,并且能够获得最大的奖赏信号。强化学习中有三种方法可以实现回报的最大化:值迭代、策略迭代、策略搜索。该文介绍了强化学习的原理、算法,并对有环境模型和无环境模型的离散空间值迭代算法进行研究,并且把该算法用于固定起点和随机起点的格子世界问题。实验结果表明,相比策略迭代算法,该算法收敛速度快,实验精度好。

【Abstract】 Reinforcement learning is learning how to map situations to actions and get the maximize reward signal. In reinforcement learning, there are three methods that can maximize the cumulative reward. They are value iteration, policy iteration and policy search. In this paper, we survey the foundation and algorithms of reinforcement learning, research about model-based value iteration and model-free value iteration and use this algorithms to solve the fixed starting point and random fixed starting point Gridworld problem. Experimental result on Gridworld show that the algorithm has faster convergence rate and better convergence performance than policy iteration.

【关键词】 强化学习; 值迭代; 格子世界;
【Key words】 reinforcement learning; value Iteration; Gridworld;
  • 【文献出处】 电脑知识与技术 ,Computer Knowledge and Technology , 编辑部邮箱 ,2014年31期
  • 【分类号】TP181
  • 【被引频次】6
  • 【下载频次】236
节点文献中: 

本文链接的文献网络图示:

本文的引文网络