节点文献

固定长度经验回放对Q学习效率的影响

Impact of Experience Replay with Fixed History Length on Q-learning

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 林明朱纪洪孙增圻

【Author】 LIN Ming,ZHU Jihong,SUN Zengqi(State Key Lab of Intelligence Technology and System,Department of Computer Science and Technology,Tsinghua University,Beijing 100084)

【机构】 清华大学计算机系智能技术与系统国家重点实验室清华大学计算机系智能技术与系统国家重点实验室 北京100084北京100084

【摘要】 提出了一种固定长度经验回放的思想,并将该思想与一步Q和Peng Q(λ)学习算法相结合,得到了相应的改进算法。该文采用不同的回放长度L将改进的算法应用在网格环境和汽车爬坡问题中进行了仿真。结果表明,改进的一步Q学习算法在两个例子中都比原算法具有更好的学习效率。改进的Peng Q(λ)学习在马尔可夫环境中对选择探索动作非常敏感,增大L几乎不能提高学习的效率,甚至会使学习效率变差;但是在具有非马尔可夫属性的环境中对选择探索动作比较不敏感,增大L能够显著提高算法的学习速度。实验结果对如何选择适当的L有着指导作用。

【Abstract】 In order to improve the learning efficiency of Q-learning,an idea of experience replay with fixed history length is proposed.This idea is integrated into one-step Q and Peng Q(λ)-learning respectively.The improved algorithms are investigated with different history length L in two learning tasks: grid world and mountain car problem.Empirical results show that improved one-step Q-learning has better efficiency than original one-step Q in both tasks.The improved Peng Q(λ) is quite sensitive to exploratory actions in Markovian environment.Increasing L can hardly enhance the performance of the algorithm,and the performance may deteriorate.However,improved Peng Q(λ) is less sensitive to exploratory actions in non-Markovian environments,and increasing L monotonously speeds up policy learning.The experimental findings also provide guidance to appropriate history length L

【关键词】 经验回放再励学习Q学习
【Key words】 Experience replayReinforcement learningQ-learning
  • 【文献出处】 计算机工程 ,Computer Engineering , 编辑部邮箱 ,2006年06期
  • 【分类号】TP181
  • 【被引频次】3
  • 【下载频次】141
节点文献中: 

本文链接的文献网络图示:

本文的引文网络