节点文献
基于期望值函数的离策略深度Q神经网络算法
Off-policy Algorithm of Deep Q-Network Based on Expected Value Function
【摘要】 深度Q神经网络算法的值函数迭代算法大多为Q学习算法,这种算法使用贪婪值函数作逼近目标,不利于深度Q神经网络算法获得长期来看更好的策略。通过以期望思想求解的期望值函数取代贪婪值函数作为更新目标,提出了基于期望值函数的离策略深度Q神经网络算法,并结合DQN算法神经网络更新方法,给出期望值函数能够作用于DQN算法的解释。通过使用该算法能够快速获得长期回报较高的动作和稳定的策略。最后分别在CarPole-v1和Acrobot仿真环境中对期望值函数的离策略深度Q神经网络算法和深度Q神经网络算法进行获取策略的稳定性对比实验,结果表明,基于期望值函数的离策略深度Q神经网络算法能够快速获得长期回报较高的动作,并且该算法表现更为稳定。
【Abstract】 The Q-learning algorithm is mostly used for the iterative algorithm of the value function in the Deep Q-Network( DQN). In the Q-learning algorithm,the greedy value function is used to approach the target,which may make against the Q-learning algorithm to get better strategic problems in the long run. By using expectation value function instead of greedy value function,the DQN based on expected value function algorithm is proposed. Combined with the DQN algorithm neural network updating method,the explanation that the expectation value function can act on the DQN algorithm is given. With this algorithm,it could obtain high long-term returns and stabled strategies quickly. Dealting with comparative experiment between the DQN based on expected value function algorithm and DQN algorithm on the expected value function separately in the CarPole-v1 and Acrobot simulation environment. The experimental results show that the DQN based on expected value function algorithm is a more stable one and it is faster to get a higher return action policy.
【Key words】 deep Q-network; expected value function; off-policy; strategy performance;
- 【文献出处】 四川理工学院学报(自然科学版) ,Journal of Sichuan University of Science & Engineering(Natural Science Edition) , 编辑部邮箱 ,2019年01期
- 【分类号】TP183
- 【被引频次】3
- 【下载频次】146