节点文献

基于模型的强化学习中的环境动力学建模研究

Research on Environment Dynamics Modeling in Model-based Reinforcement Learning

【作者】 王聪

【导师】 郝建业; 李栋;

【作者基本信息】 天津大学 , 工程(专业学位), 2021, 硕士

【摘要】 强化学习方法目前在多个任务中取得了超越人类的表现,在各个领域收到了广泛的关注。然而,强化学习方法的学习能力和决策能力依赖于大量的环境采样,这使得强化学习在多个领域中的落地变得困难。为了解决强化学习的样本利用率低下的问题,基于模型的强化学习方法应运而生。基于模型的强化学习方法拥有较高的样本利用率,可以依赖少量的数据,完成策略的训练。然而基于模型的强化学习方法由于模型误差的存在,其策略的训练效果往往会受到限制。为了提高基于模型的强化学习方法中的环境动力学模型的准确性,以往的方法一般通过设计各种精密的网络结构来对环境进行拟合。然而,这些工作并没有充分的考虑到环境动力学自身的性质,将环境动力学视为一个黑箱并进行拟合,从而导致环境动力学的拟合准确性较低。为了解决这个问题,本文提出了一个新的环境动力学模型建模框架:环境动力学的分解预测框架。在本文提出的框架中,环境动力学将以分解预测的方式来建模。本文的框架包含两个关键组成部分:环境子动力学发现和环境分解预测模型。其中,环境子动力学发现模块使用多种方式对环境动力学进行分解,将环境动力学分解为多个简单的子动力学。环境分解预测模型模块根据环境子动力学发现模块所提供的分解结果,对环境进行分解建模,并提供更加准确的环境动力学模型,用于策略的产生。环境动力学的分解预测框架可以很容易地与现有的基于模型的强化学习方法相结合,实验结果表明,本文所提出的框架显著降低了模型误差,提高了基于模型的强化学习算法在各种连续控制任务中的性能。

【Abstract】 Reinforcement learning has achieved superior performance in many tasks and received extensive attention in various fields.However,the learning ability and decision-making ability of reinforcement learning methods depend on a large amount of environmental sampling,which makes it difficult to implement reinforcement learning in many fields.In order to solve the problem of low sample-efficiency,modelbased reinforcement learning came into being.Model-based reinforcement learning has a high sample-efficiency and can rely on a small amount of data to complete policy training.However,due to the existence of model error,the training effect of modelbased reinforcement learning is often limited.In order to improve the accuracy of the dynamics model in the model-based reinforcement learning methods,the previous methods usually fit the dynamics by designing various precise network structures.However,these works did not consider the existence of the environment properties,and regarded environmental dynamics as a black box,resulting in low fitting accuracy of environmental dynamics.To solve this problem,we propose a new modeling framework of environmental dynamics: the environment dynamics decomposition framework.In this framework,we model the dynamics of the environment in a decomposing manner.Our framework consists of two key components: the subdynamics discovery and the dynamics decomposition prediction.Sub-dynamics discovery decomposes the environment dynamics into many sub-dynamics,and dynamics decomposition prediction constructs the decomposed world model following the sub-dynamics.Our framework can be easily combined with existing Model-based reinforcement learning algorithms and empirical results show that our framework significantly reduces the model error and boosts the performance of the Model-based reinforcement learning algorithms on various continuous control tasks.

  • 【网络出版投稿人】 天津大学
  • 【网络出版年期】2024年 06期
  • 【分类号】TP181
节点文献中: 

本文链接的文献网络图示:

本文的引文网络