节点文献
基于Q学习的有限时间随机线性二次最优控制
Finite-time stochastic linear quadratic optimal control based on Q-learning
【摘要】 针对系统状态和控制均依赖于噪声的随机线性离散时间系统,采用基于值迭代的Q学习迭代算法求解模型参数部分未知的有限时间随机线性二次(SLQ)最优控制问题。首先给出SLQ最优控制问题可达性条件和适应性条件,并通过矩阵拉格朗日乘子算法得到最优控制增益矩阵序列以及相应的随机代数Riccati方程(SARE)。其次,以值迭代算法为基础定义Q函数,利用Q学习迭代算法获得每个最优控制增益矩阵所对应的迭代控制增益矩阵序列和H矩阵序列。该算法依赖于系统状态信息,摆脱了系统模型参数部分未知的限制,并证明控制增益矩阵序列收敛到各自的最优控制增益矩阵,H矩阵序列收敛到各自的最优H矩阵。最后通过一个仿真实例说明了Q学习迭代算法的有效性。
【Abstract】 A Q-learning iteration algorithm based on value iteration is adopted to obtain the finite-time stochastic linear quadratic(SLQ) optimal control for stochastic linear discrete-time systems of partially unknown parameters with state and control dependent on noises. First, the condition of the attainability and well-posedness for the SLQ optimal control problem is given. In the meantime, the optimal control gain matrix sequence and the corresponding stochastic algebra Riccati equations(SAREs) are obtained by the matrix Lagrange multiplier algorithm. Secondly, a Q function is defined by the value iteration algorihtm, and a Q-learning iteration algorithm is introduced to get the iteration control gain matrix sequence and H matrix sequence for every optimal control gain matrix. The algorithm relies on the system state information, which partially gets rid of the restriction of the system unknown parameters. Then, the convergence analysis of the iteration algorithm is presented to prove that the control gain matrix sequences converge to the respective optimal control gain matrix and the H matrix sequences converge to the respective optimal H matrix. Lastly, one simulation example is provided to verify the effectiveness of the theoretical discussions.
【Key words】 Q-learning; optimal control; stochastic algebra Riccati equation; control gain matrix;
- 【文献出处】 沈阳师范大学学报(自然科学版) ,Journal of Shenyang Normal University(Natural Science Edition) , 编辑部邮箱 ,2020年03期
- 【分类号】O232
- 【被引频次】1
- 【下载频次】240