节点文献

基于神经网络的非仿射非线性系统近似最优控制研究

Researches on Approximate Optimal Control of Nonaffine Nonlinear Systems Based on Neural Networks

【作者】 张欣

【导师】 张化光;

【作者基本信息】 东北大学 , 控制理论与控制工程, 2012, 博士

【摘要】 非线性系统的最优控制问题近些年来一直是控制领域的一个研究热点.然而,非线性系统的复杂性导致已有的最优控制方法在应用时存在各自局限性,解析解很难被获得.因此,作为一种近似求解最优控制问题的新算法,近似动态规划方法自其诞生之日就成为了解决非线性系统最优控制问题的有效方法之一.此外,近似动态规划方法获得广泛关注的另外一个重要原因在于其不仅能够成功地避免“维数灾”问题,而且可以进一步获得近似最优的闭环反馈控制律.本文采用近似动态规划理论与算法对非仿射非线性系统进行了深入的分析和设计.其中主要讨论了以下几个问题:未知离散非仿射非线性系统的最优控制问题和最优跟踪问题,执行器带有非对称死区的离散非仿射非线性系统及未知连续非仿射非线性系统的最优跟踪控制问题,连续非仿射非线性二人零和微分对策问题.针对以上几个问题,我们不仅提出了相应的收敛性判据,而且提供了一些新的设计思路和有效的设计效果.总的概括起来,本文主要工作如下:1.针对未知离散非仿射非线性系统的最优控制问题,提出了一种新型的基于近似动态规划的最优控制方案.采用递归神经网络作为神经网络辨识器重构未知系统动态.根据Lyapunov理论,证明了该神经网络辨识器具有良好的辨识效果.在建立的神经网络辨识器基础上,采用近似动态规划的方法设计最优控制器.引入了评价网和控制网这两个神经网络来执行迭代启发式动态规划算法.在考虑了神经网络近似误差的基础上,严格证明了控制网估计误差及其权值的一致最终有界性.2.针对一类未知离散非仿射非线性系统的最优跟踪控制问题,提出了在线的近似最优跟踪控制策略.设计了一个在线神经辨识器,用来构建未知系统模型.然后利用近似动态规划方法设计最优跟踪控制器.该控制器由一个稳态控制项和一个最优反馈项组成.稳态控制项保证系统在稳态阶段具有良好的跟踪性能,最优反馈控制项保证在暂态阶段镇定状态跟踪误差且最小化系统性能指标函数.根据Lyapunov理论,证明了所提出的控制策略能够保证跟踪误差及神经网络的权值是一致最终有界的.3.针对一类执行器带有非对称死区的非线性系统,提出一种基于神经网络的增强学习方法设计最优跟踪控制器.首先定义一个滤波跟踪误差,设计出控制器的基本形式,并证明了采用此控制器的闭环控制系统是稳定的.在此基础上,提出了基于增强学习的自适应控制方案.此方案设计了两个神经网络:评价网和控制网.评价网用来近似系统的性能指标,控制网不仅用来近似未知的非线性系统动态,并且实现最小化系统性能指标函数的目的.这两个神经网络的训练都是在线训练,不需要离线进行.根据Lyapunov理论,证明了闭环跟踪误差以及神经网络权值的一致最终有界性.4.针对一类未知连续非仿射非线性系统的最优跟踪控制问题,提出了一种鲁棒近似最优跟踪控制方案.该控制方案不需要与系统相关的动态信息已知,而是通过建立递归神经网络模型来重构未知系统动态.通过在模型中加入一个与建模误差相关的可调项,使得建模误差渐近收敛到零.然后基于所获得的递归神经网络模型,利用近似动态规划方法设计鲁棒近似最优跟踪控制器.该控制器由一个稳态控制项、一个最优反馈项和一个鲁棒项组成.根据Lyapunov理论证明了所提出的控制策略能够保证跟踪误差渐近收敛到零,并保证了所获得控制输入在最优控制输入的一个小的邻域内.5.针对一类非仿射非线性零和微分对策问题,提出了一种新的迭代方法用于求解最优控制策略.该迭代方法首先将非线性零和问题分解成一系列的线性零和问题,相应的Hamilton-Jacobi-Issue方程分解成一系列Riccati方程.然后在得到的线性零和微分对策的状态序列及其相应的Riccati微分方程序列之间进行迭代.在局部Lipschitz条件下,证明了迭代序列的收敛性,控制序列收敛于非线性零和微分对策问题的最优控制对.进而给出了求解非仿射非线性零和微分对策问题的最优控制策略的必要条件.6.针对一类未知非线性系统的二人零和微分对策问题,提出了一种基于神经辨识器的采用近似动态规划算法的控制方案.该方案首先通过设计一个基于递归神经网络的辨识器来近似未知非线性系统动态,并在模型中增加了一个新型的调整项.根据Lyapunov理论,证明了该递归神经网络模型的动态与原未知系统动态的误差为零.基于此模型并应用近似动态规划的方法,给出了在零和微分对策问题鞍点存在或者不存在的情况下最优性能指标和最优控制对的求解方法.最后,指出了目前近似动态规划理论研究中存在的一些问题和进一步的发展方向,并对未来的研究工作进行了展望.

【Abstract】 The optimal control problem of nonlinear systems is one of the principal and dif-ficult domain in the control field. The most existed optimal control methods due to their respective limitations, the analytical optimal solution is hard to gain. Hence, approximate dynamic programming as an effective way to deal with the optimal control problem of nonlinear systems, can overcome the "curse of dimensionality", and meanwhile obtain the approximate optimal close-loop feedback control law, has gained much attention from a lot of researchers. So, it is of great importance on nonlinear optimal control for the further research on the theory and algorithm of approximate dynamic programming. By employing the approximate dynamic programming, this dissertations makes the further research on the optimal stabi-lization and tracking control of unknown discrete-time nonaffine nonlinear systems, the optimal tracking control of discrete-time nonaffine nonlinear systems with non-symmetric dead-zone inputs, the optimal tracking control of unknown continuous-time nonaffine nonlinear systems, the zero-sum games of continuous-time nonaffine nonlinear systems. The main research of the dissertation can be briefly described as follows:1. A novel neuro-optimal control scheme is proposed for unknown discrete-time nonaffine nonlinear systems by using adaptive dynamic programming method. A neuro identifier is established by employing recurrent neural networks model to reconstruct the unknown system dynamics. By using Lyapunov theory, it is proved that the identification error converges to a small neighborhood around zero by adjusting the design parameters. Based on the established recurrent neural networks model, the approximate dynamic programming method is utilized to design the approximate optimal controller. Two neural networks are used to implement the iterative algorithm. The action neural network error and weight estimation errors are proved uniformly converge to a bounded region near the origin.2. Proposed a novel online optimal tracking control scheme for unknown general nonlinear discrete-time systems by using approximate dynamic programming method. First, an online neuro-identifier is established by employing a re-current neural network model to reconstruct the unknown system dynamics. The convergence of the identification error is proved. The optimal tracking controller is composed of the steady-state controller and the optimal feedback controller. An approximate dynamic programming method is proposed to solve the optimal feedback controller forward-in-time using online approximations. Two neural networks are used. Novel weight update rules for the critic and action neural networks are derived, and the weights are tuned online. The uniform ultimate boundedness of closed-loop system is demonstrated while considering the neural network approximation errors.3. A novel adaptive-critic-based neural network(NN) controller using reinforce-ment learning is presented for a class of nonlinear systems with non-symmetric dead-zone inputs. The adaptive critic NN controller uses two NNs:the critic NN is used to approximate the strategic utility function, and the output of action NN is to approximate the unknown nonlinear function and to minimize the strategic utility function. The tuning of the NNs is performed online with-out an explicit offline learning phase. The uniformly ultimate boundedness of the close-loop tracking error is derived by using the Lyapunov approach.4. For the first time, a novel robust approximate optimal tracking control scheme is proposed for unknown general nonlinear systems by using approximate dy-namic programming method. By a recurrent neural network model to recon-struct the unknown system dynamics. Via adding a novel adjustable term related to the modeling error, the resultant modeling error is first guaranteed to converge to zero. Then, the approximate dynamic programming method is utilized to design the approximate optimal tracking controller, which consists of the steady-state controller and the optimal feedback controller. Further, a robustifying term is developed to compensate for the neural network ap-proximation errors. Based on Lyapunov approach, stability analysis of the closed-loop system is performed to show that the proposed controller guaran-tees the system state asymptotically tracking the desired trajectory.5. Proposed a new iteration approach to solve the optimal strategies for finite-horizon continuous-time nonaffine nonlinear system quadratic zero-sum game. Through iteration algorithm between two sequences which are a sequence of state trajectories of linear quadratic zero-sum games and a sequence of cor- responding Riccati differential equations, the optimal strategies for the non-affine nonlinear zero-sum game are given. Under very mild conditions of local Lipschitz continuity, the convergence of approximating linear time-varying se-quences is proved.6. Proposed an approximate dynamic programming approach for a class of un-known nonlinear zero-sum game. A neuro-identifier based on a recurrent neu-ral network is used to approximate the unknown system dynamics. A novel adjustable term related to the modeling error is added to the recurrent neural network model, which guarantees the modeling error convergent to zero. Then, an approximate dynamic programming approach is given to solve the optimal performance index and the optimal control pair under the saddle point of the zero-sum game exists or not.Finally, concluding remarks are given. Some unsolved problems and develop-ment direction for the approximate dynamic programming are proposed. Further-more, the prospects of the further study are given.

  • 【网络出版投稿人】 东北大学
  • 【网络出版年期】2015年 07期
节点文献中: 

本文链接的文献网络图示:

本文的引文网络