节点文献

基于强化学习的无人水面艇系统自适应最优控制研究

Adaptive Optimal Control for Unmanned Surface Vehicles Using Reinforcement Learning

【作者】 陈林;

【导师】 戴诗陆;

【作者基本信息】 华南理工大学 , 控制科学与工程, 2024, 博士

【摘要】 无人水面艇是一种能够在水面环境下安全自主航行,并可以自主完成多样化作业任务的智能运动平台,在海洋环境监测、海洋测绘、水上搜救等民用和军事领域得到广泛应用。在无人水面艇航行关键技术领域中,轨迹跟踪和协同编队控制问题是备受关注的研究热点。无人水面艇轨迹跟踪控制是运动控制的基础。为了扩大无人水面艇的作业范围以及提高无人水面艇作业时的容错能力和可靠性,无人水面艇编队控制是实现协同作业的基本保障。在保证无人水面艇正常工作的基础上降低能源消耗是非常有必要的,已经成为了最优控制领域中的研究热点。然而,无人水面艇系统具有强非线性、强耦合、模型不确定、欠驱动等特性,在实际作业时会受到风、浪、洋流等环境扰动的影响,并且在控制设计中还需要考虑性能约束以保证无人水面艇的航行安全,所以对无人水面艇系统最优控制的研究是具有挑战性的。本文在现有研究工作的基础上,对复杂海洋环境下无人水面艇系统最优控制问题展开研究,建立了基于动态规划思想的强化学习理论,提出了新的解决复杂海洋环境下无人水面艇系统的鲁棒轨迹跟踪控制、指定性能编队控制以及有限时间包含控制等一系列问题的最优控制设计方法,不仅提高了系统的鲁棒性,而且在优化系统性能的同时保证了跟踪误差满足指定性能约束。主要研究内容如下:首先,研究具有模型不确定且受外界时变扰动的欠驱动无人水面艇系统在线最优轨迹跟踪控制问题。采用神经网络对系统的不确定动态进行在线学习,并使用自适应控制技术来估计未知扰动的上界。结合反步法和执行-评价结构的强化学习,设计最优虚拟控制输入和最优实际控制输入为相应子系统的最优控制解,从而优化整个系统的控制性能。同时,该算法可以实现神经网络逼近器权值、执行网络权值和评价网络权值的同步在线更新。将动态面控制技术引入到最优控制设计中避免近似最优虚拟控制器的导数,从而可以得到一个简单的评价网络权值更新率,降低了计算负担。其次,研究具有领导者-跟随者结构的无人水面艇系统最优编队控制问题。为了保证编队误差满足预先指定的暂态及稳态性能,选取合适的性能函数来约束编队误差的上下界。同时,通过恰当地设计性能函数的最大超调量,使得系统的控制性能约束条件包括避碰约束条件。引入barrier函数对受约束的跟踪误差进行转化,再利用变换后的跟踪误差和控制输入来设计最优值函数。然后,采用执行-评价框架的强化学习算法,以最小化Bellman残差为学习目标,设计最优编队控制器,不仅使得编队误差满足指定性能约束,而且确保编队中每个跟随者不与其领导者发生碰撞。然后,研究指定性能约束下具有模型不确定且受外部时变扰动的无人水面艇系统最优一致性控制问题。设计自适应神经网络辨识器对未知的系统动态进行辨识/学习,并利用扰动观测器来估计外部时变扰动。设计合适的指数衰减形式的性能函数以约束系统的一致性误差。利用评价网络设计最优控制器,简化了算法结构。在设计最优控制器的过程中,利用经验回放方法,记录的历史数据用于调节评价网络权值,从而避免了为了使评价网络权值参数收敛而加入持续激励信号。此外,神经网络辨识器权值和评价网络权值是同步调整的。最后,研究含有模型不确定且受外部时变扰动的欠驱动无人水面艇系统有限时间最优包含控制问题。所考虑的欠驱动无人水面艇的惯性矩阵是非对角的,会导致转向力矩同时作用于横荡速度和艏摇角速度方向上的两个动态方程,使得最优控制器设计以及稳定性分析更加困难。为了解决这个问题,首先研究单个欠驱动无人水面艇系统最优轨迹跟踪控制问题。在此基础上,进一步研究欠驱动无人水面艇系统有限时间最优包含控制问题。结合反步法和简化的强化学习设计最优控制器,实现了包含误差在满足指定性能约束下有限时间收敛到原点的小邻域内。通过对比仿真,证明了所提出的最优控制方法能够优化系统的综合性能指标(即跟踪误差和控制输入)。

【Abstract】 Unmanned surface vehicle(USV)is an intelligent motion platform that can navigate safely and autonomously in the marine environment and complete diverse operation tasks independently,it has been widely used in civil and military fields,such as marine environmental monitoring,ocean surveying,and water rescue.In the technical field of the navigation control of USV,both trajectory tracking control and cooperative formation control have been attracted considerable attention.The trajectory tracking control is the basis of motion control of USV.To expand the operating range of a single USV and improve its fault tolerance and reliability,the formation control of multiple USVs is the basic guarantee to achieve operations collaboratively.Reducing energy-consumption on the basis of ensuring the normal operation of USVs is very necessary,and it has become one of the research hotspots in the optimal control field.However,the USV systems have the characteristics of strong nonlinearity,strong coupling,modeling uncertainty and underactuation,and they are affected by the environmental disturbances,such as wind,waves,and ocean currents.Additionally,the performance constraints need to be considered in control design to guarantee the navigation of USV systems safety.Thus,research on optimal control of USV systems is challenging.On the basis of existing research works,this thesis investigates the optimal control problem of USV systems in the complex marine environments,establishes a reinforcement learning(RL)theory based on dynamic programming,and proposes new optimal control design methods to solve a series of problems,such as robust trajectory control,formation control with performance constraints and finite time containment control of USV systems,which not only improves the robustness of the systems,but also ensures that the tracking errors satisfy the predefined performance constraints while optimizing the system performance.The main research contents are given below.Firstly,the online optimal trajectory tracking control problem of an underactuated USV system with modeling uncertainties and external time-varying disturbances is investigated.Neural networks(NNs)are employed to learn the uncertain system dynamics online and adaptive control technique is used to estimate the upper bounds of the unknown disturbances.Combining backstepping method with actor–critic RL algorithm,the optimal virtual control input and the optimal actual control input are designed as the optimal control solutions of the corresponding subsystems,so as to optimize the control performance of the whole systems.Meanwhile,this algorithm can update the NN approximator weight,the actor network weight,and critic network weight synchronously.Dynamic surface control technique is introduced into the optimal control design to avoid the derivative of the approximate optimal virtual controller,which yields a simple critic network weight updating law and reduces the computation burden.Secondly,the leader-follower optimal formation control problem for multiple USVs is studied.To provide the transient and steady-state performance specifications on formation errors,suitable performance functions in the form of exponential decay are designed to constrain the upper and lower bounds of formation errors.By appropriately choosing the maximum allowable overshoots of performance functions,the collision avoidance constraint between the follower and its leader is incorporated into the predefined performance constraint.Barrier function is introduced to transform the constrained tracking error,and the optimal value function is designed by using the transformed tracking error and control input.With the learning objective of minimizing Bellman error,an optimal formation controller is designed by using the RL algorithm of actor–critic framework,which not only guarantees the prescribed performance constraint,but also avoids collision between each USV and its leader.Thirdly,the optimal consensus control problem of a group of USVs subject to prescribed performance constraints is investigated,in the presence of modeling uncertainties and external time-varying disturbances.An adaptive NN identifier is designed to identify/learn the unknown system dynamics,and the disturbance observer is used to estimate the disturbances.Suitable performance functions in the form of exponential attenuation are designed to constrain the consensus errors.After that,the optimal consensus controller is designed based on the critic network,which simplifies the structure of the developed algorithm.In the process of designing the optimal controller,the recorded historical data is used to adjust the critic network weights using experience replay method,which avoids the persistence of excitation condition that is used to guarantee the parameter convergence of critic network weights.In addition,the identifier network weight and critic network weight are adjusted synchronously.Finally,the finite-time optimal containment control problem for a group of underactuated USVs with modeling uncertainties and external time-varying disturbances is investigated.Note that the non-diagonal inertia matrix considered in the underactuated system model results in the yaw moment acting simultaneously on the sway and yaw dynamics,which makes the design of optimal controller and stability analysis more difficult.To overcome the difficulties raised by underactuation and non-diagonal inertia matrix,the optimal trajectory tracking control problem of a single underactuated surface vehicle system is first studied.On this basis,the finite-time optimal containment control problem of underactuated USVs is further discussed.By combining backstepping method with simplified RL,an optimal controller is designed to guarantee that the containment errors converge to a small neighborhood of the origin in finite time under the satisfaction of the specified performance constraints.Comparative simulation results prove that the presented optimal control method can optimize the performance indices(i.e.,tracking error and control input).

  • 【分类号】U664.82
节点文献中: 

本文链接的文献网络图示:

本文的引文网络