节点文献
基于改进型TD3强化学习的高速飞行器姿态控制
Attitude Control of High-speed Vehicles Based on Improved TD3 Reinforcement Learning
【摘要】 针对高速飞行器再入段面临的强非线性、高不确定性以及参数快时变等挑战,结合航天器智能化发展需求,提出了一种改进型的双延迟深度确定性策略梯度(Twin Delayed Deep Deterministic Policy Gradient,TD3)端到端智能姿态控制方法。为解决TD3算法在姿态控制学习过程中存在训练不稳定、收敛困难的问题,在其马尔可夫决策过程中,设计了混合奖励机制,融合连续跟踪误差惩罚和稀疏任务完成奖励,协同引导智能体收敛;在其训练过程中,引入基于现代控制理论的先验知识约束,提出了基于行为克隆的Actor网络优化更新策略,以平衡专家经验模仿与累计回报最大化目标。仿真结果表明,在14种参数偏差组合的工况下,所提方法能够精确跟踪三通道姿态指令。
【Abstract】 To address the challenges of strong nonlinearity, high uncertainty, and rapid time-varying parameters during the reentry phase of high-speed vehicles, this study proposes an end-to-end intelligent attitude control method based on an improved Twin Delayed Deep Deterministic Policy Gradient algorithm, aligned with the demands of intelligent spacecraft development. To overcome the issues of training instability and convergence difficulties in TD3-based attitude control learning, two key innovations are introduced: a hybrid reward mechanism combining continuous tracking error penalties and sparse task-completion rewards is designed within the Markov Decision Process framework to synergistically guide agent convergence. Prior knowledge constraints derived from modern control theory are incorporated into the training process, proposing a behavior cloning-based optimization strategy for the Actor network to balance expert experience imitation and cumulative reward maximization. Simulation results show that the proposed method can accurately track the three-channel attitude commands under 14 combinations of parameter deviations.
【Key words】 high-speed vehicles; attitude control; deep reinforcement learning; behavior cloning; strongly adaptive control;
- 【文献出处】 导弹与航天运载技术(中英文) ,Missiles and Space Vehicles , 编辑部邮箱 ,2025年06期
- 【分类号】TP18;V249.1;V448
- 【下载频次】52