节点文献
基于强化学习的双臂空间机器人应急姿态控制
Attitude Control for Emergency Recovery Based on Reinforcement Learning Method for Dual-arm Space Robots
【摘要】 针对双臂空间在轨服务机器人在遭遇极端异常情况下飞轮、发动机故障等传统姿态控制失效问题,提出了一种基于强化学习的双臂空间机器人应急姿态控制算法。与传统姿态控制算法相比,本文仅通过飞行器所配置的两条机械臂进行有限的飞行器姿态恢复。通过搭建算法训练的物理环境,应用无模型的近端策略优化(Proximal policy optimization,PPO)算法进行姿态控制,结合在轨操作中机械臂运动学约束,设计奖励函数优化飞行器姿态控制精度。为验证上述策略有效性,在MuJoCo仿真环境中进行星体姿态恢复数值仿真,并针对不同星体质量、不同末端负载质量等工况进行算法适应性评估,结果表明该强化学习方法能满足飞行器进行有限姿态控制的需求,无需参数调节且具有一定鲁棒性。
【Abstract】 Aiming at the traditional attitude control failure in on-orbit service dual-arm space robots under extreme abnormal conditions such as flywheel and engine malfunctions, an emergency attitude control algorithm for dual-arm space robots based on reinforcement learning is proposed. This approach achieves limited attitude recovery of the spacecraft using only the two robotic arms configured on the spacecraft which differs from traditional attitude control algorithms. A physical environment for algorithm training is constructed and a model-free proximal policy optimization(PPO) algorithm is used for attitude control. By incorporating the kinematic constraints of manipulators movements during on-orbit operations, the reward function is designed to optimize the precision of spacecraft attitude control. To validate the effectiveness of the proposed strategy, numerical simulations of the space robot attitude recovery are conducted in the MuJoCo environment. The adaptability of the algorithm is evaluated under various conditions, including various masses of the base, various masses of the end. Results demonstrate that the reinforcement learning method is suitable for spacecraft limited attitude control and show a certain robustness without the need of parameter fine-tuning.
【Key words】 robot; reinforcement learning; dual-arm collaboration; attitude control;
- 【文献出处】 南京航空航天大学学报(自然科学版) ,Journal of Nanjing University of Aeronautics & Astronautics(Natural Science Edition) , 编辑部邮箱 ,2025年03期
- 【分类号】TP18;TP242;V448.22
- 【下载频次】69