节点文献

基于深度强化学习的扑翼飞行机器人飞行策略研究

Flight Strategies of Flapping-Wing Aerial Robots Based on Deep Reinforcement Learning

【作者】 赵雪娜;

【导师】 贺威;

【作者基本信息】 北京科技大学 , 控制科学与工程, 2025, 博士

【摘要】 扑翼飞行机器人受生物飞行机制启发,具有机动性高、能耗低、隐蔽性强等优势,在复杂环境下具有广阔应用前景。然而,其强耦合、强非线性、欠驱动和非定常气动特性增加了高效稳定自主控制的难度。本文聚焦于模仿鸟类在自然环境中展现的多样飞行动作与控制策略,以提升扑翼飞行机器人的轨迹跟踪精度、机动性与续航能力。本文围绕仿猎鹰扑翼飞行机器人的动力学建模、多模态轨迹跟踪控制、高机动控制以及复杂风场下的能量获取机制开展研究。主要内容如下:1)针对扑翼飞行机器人非定常气动特性的建模挑战,开展了高精度建模与验证研究。基于非定常涡格法建立考虑柔性变形影响的扑动翼气动力模型,并经转台实验验证;采用人工神经网络构建数据驱动的尾翼气动力模型。基于六自由度动力学模型开发虚拟仿真平台,并结合实飞数据完成开环与闭环分析验证,为控制策略训练与测试提供支撑。2)针对扑翼飞行机器人强非线性模型与多模态飞行任务带来的控制挑战,基于前向动力学模型,分别设计了非线性模型预测控制与深度强化学习两种多模态轨迹跟踪控制器。构建包含“爬升—巡航—下降—着陆”的多模态期望轨迹,并对比传统PID控制,结果表明两种方法在适应性与控制性能方面具有一定优势。3)针对扑翼飞行机器人在非稳态、高机动任务中的控制难题,开展了基于深度强化学习的高机动控制研究。面向俯冲栖停与急转弯任务,基于虚拟仿真平台,采用端到端的策略学习方法,设计启发式奖励结构,引导策略学习高机动飞行动作。基于近端策略优化算法训练的策略在复杂高机动任务中表现出良好稳定性与泛化能力,提升了系统的自主高机动飞行能力。4)针对扑翼飞行机器人以最小能耗完成目标巡航任务的控制需求,研究扑翼飞行机器人在混合风场中的能量增益策略。建立风场下的动力学模型与典型风场模型,基于近端策略优化训练静态翱翔与动态翱翔策略,并设计基于能量状态与任务需求的飞行模式切换机制,实现路径跟踪与能量获取的协同控制,进一步提升系统的续航能力。

【Abstract】 Flapping-wing aerial robots(F WARs),inspired by biological flight mechanisms,offer advantages such as high maneuverability,low energy consumption,and strong concealment,showing broad application potential in complex environments.However,its strong coupling,high nonlinearity,underactuation and unsteady aerodynamic characteristics significantly increase the difficulty of dynamic modeling and autonomous control.This paper explores the diverse flight maneuver and control strategies exhibited by birds in natural environments to enhance the trajectory tracking accuracy,high maneuverability and endurance of FWARs.This paper focuses on the dynamic modeling,multi-mode trajectory tracking control,high maneuverability control and energy gain strategies under complex wind fields of the falconlike FWAR.The main contents are listed as follows:1)To address the challenges of unsteady aerodynamic modeling,high-precision dynamic modeling and validation are conducted.An unsteady vortex lattice method(UVLM)incorporating flexible wing deformation is established and validated through turntable experiments.A data-driven tail aerodynamic model is further constructed using artificial neural network(ANN).A virtual simulation environment is developed using a six-degree-of-freedom(6-DoF)dynamics model and validated through both open-loop and closed-loop analyses using flight test data.This platform provides a reliable foundation for training and testing control strategies.2)To address the control challenges posed by the strong nonlinearity of F WARs and the complexity of multi-mode flight tasks,two trajectory tracking approaches—nonlinear model predictive control(NMPC)and deep reinforcement learning(DRL)—are developed based on a forward dynamics model.A multi-mode reference trajectory comprising climb,cruise,descent,and landing phases is constructed.Compared with traditional PID control,both methods demonstrate significant advantages in adaptability and control performance.3)To address the control challenges of FWARs in unsteady and highly maneuverable flight tasks,this paper investigates a DRL-based high-maneuverability control approach.For dive-perch and agile turning tasks,an end-to-end policy learning framework is developed within a virtual simulation environment,with heuristic reward structures designed to guide the learning of maneuverable flight behaviors.The policy trained using Proximal Policy Optimization(PPO)demonstrates consistent stability and generalization in complex high-agility task,effectively enhancing the autonomous high-maneuverability flight capabilities of the F WAR.4)To address the control requirement of completing target missions with minimal energy consumption,this paper investigates energy harvesting and flight mode switching strategies for the FWAR operating in hybrid wind fields.A dynamic model under wind fields and wind field models are established.Static and dynamic soaring policies are trained using PPO,and a flight mode switching strategy is designed based on energy state and mission demands to achieve the coordinated control of path tracking and energy gain,further improving the endurance of the FWAR.

  • 【分类号】TP242;TP18
节点文献中: 

本文链接的文献网络图示:

本文的引文网络