节点文献

考虑多维安全的深度强化学习驱动微电网实时调度方法研究

Microgrid Real-Time Dispatch Driven by Deep Reinforcement Learning considering Multi-dimensional Security

【作者】 冯斌;

【导师】 郭创新; Zhe Chen;

【作者基本信息】 浙江大学 , 电气工程, 2024, 博士

【摘要】 随着国家双碳目标的提出,分布式能源并网规模大幅增加,各种不确定性因素显著增强,给微电网调度运行带来巨大的挑战。传统物理模型依赖于对实际物理环境的精确建模,但随着系统复杂度提升和不确定性增多,建模与求解变得更加复杂和耗时;而深度强化学习(Deep Reinforcement Learning,DRL)可以通过历史经验自适应地学习调度策略并实时决策,以数据驱动的方式更好地应对不确定性和复杂度。在此背景下,本文围绕考虑多维安全的DRL驱动微电网实时调度问题,从DRL驱动的微电网实时经济调度出发,着力解决数据隐私安全、模型训练安全、调度运行安全等多维安全挑战,主要研究工作如下:针对微电网实时调度问题,对分布式能源、负荷、储能等微电网组成要素进行分析,并深入分析微电网实时调度的目标函数和约束条件,建立微电网实时调度问题的数学模型,为DRL方法的应用奠定模型基础。同时,对DRL方法应用于微电网实时调度问题的适用性进行探讨。微电网实时调度问题被建模为马尔可夫决策过程,明确微电网实时调度问题中的智能体、环境、状态、动作、奖励以及状态转移函数。算例分析验证了DRL方法在满足负荷需求、降低运行成本等方面的有效性。针对微电网实时调度中的数据隐私安全问题,提出了基于联邦深度强化学习(Federated DRL,FDRL)的微电网实时经济调度方法。微电网运营商利用价格型需求响应负荷和温控负荷的灵活性,通过合理调控温控负荷、储能、售电价格和外部电网购售电量以最大化微电网运营商经济收益。所提FDRL方法利用联邦学习聚合不同微电网的梯度,保证了各微电网内部数据的隐私安全,并提升了方法收敛速度。同时,利用Transformer编码器架构对历史数据的时序特征进行提取,提升了调度策略的经济性。算例分析验证了所提FDRL方法能够在保持室内温度符合设定范围的前提下,有效调控温控负荷、储能、售电价格,并合理择时购售电,从而显著提升微电网运营商的经济收益。针对微电网实时调度中的模型训练安全问题,提出了基于鲁棒联邦深度强化学习(Robust FDRL,RFDRL)的微电网实时自平衡调度方法。所提RFDRL方法采用鲁棒梯度过滤器排除故障梯度,减少故障干扰对训练的影响以保证模型训练过程的安全。同时,为提高样本利用效率,使用随机受控的随机梯度优化加速模型收敛。算例分析验证了所提RFDRL方法采用鲁棒梯度过滤器能够高效排除故障梯度;采用随机受控的随机梯度优化相对于传统方法能实现更快的收敛速度。在不同故障干扰下进行的算例分析验证了所提RFDRL方法相对其他方法可以更好地实现微电网的实时自平衡调度。针对微电网实时调度中的调度运行安全问题,提出了基于安全深度强化学习(Safe DRL,SDRL)的微电网实时交流最优潮流计算方法。该方法结合了基于原始-对偶优化的近端策略优化、行为克隆预训练和多进程训练。其中,基于原始-对偶优化的近端策略优化方法避免在成本奖励和安全代价之间人为设置权重,通过自适应权衡奖励和代价,确保了安全约束不越限并降低运行成本;行为克隆预训练对神经网络参数进行初始化以加速后续DRL训练过程;多进程训练则进一步对学习过程进行加速。在改进的CIGRE低压微电网上进行实验,对比了经济性、约束满足程度和求解时间,验证了所提基于SDRL的微电网实时交流最优潮流计算方法能够提供更优调度方案。

【Abstract】 With the aim of national dual carbon goals,distributed renewable energy has increased significantly.This growth,combined with various uncertainties poses substantial challenges to microgrid dispatch.Traditional physical models often require precise modeling of the actual physical environment,but as system complexity and uncertainty increase,the modeling and solving processes become more complex and time-consuming.In contrast,deep reinforcement learning(DRL)can adaptively learn dispatch strategies and make real-time decisions through historical experience,addressing higher uncertainties and complexities with data-driven methods.Thus,this paper focuses on DRL-driven real-time dispatch methods for microgrids that consider multi-dimensional security.Starting from the real-time economical dispatch of microgrids based on DRL methods,this paper mainly addresses three major issues:data privacy security,model training security,and dispatch operation security.The main research works are as follows:For the microgrid real-time dispatch problem,the elements of microgrids such as distributed energy,loads,and storage are analyzed,along with a deep analysis of the objective functions and constraints of microgrid real-time dispatch.A mathematical model for the real-time dispatch problem of microgrids is established to lay the foundation for the application of the DRL method.Also,the applicability of DRL methods in microgrid dispatch problems is analyzed.The microgrid real-time dispatch problem is modeled as a Markov decision process,clarifying the agents,environment,states,actions,rewards,and state transition functions.Case studies verify the effectiveness of DRL methods in meeting load demands and reducing operational costs.For data privacy security in microgrid real-time dispatch,a real-time economical dispatch method for microgrids based on federated deep reinforcement learning(FDRL)is proposed.Microgrid operators utilize the flexibility of price-based demand response loads and thermal loads,controlling storage,thermal loads,selling prices,and external grid transactions to maximize the profits of microgrid operators.The FDRL method aggregates the gradients of different microgrids using federated learning,ensuring privacy within each microgrid and enhancing convergence speed.Furthermore,the encoder architecture in Transformers is used to extract temporal features from historical data,enhancing the economic efficiency of the dispatch strategy.Case studies show that the proposed FDRL method can effectively control storage,thermal loads,selling prices,and maintaining indoor temperatures within the set region while purchasing and selling electricity at opportune times,significantly enhancing the profits for microgrid operators.For model training security in microgrid real-time dispatch,a real-time microgrid self-balancing dispatch method based on robust federated deep reinforcement learning(RFDRL)is proposed.The proposed RFDRL method uses the robust gradient filter to filter out faulty gradients based on gradient filtering rules,reducing the impact of fault disturbances on training.At the same time,the stochastic controlled stochastic gradient is used to accelerate model convergence and improve sample efficiency.Case studies show that the proposed RFDRL method effectively excludes faulty gradients,achieves faster convergence compared to traditional methods,and provides better performance in the real-time self-balancing dispatch of microgrids under various fault disturbances.For dispatch operation security in microgrid real-time dispatch,a real-time AC optimal power flow calculation method for microgrids based on safe deep reinforcement learning(SDRL)is proposed.This method combines primal-dual optimization-based proximal policy optimization,behavior cloning pretraining and multi-process training.The primal-dual optimization-based proximal policy optimization method avoids manually setting weights between cost rewards and safety costs.It adaptively balances rewards and costs to ensure safety constraints are not exceeded and reduce operational costs.Behavior cloning pretraining initializes the policy network parameters to accelerate DRL training convergence,and multi-process training further accelerates the learning process.Experiments conducted on the modified CIGRE low-voltage microgrid compare economic efficiency,constraint satisfaction,and solution times,verifying that the proposed SDRL-based method provides superior dispatch solutions for microgrids.

  • 【网络出版投稿人】 浙江大学
  • 【网络出版年期】2025年 07期
  • 【分类号】TM73;TP18
节点文献中: 

本文链接的文献网络图示:

本文的引文网络