节点文献

基于LLM Prompt推理的虚拟电厂MAPPO交易策略

Virtual Power Plant Trading Strategy in the Electricity Market Based on Prompt-LLM & MAPPO

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 丁逸文付蓉葛辉

【Author】 DING Yiwen;FU Rong;GE Hui;College of Automation, Nanjing University of Posts and Telecommunications;Institute of Carbon Neutral Advanced Technology, Nanjing University of Posts and Telecommunications;

【通讯作者】 付蓉;

【机构】 南京邮电大学自动化学院南京邮电大学碳中和先进技术研究院

【摘要】 为提升虚拟电厂在电能量市场中的竞价收益与策略稳定性,提出一种融合大语言模型(large language model, LLM)语义推理与多智能体近端策略优化(multi-agent proximal policy optimization, MAPPO)的智能竞价方法。首先,以虚拟电厂作为激励与决策主体,提出基于LLM语义推理的虚拟电厂多智能体竞价决策架构,将历史市场出清结果、储能运行状态、电价变化趋势及竞价规则作为运行状态输入,通过提示词(Prompt)指令推理机制引导LLM对相关信息进行语义解析,从而识别虚拟电厂竞价过程所处的任务阶段,并生成结构化的奖励引导信号;然后,将LLM生成的语义状态与奖励信息分别嵌入MAPPO的策略(Actor)网络与价值(Critic)网络中,用于辅助竞价策略生成与价值函数评估,增强强化学习算法对市场规则与长期收益特性的感知能力;接着,将市场出清结果与语义奖励共同参与优势函数计算,并在近端策略优化裁剪机制下实现策略的稳定训练与收敛;最后,基于IEEE-39节点系统构建算例,设置包含多个虚拟电厂参与电能量市场竞价的多智能体仿真环境,计算并对比标准MAPPO与所提方法在不同电价波动与储能初始状态下的运行结果。算例表明:正常电价波动场景下,所提策略较标准MAPPO策略在训练末期收益提升约19%;在高波动场景下,市场出清价格标准差降低约16.3%。这验证了所提方法能够有效提升虚拟电厂竞价策略的收敛速度、收益水平与运行稳定性。

【Abstract】 To improve the bidding revenue and strategy stability of the virtual power plants(VPPs) in the electric energy market, this paper proposes an intelligent bidding method that integrates large language model(LLM) based semantic reasoning with multi-agent proximal policy optimization(MAPPO). Firstly, with the VPP modeled as the incentive and decision-making agent, an LLM driven multi-agent bidding decision-making architecture is developed for VPPs. Historical market clearing results, energy storage operating states, electricity price variation trends and bidding rules are taken as state inputs, and a Prompt-based reasoning mechanism is employed to guide the LLM in semantically analyzing the relevant information, thereby identifying the task stage of the VPP bidding process and generating structured reward guidance signals. Then, the semantic states and reward information generated by the LLM are embedded into the Actor and Critic networks of MAPPO respectively to assist bidding strategy generation and value function evaluation, thus enhancing the ability of the reinforcement learning algorithm to perceive market rules and long-term revenue characteristics. Next, market clearing results and semantic rewards are jointly incorporated into the advantage function calculation, and stable policy training and convergence are achieved under the proximal policy optimization clipping mechanism. Finally, a case study is constructed based on the IEEE 39 bus system, in which a multi-agent simulation environment involving multiple VPPs participating in electric energy market bidding is established, and the performance of the proposed method is compared with that of standard MAPPO under different electricity price fluctuations and initial energy storage states. The results show that under normal electricity price fluctuation scenarios, the proposed strategy improves the late-stage training revenue by approximately 19% compared with standard MAPPO, and under highly volatile scenarios, the standard deviation of market clearing prices is reduced by approximately 16.3%. These results verify that the proposed method can effectively improve the convergence speed, revenue performance, and operational stability of VPP bidding strategies.

【基金】 国家自然科学基金项目(52077106)
  • 【文献出处】 广东电力 ,Guangdong Electric Power , 编辑部邮箱 ,2026年04期
  • 【分类号】F426.61;TM73;TP18
  • 【下载频次】79
节点文献中: 

本文链接的文献网络图示:

本文的引文网络