节点文献

基于多Agent深度强化学习的多边缘任务卸载和服务缓存联合优化算法研究

Research on Joint Optimization Algorithm for Multi-edge Task Offloading and Service Caching Based on Multi Agent Deep Reinforcement Learning

【作者】 张玉;

【导师】 吕增威; 张文化;

【作者基本信息】 合肥工业大学 , 计算机技术(专业学位), 2024, 硕士

【摘要】 多边缘协作计算方式可以将任务所需的服务环境存储在距离用户更近的边缘端,并通过多边缘网络进行协作卸载计算,以克服传统云计算方式存在的传输距离长、响应时间慢等问题。其中,边缘设备通过权衡系统资源使用状况和网络条件,为计算任务选择合适的卸载位置以最小化总时延。同时,由于计算任务的执行依赖特定的服务环境,这使得计算任务的卸载过程与服务环境缓存过程表现出耦合关系。因此,需要对计算任务的卸载策略和服务环境的缓存策略进行联合优化。由于联合优化问题的规模和复杂度非常高。现有研究通常将上述联合优化问题分解为两个相互关联的子问题,并分别求解子策略。然而,这种独立的策略求解方式忽略了两个策略优化过程之间的耦合关系,且在复杂场景下求解效率低下,难以获得最优的联合策略。同时,由于任务卸载过程受到服务环境缓存状态的约束,导致卸载地的选择较为固定,从而可能造成多边缘系统负载失衡。为解决现有研究中存在的问题,本文基于多Agent深度强化学习方法进行了如下相关研究:1)针对多边缘场景下任务卸载和服务缓存过程耦合问题,设计了一种多智能体双演员-共享评论家网络架构求解任务卸载和服务缓存联合优化策略。首先,利用强化学习方法将该场景下的任务卸载和服务缓存过程建模为马尔可夫博弈模型。同时,针对模型中智能体多种动作空间需求,设计了一种基于多Agent深度强化学习决策框架求解联合策略。该框架基于双演员网络实现智能体任务卸载动作和服务缓存动作的独立输出,并使用共享评论家网络进行智能体状态值估计保证策略优化目标一致性。最后,通过仿真实验验证了所提网络架构的有效性和先进性。2)针对多边缘场景下传统主动式服务缓存带来的资源利用率低下问题,设计了一种多智能体场景下基于自注意力机制的多演员-集中式评论家网络架构求解任务卸载和服务缓存联合优化策略。基于新型虚拟化技术搭建了一种可迁移服务环境下的计算任务调度模型。同时,为解决该模型存在的调度过程耦合以及缓存内容替换问题,将上述过程中的任务卸载、服务环境迁移和内容替换过程融合为任务调度动作,进一步设计了一种基于多Agent深度强化学习的多边缘协作任务调度算法。该算法使用集中式评论家网络进行全局状态值估计保证协作策略全局最优,并通过引入自注意力机制帮助系统关注不同智能体对系统的贡献,以优化策略学习过程。仿真实验表明,使用所提算法可有效提升多边缘协作系统的负载均衡水平并获得更低的平均任务执行时延。

【Abstract】 The multi edge collaborative computing method can store the service environment required for tasks at the edge closer to the user,and perform collaborative offloading calculations through multi edge networks to overcome the problems of long transmission distance and slow response time in traditional cloud computing methods.Among them,edge devices select appropriate offloading locations for computing tasks to minimize total latency by balancing system resource usage and network conditions.Meanwhile,due to the dependency of computing tasks on specific service environments,the unloading process of computing tasks exhibits a coupling relationship with the caching process of service environments.Therefore,it is necessary to jointly optimize the offloading strategy of computing tasks and the caching strategy of service environments.Due to the high scale and complexity of joint optimization problems.Existing research typically decomposes the aforementioned joint optimization problem into two interrelated sub problems and solves each sub strategy separately.However,this independent strategy solving method ignores the coupling relationship between the two strategy optimization processes,and is inefficient in solving complex scenarios,making it difficult to obtain the optimal joint strategy.Meanwhile,due to the constraint of service environment cache status during task offloading,the selection of offloading location is relatively fixed,which may cause load imbalance in multi edge systems.To address the existing research issues,this paper conducted the following related studies based on multi-agent deep reinforcement learning methods:1)A multi-agent double-actor shared-critic network architecture is designed to address the coupling problem of task offloading and service caching processes in multi edge scenarios,and to solve the joint optimization strategy of task offloading and service caching.Firstly,using reinforcement learning methods,model the task offloading and service caching process in this scenario as a Markov game model.At the same time,a reinforcement learning decision framework based on multi-agent deep reinforcement learning was designed to solve joint strategies in response to the multiple action space requirements of intelligent agents in the model.This framework is based on a double-actor network to achieve independent output of agent task offloading actions and service caching actions,and uses a shared-critic network for agent state value estimation to ensure policy optimization goal consistency.Finally,the effectiveness and progressiveness of proposed network architecture are verified through simulation experiments.2)Aiming at the problem of low resource utilization caused by traditional active service caching in multi edge scenarios,a multi-agent multi-actor centralized-critic baed self-attention network is designed to solve the joint optimization strategy of task offloading and service caching in multi-agent scenarios.A computing task scheduling model in a transferable service environment was built based on new virtualization technology.At the same time,to address the coupling of scheduling processes and cache content replacement issues in the model,the task offloading,service environment migration,and content replacement processes in the above processes were integrated into task scheduling actions,and a multi edge collaborative task scheduling algorithm based on multi-agent deep reinforcement learning was further designed.This algorithm uses a centralized-critic network for global state value estimation to ensure the global optimization of collaborative policies,and introduces a self-attention mechanism to help the system pay attention to the contributions of different agents to the system,in order to optimize the policy learning process.Simulation experiments show that the proposed algorithm can effectively improve the load balancing level of multi edge collaborative systems and achieve lower average task execution latency.

  • 【分类号】TN929.5;TP18
节点文献中: