节点文献

基于深度强化学习的超密集小基站网络资源分配算法研究

Research on Resource Allocation Algorithm for Ultra-Dense Small Cell Networks Based on Deep Reinforcement Learning

【作者】 陈海波

【导师】 曹叶文;

【作者基本信息】 山东大学 , 信息与通信工程, 2023, 硕士

【摘要】 移动通信技术的快速发展和移动设备的普及,使得人们对于高速、稳定、可靠的移动通信需求越来越大。超密集小基站网络(Ultra-Dense Small Cell Network,UDSCN)作为一种新兴的无线通信网络形态,具有部署灵活、频谱利用率高、能耗低等优势,成为提高网络容量和覆盖率的重要手段。然而,由于其基站密度高且信道环境复杂,超密集小基站网络的资源分配问题变得尤为复杂和困难。深度强化学习作为一种新兴的机器学习方法,可以在不需要精确的问题模型的情况下,从历史数据中自动学习出最优策略。将深度强化学习应用于超密集小基站网络的资源分配问题,具有重要的研究意义。本文基于深度强化学习中的策略优化算法,提出了适用于超密集小基站网络的资源分配算法和框架,分别从功率分配、联合资源分配和资源分配框架三个方面入手,为解决超密集小基站网络的资源分配问题提供新的思路和方法。本文首先提出了一种基于A2C(Advantage Actor-Critic)的功率分配算法AUPA(A2C based UDSCN Power Allocation Algorithm),旨在通过优化功率分配策略来最大化超密集小基站网络的系统速率。首先分析了小基站网络系统模型,并对功率分配问题进行建模。然后对策略梯度进行了详细的数学推导,给出了针对小基站功率分配的算法设计和算法流程。最后通过仿真实验验证,AUPA算法能够适应复杂的网络场景,有效提升网络的总体传输速率。为了进一步优化超密集小基站网络的资源分配效果,本文提出了一种基于PPO(Proximal Policy Optimization)的联合资源分配策略优化算法PURA(PPO based UDSCN Resource Allocation Algorithm)。PURA算法融合子载波分配和功率分配,以最大化系统能量效率为目标,同时考虑用户设备的服务质量(Quality of Service,QoS)需求。仿真实验结果证明,PURA算法在最大化系统速率的同时,也能够有效降低网络的总体能耗,并提高用户的QoS满足率。最后,本文还提出了一种基于O-RAN(Open Radio Access Network)架构的超密集小基站网络资源分配策略优化框架O-URAF(O-RAN based UDSCN Resource Allocation Framework),旨在为深度强化学习算法在超密集小基站网络中的应用提供高可行性的方案。O-URAF包含执行模块和训练模块,通过模块化设计,执行模块能够及时获取最优资源分配策略,而训练模块则能够利用全局信道状态信息对神经网络进行训练。仿真结果表明,O-URAF能够在极低的执行延时下提高资源分配算法性能,具有重要的实际应用价值。

【Abstract】 With the rapid development of mobile communication technology and the popularization of mobile devices,there is an increasing demand for high-speed,stable,and reliable mobile communication.Ultra-dense small cell networks(UDSCN)have emerged as a new form of wireless communication network,with advantages such as flexible deployment,high spectrum utilization,and low energy consumption,making it an important means to enhance network capacity and coverage.However,due to its high base station density and complex channel environment,resource allocation in UDSCN has become particularly complex and challenging.Deep reinforcement learning,as an emerging machine learning method,can automatically learn the optimal strategy from historical data without requiring an accurate problem model.Applying deep reinforcement learning to resource allocation in UDSCN has important research significance.This thesis proposes resource allocation algorithms for UDSCN based on deep reinforcement learning,starting from three aspects:power allocation,joint resource allocation,and resource allocation framework,providing new ideas and methods for solving the resource allocation problem in UDSCN.Firstly,an A2C(Advantage Actor-Critic)based UDSCN power allocation algorithm(AUPA)is proposed to maximize the system rate of UDSCN by optimizing the power allocation strategy.The system model of small cell networks is first analyzed,and the power allocation problem is modeled.Then,detailed mathematical derivation of the policy gradient is carried out,and the algorithm design and flow for small cell power allocation are given.Finally,through simulation experiments,the AUPA algorithm can adapt to complex network scenarios and effectively improve the overall transmission rate of the network.To further optimize the resource allocation effect of UDSCN,a PPO(Proximal Policy Optimization)based UDSCN resource allocation algorithm(PURA)is proposed,which integrates subcarrier allocation and power allocation to maximize the system energy efficiency while considering the quality of service(QoS)requirements of user devices.Simulation results show that the PURA algorithm can effectively reduce the overall energy consumption of the network while maximizing the system rate and improving the QoS satisfaction of users.Finally,an O-RAN(Open Radio Access Network)based UDSCN resource allocation framework(O-URAF)is proposed to provide a highly feasible solution for.the application of deep reinforcement learning algorithms in UDSCN.O-URAF includes an execution module and a training module.Through modular design,the execution module can obtain the optimal resource allocation strategy in real-time,while the training module can train the neural network using global channel state information.Simulation results show that O-URAF can improve the performance of the resource allocation algorithm with extremely low execution delay and has significant practical application value.

  • 【网络出版投稿人】 山东大学
  • 【网络出版年期】2024年 01期
  • 【分类号】TP18;TN929.5
节点文献中: 

本文链接的文献网络图示:

本文的引文网络