节点文献

基于深度强化学习的空天地一体化网络信息物理系统垂直切换策略

Vertical handover policy for cyber-physical systems aided by SAGIN based on deep reinforcement learning

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 武艳潘广川姚明旿杨清海梁中明

【Author】 WU Yan;PAN Guangchuan;YAO Mingwu;YANG Qinghai;LEUNG Victor C.M.;State Key Laboratory of Integrated Services Network, School of Telecommunications, Xidian University;Internet of Things Research Center, School of Computer and Software, Shenzhen University;

【通讯作者】 姚明旿;

【机构】 西安电子科技大学空天地一体化综合业务网全国重点实验室深圳大学计算机与软件学院物联网研究中心

【摘要】 针对空天地一体化网络信息物理系统模型复杂、很难获得网络拓扑先验知识和模型化假设的特点,研究其基于深度强化学习的垂直切换策略。首先,综合考虑系统稳定性、切换开销和网络使用成本约束,将垂直切换策略问题建模为约束马尔可夫决策过程(CMDP),并给出保证可行解存在的充分条件;其次,提出约束-近端策略优化(CPPO)算法解决该问题,并在基站侧引入分布式强化学习机制加速训练收敛。相较于基准策略,仿真验证了所提垂直切换策略的优越性和有效性。

【Abstract】 The vertical handover policy of space-air-ground integrated cyber-physical systems based on deep reinforcement learning was studied, in which the challenges of complicated network model and difficulties in acquiring prior knowledge for network topology and model were addressed. By jointly taking the system stability, handover cost and network-using cost into account, the vertical handover policy problem was modeled as a constraint Markov decision process(CMDP), and a sufficient condition to ensure the existence of a feasible solution was derived. Furthermore, a constraint-proximal policy optimization(CPPO) algorithm was proposed to solve the CMDP, and also the distributed learning scheme at base station sides was introduced to accelerate the speed of converging. Simulation results verify the validation and superiority of the proposed vertical handover policy as compared with the baselines.

【基金】 国家重点研发计划基金资助项目(No.2020YFB1807700);陕西省创新团队基金资助项目(No.2024RS-CXTD-01)~~
  • 【文献出处】 通信学报 ,Journal on Communications , 编辑部邮箱 ,2024年08期
  • 【分类号】TN927.2;TP18
  • 【下载频次】86
节点文献中: 

本文链接的文献网络图示:

本文的引文网络