节点文献
基于深度强化学习的红绿灯信号多目标协调控制方法
Multi-objective Coordinated Control Method for Traffic Signal Based on Deep Reinforcement Learning
【作者】 张泰;
【导师】 胡云峰;
【作者基本信息】 吉林大学 , 电子信息硕士(专业学位), 2025, 硕士
【摘要】 随着我国经济发展水平的提升和城市化的快速发展,城市机动车保有量在迅速增加,市民移动出行的持续需求给城市的交通系统带来了巨大的压力,随之带来的主要挑战就是交通拥堵现象的不断频发。近年来,十字路口信号控制问题已成为大多数城市所面临的重要问题,往往会导致旅行时间的增加并且会降低通行的安全性,尤其是在大城市地区,交通道路网络的容量通常很难满足交通流正常通行的需求,这使得交通流量难以得到有效的控制。传统的控制方法对路口车流状态的获取和信号动作决策的执行普遍缺乏实时性。为了突破传统方法的限制,并更好地应对交通流量的可变性,本文提出了基于深度强化学习(Deep Reinforcement Learning,DRL)的红绿灯信号多目标协调控制方法,提高交通信号控制的灵活性和适应性。本文的主要工作如下:首先,总结了国内外深度强化学习在交通信号控制领域的研究现状,介绍了交通信号控制、深度强化学习、多智能体强化学习的相关基础理论;在此技术上,在开源交通模拟器SUMO中搭建了多信号路口交通仿真模型,可以用于后续算法训练和控制方法结果评估。然后,针对单路口交通信号控制问题,提出一种基于D3QN_CBAM算法的多目标优化信号控制方法,在对决双Q网络(D3QN)方法的基础上,引入卷积块注意力模块(CBAM),提高了模型对交通状态的敏感性,可以帮助模型更加关注路口附近车辆的分布和动态信息;为进一步提高模型性能,将平均排队长度、平均等待时间和司机可以忍受的最长红灯时间的加权和作为奖励函数,在原始D3QN方法的基础上引入了红绿灯相位可变时间间隔,使得模型能够兼顾路口各个方向的交通需求;在此基础上,融合了双深Q网络(Double DQN)和决斗深Q网络(Dueling DQN),进一步提升模型的性能。仿真结果表明,所提方法在平均等待时间、平均停靠次数和平均队列长度等性能指标上与D3QN、最大压力算法和固定定时策略相比具有显著优势。最后,针对多路口交通信号控制问题,提出了基于多智能体深度强化学习协调控制算法N-D3QN(Dueling Double Deep Q Network with Nash equilibrium,N-D3QN)。该方法将信号灯的协调控制问题建模为多智能体强化学习系统,每个智能体都经过与环境交互状态信息进行网络训练,通过获取本地车道状态信息来选择最佳动作控制十字路口的车流,在策略学习过程中考虑相邻智能体的状态、行动和奖励的影响,最终达到纳什均衡。仿真结果表明,所提方法在平均等待时间、平均停车次数和平均排队长度等性能指标上与感应控制和IQL算法相比控制性能上具有明显提升。
【Abstract】 With the rapid development of economy and urbanization in our country,the number of urban motor vehicles is increasing rapidly.The continuous demand of citizens for mobile transportation brings great pressure to the urban transportation system.In recent years,the intersection signal control has become an important problem for most cities,which often leads to the increase of travel time and reduces the traffic safety.Traditional signal control methods can be divided into three categories:fixed time control,induction control and adaptive control.However,traditional control methods are generally lack of real-time.In order to overcome the limitations of traditional methods and better cope with the variability of traffic flow,Deep Retrieval Learning(DRL)has received increasing attention in recent years.DRL can provide greater flexibility and adaptability for traffic signal control by interacting with the environment to make adaptive adjustments.The main contributions of this paper are as follows:First,this paper summarizes the research status of deep reinforcement learning in the field of traffic signal control at home and abroad.SUMO,the open source traffic simulator,is selected as the traffic simulation platform,and the Deep Reinforcement Learning Network can communicate with the SUMO simulation environment through Traffic Control Interface.Then,in the single intersection traffic signal control problem,this paper presents a multi-objective optimal signal control method based on D3QN_CBAM algorithm.Convolution Block Attention Module(CBAM)is introduced on the basis of D3QN.To further improve the performance of the model,the weighted sum of the average queue length,the average waiting time and the maximum red light time tolerable by the driver is used as the rewarding function.In addition,the double DQN and dueling DQN technologies are used to further improve the performance of the model.Simulation results using SUMO show that the proposed method has significant advantages over D3QN,maximum pressure algorithm and fixed timing strategy in terms of average waiting time,average number of stops and average queue length.Finally,based on the improved single intersection control method,a multi-agent DLC algorithm N-D3QN(Dueling Double Deep Q Network with Nash Equilibrium,N-D3QN)is proposed to control multiple intersections.The problem is modeled as a multi-agent reinforcement learning(MARL)system.Each agent undergoes network training by interacting with the state information of the environment,controls the intersection by obtaining the information of its local lane state to select the best action,and considers the influence of the states,actions and rewards of adjacent agents in the process of strategy learning,and finally achieves the Nash equilibrium.Simulation results show that the performance of the proposed method is better than that of inductive control and IQL algorithm in average waiting time,average stopping time and average queue length.
- 【网络出版投稿人】 吉林大学 【网络出版年期】2025年 10期
- 【分类号】TP18;U491.54