节点文献
基于改进PPO探索增强的交通信号控制
Exploring Enhanced Traffic Signal Control Based on Improved PPO
【摘要】 针对当前城市交通路网通行效率低、交通信号控制效果发挥不充分以及深度强化学习近端策略优化(PPO)算法探索能力较弱的问题,提出3种基于探索能力增强的PPO交通信号控制算法,即基于噪声动作网络的PPO、基于动作熵奖赏衰减的PPO和基于状态计数的PPO。首先,定义交通状态、动作和奖赏,最大化发挥信号控制效果。其次,构建多智能体交通信号控制系统,并使用异步并行多进程参数共享训练机制,加快训练速度。最后,以南阳市部分城区交通路网车流量为例,在SUMO中进行仿真实验。结果表明:相比PPO算法,3种探索增强的PPO算法的在线学习控制和非在线学习控制均降低了车辆排队长度、行驶时长和等待时长。研究结果验证了探索增强的有效性。
【Abstract】 Aiming at the problems of low traffic efficiency and insufficient traffic signal control effect of the current urban traffic network, as well as the weak exploration ability of the deep reinforcement learning proximal policy optimization(PPO) algorithm, three kinds of PPO traffic signal control algorithms based on exploration capability enhancement were proposed, namely, PPO based on noise action network(NA-PPO), PPO based on action entropy reward attenuation(ER-PPO), and PPO based on state counting(SC-PPO). Firstly, traffic states, actions, and rewards were defined to maximize signal control effect. Secondly, a multi-agent traffic signal control system was constructed, and an asynchronous parallel multi-process parameter sharing training mechanism was used to accelerate the training speed. Finally, taking the traffic flow of some urban traffic networks in Nanyang City as an example, SUMO was used to carry out simulation experiments. The results show that compared to the PPO algorithm, the online learning control and non-online learning control of the three kinds of exploration enhanced PPO algorithms both reduce the vehicle queue length, driving time and waiting time, and the experiment results verify the effectiveness of the exploration enhancement.
【Key words】 traffic and transportation engineering; deep reinforcement learning; traffic signal control; proximal policy optimization(PPO); SUMO;
- 【文献出处】 重庆交通大学学报(自然科学版) ,Journal of Chongqing Jiaotong University(Natural Science) , 编辑部邮箱 ,2026年04期
- 【分类号】U491.54
- 【下载频次】19