节点文献

分布式深度学习下基于强化学习的参数同步优化策略

Reinforcement Learning-based Parameter Synchronization Optimization

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 虞廖志朱晓娟

【Author】 YU Liaozhi;ZHU Xiaojuan;School of Computer Science and Engineering,Anhui University of Science and Technology;

【通讯作者】 朱晓娟;

【机构】 安徽理工大学计算机科学与工程学院

【摘要】 针对在异构环境中,计算节点之间的资源配置存在显著差异,产生的掉队节点会增加参数同步的时间开销,降低模型的整体训练效率问题,提出了一种基于强化学习的参数同步优化策略。融合各节点性能状态并构建智能调度网络对节点批次大小与通信压缩比进行自适应调整,以此平衡节点的计算与通信任务负载,使各节点的整体迭代时间趋于一致,从而消除参数同步过程中的长尾效应。实验结果表明,在CIFAR100和CIFAR10数据集上的训练中,该策略在实现相同准确率的情况下,将整体训练时间缩短了20%以上,相较于主流同步策略(如ADSP和LOCALSGD),表现出更高的效率和优越的泛化能力。

【Abstract】 In heterogeneous environments,there are significant differences in resource configurations among computing nodes. The resulting straggler nodes increase the time overhead of parameter synchronization and reduce the overall training efficiency of the model. To address this issue,this paper proposes a parameter synchronization optimization strategy based on reinforcement learning. The strategy integrates the performance states of various nodes and constructs an intelligent scheduling network to adaptively adjust the batch size of nodes and communication compression ratio. This balances the computational and communication task loads of the nodes,aligns the overall iteration time of each node,and thereby eliminates the long-tail effect in the parameter synchronization process. Experimental results show that during training on the CIFAR100 and CIFAR10 datasets,the proposed strategy shortens the overall training time by more than 20% while achieving the same accuracy. Compared with mainstream synchronization strategies( such as ADSP and LocalSGD),it exhibits higher efficiency and superior generalization ability.

【基金】 安徽高校自然科学研究重点项目(KJ2020A0300);安徽省新时代育人省级质量工程项目(研究生教育)(2024zyxwixalk079)
  • 【文献出处】 兰州工业学院学报 ,Journal of Lanzhou Institute of Technology , 编辑部邮箱 ,2026年02期
  • 【分类号】TP18
  • 【下载频次】13
节点文献中: 

本文链接的文献网络图示:

本文的引文网络