节点文献
分布式机器学习中的自适应同步并行策略
Adaptive Synchronous Parallel Strategy in Distributed Machine Learning
【摘要】 分布式机器学习中的资源异构和资源不稳定性易造成掉队问题,使并行策略难以平衡同步滞后和过时梯度,导致同步开销较高,降低了模型的整体训练效率。因此,提出一种面向分布式机器学习的自适应同步并行策略。首先,利用计算节点参数版本和训练延迟时间识别掉队节点;其次,通过参数服务器比较最新、最旧参数的版本差和阈值,判断出计算节点所处状态;最后,基于小批量随机梯度下降算法,采用不同全局模型参数更新规则自适应调节不同状态的计算节点。实验结果表明,相较于其他并行策略,所提策略的收敛时间减少了9.61%~41.15%,准确率最高提升了3.29%。
【Abstract】 In distributed machine learning, the straggle problem caused by resource heterogeneity and resource instability leads to high synchronization overhead and reduces the overall model training efficiency. The existence of stragglers makes it difficult for existing parallel strategies to balance the effects of synchronization lag and stale gradient. To solve this problem, an adaptive synchronous parallel strategy for distributed machine learning is proposed. Firstly, the stragglers are identified by the version of compute node parameters and the training delay time. Secondly, the parameter server determines the status of the compute node by comparing the version difference of the latest and oldest parameters and the size of the threshold.Finally, based on the small-batch stochastic gradient descent algorithm, different global model parameter update rules are adapted to compute nodes in different states. The experimental results show that, compared to other parallel strategies, the convergence time of the proposed method is reduced by 9.61%~41.15%, and the accuracy of the proposed method is improved by 3.29%.
【Key words】 distributed machine learning; stragglers; parameter server; synchronization strategy;
- 【文献出处】 辽东学院学报(自然科学版) ,Journal of Liaodong University(Natural Science Edition) , 编辑部邮箱 ,2024年04期
- 【分类号】TP181
- 【下载频次】6