Aiming at the problem of deep neural network speeding up training on distributed multi-machine and multiGPU,this paper proposes an implementation method of remote multi-GPUs calls based on virtualization. The distributed GPU clusters deployed by remote GPU calls improve the traditional one-to-one virtualization technology and change the location of the deep neural network for parameter exchange during distributed multi-GPU training,achieve the compatibility between the two. The method utilizes the remote GP...
【基金】
国家重点研发计划项目“面向异构融合数据流加速器的运行时系统”(2016YFB1000403)
【更新日期】
2018-03-13
【分类号】
TP183;TP332
【正文快照】
2.中国科学技术大学a.苏州研究院;b.软件学院,江苏苏州215123)中文引用格式:杨志刚,吴俊敏,徐恒,等.基于虚拟化的多GPU深度神经网络训练框架[J].计算机工程,2018,44(2):68-74,83.英文引用格式:YANG Zhigang,WU Junmin,XU Heng,et al.Training Framework of Multi-GPU Deep Neur