节点文献

面向边缘环境的深度学习模型部署优化研究

Research on Optimization of Deep Learning Model Deployment for Edge Environment

【作者】 刘阳;

【导师】 滕颖蕾;

【作者基本信息】 北京邮电大学 , 电子与通信工程(专业学位), 2022, 硕士

【摘要】 近年来,深度神经网络在计算机视觉、自然语言处理等领域取得了持续的突破性进展。随着算法的进步,在云、端、边缘等场景部署神经网络应用的需求也逐渐增加。现有的深度学习模型计算复杂度高、参数量大,给在移动设备等硬件资源和功耗要求较为严格的边缘环境下的部署提出了很大挑战。深度模型压缩与加速技术能够在几乎不损失精度的情况下,大幅度的缩减参数量和计算量,降低了深度模型部署的难度。本文面向边缘环境,分别从深度模型剪枝压缩和模型协同推理加速两个角度展开了研究,具体研究内容如下:(1)基于强化学习的深度模型滤波器剪枝算法。针对模型剪枝过程中每层剪枝标准和剪枝率选择问题,提出了一种剪枝标准与剪枝率联合优化的滤波器剪枝方案。本文充分考虑剪枝敏感度及层与层之间的内在联系,重新建立滤波器剪枝的优化模型,并在满足目标稀疏度的基础上最小化模型剪枝后的精度损失,采用参数化DQN算法(Parametrized Deep Q-Networks,PDQN)求解此混合变量非线性优化问题。实验结果表明,所提方案在给定目标稀疏度下为每一层选择了合适的剪枝标准与剪枝率,减小了模型剪枝后的精度损失。(2)基于空间与通道注意力机制的剪枝算法。针对剪枝过程中通道重要性衡量标准问题,本文提出了一种基于注意力机制的通道重要性衡量方法。启发于注意力机制能够让模型更加关注重要的特征,通过在卷积层上引入空间通道注意力(Spatial Channel Attention,SC A)模块,以获取输出通道的注意力得分,并依据此注意力得分删除掉冗余通道。该算法将剪枝过程和网络训练相结合,引入注意力模块以较少的开销完成对通道重要性的评估。实验结果证明,该方案根据注意力得分选择冗余通道减小了剪枝操作对模型精度的影响。(3)复杂度感知的协同推理加速研究。针对在边缘环境下深度模型推理面临的高延迟和通信带宽不稳定等问题,本文提出了复杂度感知的协同推理方案。通过调节渐进推理过程中各早退分支的退出阈值,及协同推理过程中模型分割点,以应对边缘环境中的动态变化。并采用强化学习方法对退出阈值及分割点的调节策略进行优化。实验结果表明,该方案可以很好的适应通信带宽和输入数据复杂度的变化,满足不同类型的边缘智能应用需求。

【Abstract】 In recent years,deep neural networks have made continuous breakthroughs in the fields of computer vision and natural language processing.With the advancement of algorithms,the demand for deploying neural network applications in cloud,terminal,edge and other scenarios has gradually increased.The existing deep learning models have high computational complexity and large number of parameters,which pose great challenges for deployment in edge environments such as mobile devices with strict requirements on hardware resources and power consumption.Deep model compression and acceleration technology can greatly reduce the amount of parameters and calculations without losing accuracy,reducing the difficulty of deep model deployment.Facing the edge environment,this paper conducts research from the perspectives of deep model pruning compression and model collaborative inference acceleration.The specific research contents are as follows:(1)An automatic structure pruning algorithm for deep models based on reinforcement learning.Aiming at the selection of pruning standards and pruning rates for each layer in the model pruning process,a filter pruning scheme with joint optimization of pruning standards and pruning rates is proposed.This paper fully considers the pruning sensitivity and the internal relationship between layers,re-establishes the optimization model of filter pruning,and minimizes the accuracy loss after model pruning on the basis of satisfying the target sparsity,using parametrized deep qnetworks algorithm(PDQN)to solve this mixed variable nonlinear optimization problem.The experimental results show that the proposed scheme selects the appropriate pruning standard and pruning rate for each layer under the given target sparsity,which reduces the accuracy loss after model pruning.(2)Pruning algorithm based on spatial and channel attention mechanism.Aiming at the problem of channel importance measurement in the pruning process,this paper proposes a channel importance measurement method based on attention mechanism.Inspired by the attention mechanism can help model pay more attention to important features.By introducing the Spatial Channel Attention(SCA)module on the convolutional layer,the attention score of the output channel can be obtained and deleted according to the attention score.Remove redundant channels.The algorithm combines the pruning process and network training,and introduces an attention module to complete the evaluation of channel importance with less overhead.Experimental results demonstrate that this scheme selects redundant channels according to the attention score and reduces the impact of pruning operations on model accuracy.(3)Research on acceleration of complexity-aware collaborative inference.Aiming at the problems of high latency and unstable communication bandwidth faced by deep model inference in edge environments,this paper proposes a complexity-aware collaborative inference scheme.By adjusting the exit threshold of each early exit branch in the progressive inference process and the model split point in the collaborative inference process,it can cope with the dynamic changes in the edge environment.And the reinforcement learning method is used to optimize the adjustment strategy of the exit threshold and the split point.The experimental results show that the scheme can well adapt to the changes of communication bandwidth and input data complexity,and meet the needs of different types of edge intelligence applications.

  • 【分类号】TP18
节点文献中: 

本文链接的文献网络图示:

本文的引文网络