节点文献

剪枝量化联合的目标检测网络压缩算法

Object Detection Network Compression Algorithm Based on Pruning and Quantization

【作者】 徐磊;

【导师】 李君宝;

【作者基本信息】 哈尔滨工业大学 , 仪器科学与技术, 2021, 硕士

【摘要】 现今流行的目标检测算法通常是基于卷积神经网络的,由于计算量和参数量的限制,难以部署在嵌入式平台等计算资源受限的平台上。而随着工业界对目标检测任务速度的需求越来越高,普通的目标检测模型的检测速度难以胜任。虽然近年来嵌入式平台的计算能力也在高速发展,然而目标检测算法的计算量和参数量庞大仍是制约目标检测算法实际应用的主要因素。因此,研究针对目标检测算法的模型压缩算法对于目标检测算法在工业界中的应用有着十分重要的意义。为此,本文通过深入研究目标检测算法,优化设计了适合于嵌入式平台部署的目标检测网络。并在优化网络的基础上,研究提出了基于剪枝量化联合的目标检测网络压缩算法,能够在保持模型准确率的同时,大幅压缩模型大小,提升模型推理速度。在Jetson Xavier NX嵌入式平台上对比了压缩前后模型的性能指标,验证了基于剪枝量化联合的检测网络压缩算法的有效性。针对现今流行的目标检测网络计算量和参数量庞大,不适合于嵌入式平台部署的问题,设计了适合于嵌入式平台部署的目标检测网络。在深入分析不同目标检测算法的基础上,在数据集增强、权重优化方法、网络结构设计等方面,采用了多种针对目标检测任务的优化方法。参考成功的目标检测网络的设计方法,设计了一个适合于嵌入式平台部署的目标检测网络,并在COCO数据集上与其他目标检测算法进行了对比,验证了优化网络的性能。针对传统通道剪枝方法对于模型准确率影响大,稀疏化效果差的问题,提出了一种渐进式的稀疏化训练方法。该方法通过将通道划分为剪枝部分、待定部分和保留部分三部分,在稀疏化训练过程中对每个部分采取不同的稀疏化策略。与传统的稀疏化训练方法相比,所提方法对模型的准确率影响更小,产生的稀疏化效果更好。为了验证通道剪枝模型压缩算法的效果,在PASCAL VOC数据集和WIDER FACE数据集上对优化的检测网络进行了通道剪枝实验。通过分析不同剪枝比例下剪枝模型在准确率、计算量、参数量以及推理时延指标上的表现,选取了能够平衡准确率和推理速度的剪枝模型,同时验证了通道剪枝算法的有效性。针对单一模型压缩方法的模型压缩效果有限的问题,提出了基于剪枝量化联合的目标检测网络压缩算法。该方法在剪枝后的模型上进行参数量化,联合剪枝和量化两种模型压缩算法,能够进一步地提升模型压缩效果。通过分析对比训练后量化和训练时量化方法的优劣,选择训练时量化方法对剪枝后模型进行参数量化。为了验证基于剪枝量化联合的压缩算法的效果,在PASCAL VOC数据集和WIDER FACE数据集上进行了参数量化实验。实验结果表明,基于剪枝量化联合的压缩算法相较于单一模型压缩算法,能够在达到更高的压缩比的同时,保持更优的模型准确率。本文压缩算法在VOC数据集上压缩比可达15.24,推理时延9.12ms,准确率损失仅3.89%;在WIDER FACE数据集上压缩比达到51.57,推理时延7.58ms,准确率损失仅2.42%。

【Abstract】 Nowadays popular object detection algorithms are usually based on convolutional neural networks,which are difficult to be deployed on platforms with limited computational resources such as embedded platforms due to the limitation of computational amount and large number of parameters.However,with the increasing demand for object detection task in the industry,the detection speed of the common object detection model is not up to the standard.Although the computing power of embedded platform is developing rapidly in recent years,the large amount of computation and the large number of parameters of object detection algorithm are still the main factors that restrict the practical application of object detection algorithm.Therefore,it is of great significance to study the model compression algorithm of object detection algorithm for the application of object detection algorithm in industry.In this thesis,an object detection network suitable for embedded platform deployment is optimized and designed through in-depth study of object detection algorithm.On the basis of optimized network,a object detection network compression algorithm based on pruning and quantization is proposed,which can greatly compress the model size and improve the reasoning speed of the model while maintaining the accuracy of the model.Experimental results on Jetson Xavier NX embedded platform show that the compression algorithm based on pruning and quantization is effective.In order to solve the problem that the popular object detection network is not suitable for embedded platform deployment due to the large amount of computation and the number of parameters,a object detection network suitable for embedded platform deployment is designed.On the basis of in-depth analysis of different object detection algorithms,a variety of object detection optimization methods are used in dataset augmentation,weight optimization,network architecture design and other aspects.By referring to the successful design method of object detection network,an object detection network suitable for embedded platform deployment is designed,and the performance of the optimized network is verified by comparing it with other object detection algorithms on COCO dataset.Aiming at the problem that traditional channel pruning method has great influence on model accuracy and poor sparsity effect,a progressive sparsity training method was proposed.In this method,the channel is divided into three parts: pruning part,undetermined part and reserved part,and different sparse strategies are adopted for each part in the sparsity training process.Compared with the traditional sparsity training method,the proposed method has less impact on the accuracy of the model and produces better sparsity effect.To verify the effectiveness of the channel pruning model compression algorithm,channel pruning experiments were carried out on Pascal VOC datasets and Wider Face datasets for the optimized detection network.By analyzing the performance of the pruning model in terms of accuracy,calculation amount,number of parameters and inference latency under different pruning ratios,a pruning model that can balance accuracy and inference speed was selected,and the effectiveness of channel pruning algorithm was verified.Aiming at the problem that the model compression effect of single model compression method is limited,a object detection network compression algorithm based on pruning quantization joint is proposed.This method applies quantization on the pruned model,and the combination of pruning and quantization can further improve the compression ratio of the model.By analyzing and comparing the advantages and disadvantages of the post-training quantization method and the during-train quantization method,the post-pruned model was quantified by during-train quantization method.In order to verify the effect of the compression algorithm based on pruning and quantization,quantization experiments were carried out on Pascal VOC datasets and Wider Face datasets.The experimental results show that the compression algorithm based on pruning and quantization can achieve higher compression ratio and maintain better model accuracy compared with the single model compression algorithm.The compression ratio of the proposed algorithm on VOC dataset can reach 15.24,the inference delay is 9.12 ms,and the accuracy loss is only 3.89%.The compression ratio reaches 51.57 on Wider Face dataset,the inference delay is 7.58 ms and the accuracy loss is only 2.42%.

  • 【分类号】TP368.1;TP18;TP391.41
  • 【被引频次】5
  • 【下载频次】420
  • 攻读期成果
节点文献中: 

本文链接的文献网络图示:

本文的引文网络