节点文献
基于分组和深度可分离卷积的轻量级卷积神经网络
Lightweight Convolution Neural Network Based on Group and Depthwise Separable Convolution
【作者】 李林;
【导师】 张可;
【作者基本信息】 电子科技大学 , 工程硕士(专业学位), 2021, 硕士
【摘要】 卷积神经网络近年来成为了解决各类视觉任务的主流选择,包括图像分类,检测跟踪,动作及意图识别等领域。卷积神经网络由一系列卷积层层堆叠构成,而传统卷积层存在着参数量和计算量大的问题,同时网络深度和宽度的提升进一步加剧参数量和计算量的问题,使得这些网络模型往往无法顺利进行移动端部署。因此设计高效的卷积神经网络具有重大的学术及工程价值。针对以上问题,本文提出了一种高效的分组卷积单元,并提出了一种基于深度可分离卷积的多尺度卷积模块。主要的研究内容和成果如下:1.通过对常用卷积方式的原理梳理,并分析对比了它们的复杂度,分组卷积作为现今轻量级卷积神经网络常用的卷积结构,既拥有大量减少参数量和计算量的能力,又拥有超过深度可分离卷积的特征生成效果。针对分组卷积方式,我们提出了基于分组卷积的高效卷积结构Fuse Conv系列,在CIFAR-10和CIFAR-100数据集上的实验以及与常用的网络剪枝算法的对比,证明了该系列高效卷积结构的可行性和有效性。2.现有的轻量级网络大多摒弃了传统卷积而使用深度可分离卷积,不可避免地存在着特征提取单一的问题,使得网络对于如目标检测,超分辨率等任务而言并不友好。受多尺度卷积特征融合的启发,我们提出了一种以深度可分离卷积为基础的基于多尺度特征的卷积模块DF(Depthwise Fuse)模块,详细说明了该模块的具体结构,并以此模块构建了DFNet网络模型。通过在CIFAR-10和CIFAR-100数据集上的实验说明了所提模块和网络的有效性,同时设置了两组消融实验以更好地理解该模块。通过在BSD数据集上的实验说明了DF模块可有效嵌入其他任务中。3.网络模型的深度和宽度对于模型的特征拟合能力有着至关重要的作用,而基于深度可分离卷积的卷积模块和对应网络中,少有具体深入探究模块结构,网络深度和宽度的影响。受此启示,我们比较了多种基于深度可分离卷积的卷积模块,通过基于CIFAR-10数据集的实验结果说明了这些模块的性能,并借助这些结果提出了性能更强的卷积模块,以此卷积模块构建了全新的轻量级卷积神经网络DBN(Deep Bottleneck Network)。通过设置不同的深度和宽度参数的实验,根据实验中网络的性能确定了最佳的深度和宽度,同时为进一步提升网络模型性能,在网络中嵌入了通道注意力模块,实现了在CIFAR-10数据集上更佳的效果。在CIFAR-100数据集上与常用的一些网络结构进行了对比,所提网络取得了不错的性能。
【Abstract】 Convolution neural networks have become the mainstream choice for various visual tasks in recent years,including image classification,detection and tracking,action and intention recognition and so on.Convolution neural networks are composed of a series of convolutional layers.However,the traditional convolutional layers have the problem of large amount of parameters and calculations.At the same time,the increase in network depth and width further aggravates this problem.Due to the problems of complexity,these models cannot be successfully deployed on mobile devices.Therefore,designing an efficient convolution neural network is of vital importance in academic and industrial research.To solve the above problems,this thesis introduced an efficient group convolution unit and proposed a multi-scale convolution module based on depthwise separable convolution.The main research content and results of our work are as follows:1.After sorting out the principles of commonly used convolution methods,and analyzing and comparing their complexity,group convolution,as a commonly used convolution structure for lightweight convolution neural networks nowadays,can greatly reduce the number of parameters and the amount of calculation,furthermore it also has feature extraction capabilities that exceed the depthwise separable convolution.According to group convolution,we propose the Fuse Conv series of efficient convolution module.Experiments on the CIFAR-10 and CIFAR-100 datasets and the comparison with commonly used network pruning algorithms prove the feasibility and effectiveness of this series of efficient convolution structures.2.Most of the existing lightweight networks abandon traditional convolution and use depthwise separable convolutions.There is inevitably the problem of single feature extraction,which makes these networks unfriendly to tasks such as object detection and super resolution.Inspired by the fusion of multi-scale features,we propose a multi-scale convolution module called DF(Depthwise Fuse)module which is based on depthwise separable convolution.The specific structure of the module is explained in detail and we constructed the DFNet model based on this structure.Experiments on the CIFAR-10 and CIFAR-100 datasets illustrate the effectiveness of the proposed module and network.At the same time,two groups of ablation experiments are set up to better understand the proposed module.Furthermore,Experiments on the BSD dataset show that the DF module can be effectively embedded in other tasks.3.The depth and width of a model play a vital role in the feature extraction ability of the model.There are few specific explorations based on depthwise separable convolution and corresponding network in terms of the basic modules,network depth,and width.Inspired by this,we compare a variety of convolution modules based on deep separable convolution and choose a module through experimental results based on the CIFAR-10 data set.Finally,we propose a more powerful convolution module and a novel lightweight convolutional neural network called DBN(Deep Bottleneck Network).Through experiments with different depth and width factors,the optimal depth and width are determined according to the performance of the network in the experiment.At the same time,to further improve the performance of the network model,the channel attention module is embedded in the network to achieve the best performance on the CIFAR-10 dataset.The proposed network still outperforms some commonly used networks on the CIFAR-100 dataset.
【Key words】 convolution neural network; multi-scale features; group convolution; depthwise separable convolution;