节点文献
Tiny YOLO卷积神经网络加速器设计与优化
Design and Optimization of Tiny YOLO Convolutional Neural Network Accelerator
【作者】 刘杰;
【作者基本信息】 天津大学 , 集成电路工程, 2019, 硕士
【摘要】 近年来,卷积神经网络由于其较高的分类精度而广泛应用于图像分类、目标检测以及场景分割等计算机视觉任务。卷积神经网络的分类精度随着网络层数的增加而提高。然而,伴随着网络加深,网络规模变大,需要的计算量剧增,采用软件运行卷积神经网络算法将会是一项非常耗时的工作。各种硬件加速器应运而生,以提高卷积神经网络模型的计算性能并满足嵌入式设备对于实时性、低功耗的要求。其中,现场可编程门阵列由于其强大的并行计算能力、高能效比、高灵活性等特点成为了实现硬件加速器的理想平台。本文采用具有典型卷积神经网络结构的Tiny YOLO算法进行硬件实现,提出了一种基于图像细粒度分块策略的加速器架构。在硬件设计中,提出应用于Line Buffer结构的Padding硬件实现方案,避免了软件方案存在的时间冗余和空间冗余问题。为了进一步提升加速器性能,本文通过改变首层卷积计算模式提高计算并行度;通过优化Line Buffer结构提高数据传输效率;通过乒乓技术和全流水设计思想降低系统延时。实验结果表明,在150MHz时钟频率下,本设计实现的性能为270.16 GOP/s。相比于CPU,加速比为6倍;相比于GPU,性能功耗比是其9倍;与国内外一些基于FPGA的研究成果对比,本文实现的性能是其1.3~1.7倍。
【Abstract】 In recent years,convolutional neural networks have been widely adopted in computer vision tasks such as image classification,target detection and scene segmentation due to their high classification accuracy.The classification accuracy of convolutional neural networks increases as the number of network layers increases.However,as the network deepens,the network becomes larger,and the amount of computation required increases dramatically.It is a very time-consuming task to implement convolutional neural network algorithms using software.Various hardware accelerators have emerged to improve the computational performance of the CNN models and meet the real-time and low-power requirements of embedded devices.Among them,field programmable gate array has become an ideal platform for hardware accelerators due to its powerful parallel computing capability,high energy efficiency and high flexibility.In this paper,the Tiny YOLO algorithm with typical CNN network structure is implemented in hardware for acceleration.An accelerator architecture based on image fine-grained block strategy is proposed.In the hardware design,the padding scheme applied to the Line Buffer structure is proposed,which avoids the time redundancy and space redundancy problems of the software solution.In order to further improve the performance of the accelerator,this paper improves the computational parallelism by changing the calculating order of the first layer;improves the data transmission efficiency by optimizing the Line Buffer structure;and reduces the system latency through the ping-pong technique and the full pipeline design.The experimental results show that the performance achieved by this design is 270.16 GOP/s under 150 MHz working frequency.Compared to the CPU implementation,the speedup ratio is 6times;compared to the GPU implementation,the performance-to-power ratio is 9times;more importantly,the accelerator shows a 1.3x~1.7x speedup compared with the state-of-the-art technique based on FPGA.
【Key words】 Convolutional neural network; Hardware accelerator; Field programmable gate array;