节点文献
面向深度学习的SoC架构设计与仿真
Design and simulation of a deep learning SoC architecture
【摘要】 互联网时代信息量的爆炸式增长、深度学习的普及使传统通用计算无法适应大规模、高并发的计算需求。异构计算能够为深度学习释放更强的计算能力,达到更高的性能要求,并可应用于更广阔的计算场景。针对深度学习算法,设计仿真了一款完整的异构计算SoC架构。首先,通过对常用深度学习算法,如GoogleNet、LSTM、SSD,进行计算特征分析,将其归纳为有限个共性算子类,并用图表及结构框图的形式进行展示,同时生成最小算子级别伪指令流。其次,根据提取的算法特征,进行面向深度学习的硬件加速AI IP核设计,构建异构计算SoC架构。最后,通过仿真建模平台进行实验验证,SoC系统的性能功耗比大于1.5TOPS/W,可通过GoogleNet算法对10路1 080p 30fps视频逐帧处理,且每帧端到端的处理时间不超过30ms。
【Abstract】 The explosive growth of information volume in the Internet era and the popularization of deep learning have made traditional general-purpose computing unable to meet large-scale,high-concurrency computing requirements.Heterogeneous computing can release greater computing power for deep learning,satisfy higher performance requirements,and be applied to a wider range of computing scenarios.We design and simulate a complete heterogeneous SoC architecture for deep learning.Firstly,we analyze the computational features of commonly used deep learning algorithms such as GoogleNet,VGG and SSD,and summarize them into a limited number of deep learning common operator classes which are displayed in charts and structure diagrams.At the same time,the pseudo instruction stream at the minimum operator level is generated.Then,based on extracted algorithm features,a hardware-accelerated AI IP core for deep learning is designed,and a heterogeneous computing SoC architecture is constructed.Finally,experimental verification on the simulation modeling platform shows that the performance to power ratio of the SoC system is greater than 1.5 TOPS/W.The 10-channel 1080 p 30 fps video can be processed frame by frame by the GoogleNet algorithm,and the end-to-end processing time of each frame is no more than 30 ms.
【Key words】 heterogeneous computing; deep learning; acceleration unit; simulation modeling;
- 【文献出处】 计算机工程与科学 ,Computer Engineering & Science , 编辑部邮箱 ,2019年01期
- 【分类号】TP18
- 【被引频次】1
- 【下载频次】233