节点文献
基于片上神经网络的硅基集成光子计算芯片
Silicon-Based Integrated Photonic Computing Chip for On-Chip Neural Networks
【作者】 黄莹;
【导师】 储涛;
【作者基本信息】 浙江大学 , 电子信息, 2023, 博士
【摘要】 随着大数据、云计算时代的发展,需要计算的数据量急速增长,人工智能(Artificial Intelligence,AI)应用日益广泛。传统的微电子处理器难以满足目前AI技术的巨大算力需求,而光信号因其具备大带宽、超高速、低功耗、低延时和高并行度等优势,使其成为解决AI计算中功耗和计算效率等问题的一种有效手段。同时硅基光电子芯片因其具有CMOS(Complementary Metal-Oxide-Semiconductor)工艺高度兼容性、高集成度、小尺寸及低成本等优势,被认为是实现超高速、低功耗光计算的理想集成芯片平台,在AI的高算力计算场景中具有重要的潜在应用。因此,本文将面向光计算中主要的光子神经网络硬件处理器实现方式展开研究,重点探究神经网络中的线性矩阵乘法计算架构和非线性激活单元的硅基集成芯片的实现方式,以及对系统应用进行验证,实现适用于片上神经网络的硅基集成光子计算芯片。本文首先对光计算、硅基光子神经网络的研究背景和研究现状等做了简要介绍,并结合传统的相干光线性矩阵计算架构对研究理论、仿真方法、制造工艺以及系统应用展开论述。而后重点针对高性能、高集成度的线性矩阵计算架构和非线性激活单元的片上集成实现方式进行了一系列的研究。本文的主要研究内容和创新点为:1.本文基于单臂调制马赫-曾德尔干涉仪(Mach-Zehnder Interferometer,MZI)单元提出了一种易扩展、单元级联数目少的类扩张双层网络(Similar to Dilated Double-layer Network,SDDLN)硅基集成非相干光线性矩阵计算架构,有效解决了相干光网络中的架构扩展性较差、片上传输损耗较大、驱动数目较多以及相位误差积累影响较大等问题。本研究进一步研制了一种可重构的8×8 MZI阵列芯片。在实验中,以5 GHz的频率分别实现了一个2×4的非负矩阵和一个2×2的实数矩阵实现的光学线性矩阵乘加计算,算力分别达到40 GMAC/s和20 GMAC/s。并将此芯片应用于MNIST(Modified National Institute of Standards and Technology)手写数字识别分类任务,分别获得了89.5%和85.5%的推理准确度。更进一步,基于非相干型SDDLN阵列和波分复用器件提出了一种矩阵-矩阵乘法加速系统。实验演示了一个2×2的计算系统,并在MNIST手写数字识别分类任务中最高取得了92%的推理准确率。该模型的算力达到40 GMAC/s,相较于同维度传统网络,算力提升了一倍。2.本文基于锗硅光电探测器与强度调制器件提出了一种低阈值、片上可编程以及输入光源独立的光控非线性激活单元结构(Optically Controlled Nonlinear Activation Function,ONAF),有效解决了多层级联的光子神经网络中片上光信号强度衰减以及能量较高的问题。本文设计了基于PIN(P-doped–Intrinsic–N-doped)型MZI电光开关和微环谐振器电光开关,成功实现了片上可编程的整流线性单元函数(Rectified Linear Unit,Re LU)和Sigmoid型非线性激活函数。这两种函数的激活阈值低至0.2 m W,较当时主流方案降低了93.67%,显著提高了网络的能量效率。同时,其在MNIST手写数字识别任务中均获得了约95%的推理准确率。3.本文采用基于推挽驱动的MZI单元,提出了Cross-Bar架构的线性矩阵计算网络,并与ONAF相结合成功研制出4×4通用型的硅基单片集成光子计算芯片。此设计克服了基于波分复用技术的微环权重库系统在信道数量上的限制,以及传统三角和矩形MZI阵列中的驱动单元数目过多问题。基于非相干光原理,该芯片能进行实数域的线性矩阵计算和非线性激活运算,其中驱动数目减少了50%、单元功耗减少了81.7%。系统仿真结果显示,在加载电压精度达到小数点后3位的情况下,线性矩阵计算结果的均方根误差为0.041%。实验表明,实数域的非线性激活单元实现的反Sigmoid的最低激活阈值为0.1m W。此外,为开发具有更高算力的线性矩阵计算方案,本文结合偏振波长多维(解)复用器件和偏振不敏感耦合器等器件,首次提出了一个集波分复用、偏振复用、时分复用及空分复用于一体的四维计算架构。该设计不仅增强了计算维度,还解决了传统片上光子线性矩阵计算误差较大的问题。经系统仿真,该4×2×10×2的计算架构在加载的相位精确到小数点后3位时,其均方根误差为0.079%,这验证了该系统架构在矩阵计算方面的高鲁棒性。
【Abstract】 In the last two decades,due to increasingly maturing development of big data and cloud computing,the applications based on artificial intelligence(AI)are emerging in many fields,concomitant with a rapid increase in the computational demands on data processing.Conventional microelectronic processors fall short of meeting the growing computational requirements of modern AI.Optical computing,with its intrinsic broad bandwidth,ultra-high speeds,and high parallelism,is paving a way to solve the current bottlenecks in AI computing such as energy consumption and calculating efficiency.Meanwhile,silicon photonic chips,known for its CMOS compatibility,cost-efficiency,and high integration,is considered as a promising solution for the high-performance computing needs in AI.Therefore,this dissertation aims to explore key methods for implementing photonic neural network(PNN)processors in the realm of optical computing.Specifically,we examined how silicon photonic chips are applied for both linear matrix multiplication architectures and nonlinear activation function units in the neural network.Furthermore,we validated the applicability of the system,focusing on the optimization of specific integrated silicon photonic chips for on-chip neural networks.This dissertation describes out efforts towards a comprehensive study on optical computing and silicon-based PNNs.It first offers a comprehensive overview of the research background as well as the current state-of-the-art advancements.Subsequently,it delves deeply into advanced onchip integration strategies for high-performance,high-density linear matrix computational frameworks and non-linear activation units.The key contributions and novel aspects are summarized as follows:1.We proposed an easily scalable silicon-based integrated architecture for incoherent light linear matrix computation with reduced unit cascading,utilizing the single-arm modulated MachZehnder Interferometer(MZI)units within a Similar to Dilated Double-layer Network(SDDLN).This architecture effectively tackled common challenges in coherent optical networks,such as limited scalability,high on-chip transmission loss,excessive drive requirements,and significant phase error accumulation.Furthermore,we developed a reconfigurable 8 × 8 MZI array,demonstrating its capability for high-throughput optical linear matrix-vector multiplications at 5 GHz.Specifically,the array achieved computational speeds of 40 GMAC/s for 2 × 4 non-negative value matrices and 20 GMAC/s for 2 × 2 real-value matrices.When applied to the MNIST handwritten digit classification task,these two matrices achieved inference accuracies of 89.5% and 85.5%,respectively.Moreover,we implemented an accelerated matrix-matrix multiplication system based on the aforementioned SDDLN array and several wavelength division multiplexers.Through empirical validation,our 2 × 2 optimized system achieved inference accuracy of 92% on the MNIST classification task and demonstrated a computational throughput of 40 GMAC/s,effectively doubling the performance of comparable conventional MZI array.2.We proposed a low-threshold,on-chip programmable Optically Controlled Nonlinear Activation Function(ONAF)architecture,based on germanium-silicon photodetectors and intensity modulators.The ONAF architecture successfully addressed the challenges of on-chip optical signal attenuation and high energy requirement in cascaded PNNs.Utilizing a P-doped–Intrinsic–N-doped(PIN)MZI type and micro-ring resonator type electro-optic switch,we implemented on-chip programmable Rectified Linear Unit(Re LU)and Sigmoid nonlinear activation functions.The activation thresholds of both functions are as low as 0.2 m W,representing a 93.67% reduction compared to conventional approaches and significantly enhancing energy efficiency in PNNs.Concurrently,both the implemented Re LU and Sigmoid functions achieved approximately 95%inference accuracy in the MNIST classification task.3.We proposed a linear matrix computation network with a Cross-Bar architecture,utilizing push-pull-driven MZI units.Integrated with the aforementioned ONAF architecture,we implemented a 4 × 4 universal silicon-based integrated photonic computation chip.This design effectively circumvented the channel-count constraints in micro-ring weight banks based on wavelength division multiplexing(WDM),as well as the excessive number of drive units on conventional triangular and rectangular MZI arrays.Based on the characteristic of incoherent light,the chip is capable of performing linear matrix computations and nonlinear activation operations in the realnumber domain.Notably,it reduced the number of driving units by 50% and per-unit power consumption by 81.7%.System simulation results revealed that the root mean square error for linear matrix calculations is 0.041% with a voltage loading accuracy extending to three decimal places.Concurrent experiments showed that the minimum activation threshold of the inverse-Sigmoid implemented in the real-number domain nonlinear activation unit is 0.1 m W.Furthermore,to develop a linear matrix computational scheme with enhanced computational capabilities,we presented a groundbreaking four-dimensional computational architecture that combines devices such as polarization-wavelength multidimensional(de)multiplexers and polarization-insensitive couplers.This framework synergistically integrated WDM,polarization multiplexing,time division multiplexing,and spatial division multiplexing.Such an integration not only expanded the computational dimensions but also reduced the error in conventional on-chip photonic linear matrix computations.Our simulations of a 4 × 2 × 10 × 2 system showed that when the loaded phase is accurate to three decimal places,the root mean square error is only 0.079% to ideal value,thereby validating its high robustness in matrix computations.
- 【网络出版投稿人】 浙江大学 【网络出版年期】2025年 04期
- 【分类号】TN40;TP183