节点文献
基于STT-MRAM的模拟域存算融合方法和关键电路设计
STT-MRAM-Based Analog Domain Computing-in-Memory Method and Key Circuit Design
【作者】 郭亚楠;
【导师】 蔡浩;
【作者基本信息】 东南大学 , 微电子学与固体电子学, 2023, 硕士
【摘要】 随着人工智能(Artificial Intelligence,AI)的快速发展和广泛应用,计算量和数据量呈指数级增长。然而,冯·诺依曼架构下的存储墙问题和大容量存储的读写功耗限制了计算速度和能效,无法满足物联网(Internet of Things,Io T)应用场景下的速度和功耗需求。自旋转移力矩磁性随机存取存储器(Spin Transfer Torque Magnetic Random Access Memory,STT-MRAM)因其在非易失性、存储密度、访存速度和功耗等方面的优势备受瞩目。同时,模拟域存算融合(Computing-In-Memory,CIM)具有高并行度和高能效等特点。基于STT-MRAM的模拟域CIM能够在满足算力需求的同时提高计算能效,从而打破存储墙的限制。本文基于STT-MRAM提出了适用于多比特神经网络的模拟域存算融合(Multi Bit Analog Domain CIM,MBAD-CIM)方法和宽电压高良率写入策略,以上方案分别应用于MBAD-CIM宏单元和一颗4Mb的MRAM存储芯片中。在MBAD-CIM中:提出了具有锁存功能的存储单元(Latching In Bit-cell,LIB),以提高隧穿磁阻比(Tunneling Magnetoresistance Ratio,TMR);采用反馈电流镜(Current Mirror with Feedback,CMF)和双脉宽调制(Dual Pulse Width Modulation,DPWM)输入方式,提高模拟域计算的线性度和速度;设计了3-bit/4-bit精度可调模数转换器(Analog to Digital Converter,ADC)以减小量化功耗;提出了正负分离的存算阵列并采用逐位存储的数据映射方式以提高计算并行度。在MBAD-CIM中引入锁存-计算-量化-累加的流水线设计方法,提高CIM的吞吐率。在宽电压写入策略中:采用双电压方案解决磁隧道结(Magnetic Tunnel Junction,MTJ)翻转电压与先进互补金属氧化物半导体(Complemetary Metal Oxide Semiconductor,CMOS)工艺的工作电压不兼容的问题;提出非对称写驱动电路(Asymmetric Write Drive Circuit,AWDC),以减少MTJ的击穿风险和数据的写入功耗。MBAD-CIM宏单元采用28 nm CMOS工艺和垂直磁各向异性(Perpendicular Magnetic Anisotropy,PMA)MTJ模型联合设计,4Mb的MRAM存储芯片采用40 nm CMOS工艺完成版图设计和后仿真。仿真结果显示:在MBAD-CIM中,LIB能够将存储单元中的阻值转换为晶体管的导通截止电阻,TMR等效提高36000倍;与传统电流镜相比,CMF的积分非线性(Integral Nonlinearity,INL)减少了67.9%;3-bit/4-bit精度可调ADC的采样周期为4ns,转换功耗为108 f J~135 f J;完成4-bit激活值1-bit权重4-bit输出(4b IN-1b W-4b OUT)的乘累加操作时,计算能效为8.71 TOPS/W~56.5 TOPS/W,与传统存算分离的MRAM相比,能效提升至6.6倍,计算延时减少了78.4%;在MNIST识别任务中,4b IN-4b W-8b OUT(有符号权重)操作的计算能效是11.1 TOPS/W,网络的推理准确率为92.8%。4Mb MRAM的版图面积为10.23 mm~2,单个存储单元的写入良率为99.99%,AWDC的写入功耗相比传统设计减少了28.9%。
【Abstract】 With the rapid development and wide application of artificial intelligence(AI),the volume of computation and data is growing exponentially.However,the memory wall problem under Von Neumann architecture and the read-write power consumption of mass storage limit the computing speed and energy efficiency,which cannot meet the speed and power requirements under the Internet of Things(Io T)application scenario.Spin transfer torque magnetic random access memory(STT-MRAM)has advantages in non-volatility,memory density,memory access speed,and power consumption.Computing-in-memory(CIM)in analog domain is characterized by high parallelism and high energy efficiency.The analog domain CIM based on STT-MRAM can improve the computing energy efficiency while meeting the computing power demands,thus breaking the limitations of the memory wall.Based on STT-MRAM,this thesis proposes a multi bit analog domain CIM(MBAD-CIM)method for neural networks and a wide voltage high yield write strategy.The above schemes are applied to MBAD-CIM macro and a 4Mb MRAM memory chip respectively.In MBAD-CIM:the design of latching in bit-cell(LIB)is proposed to improve tunneling magnetoresistance ratio(TMR);a current mirror with feedback(CMF)and a dual pulse width modulation(DPWM)input mode are designed to improve the linearity and speed of analog domain calculation;a 3-bit/4-bit precision adjustable ADC is designed to reduce quantization power consumption;a positive-negative separation CIM array and the data mapping mode of bit-by-bit storage are adopted to improve the parallelism of computing.The MBAD-CIM introduces the pipeline design of latch-compute-quantize-accumulate to improve the throughput of CIM.In the wide voltage write strategy:a dual-voltage scheme is used to solve the problem of incompatible between the magnetic tunnel junction(MTJ)flip voltage and the operating voltage of the advanced CMOS process;an asymmetric write drive circuit(AWDC)is proposed to reduce the breakdown risk of MTJ and the write power consumption of data.The MBAD-CIM macro is designed with 28 nm CMOS process and the perpendicular magnetic anisotropy(PMA)MTJ model,and the 4Mb MRAM is completed in 40 nm CMOS process.LIB can convert the magneticresistance into the on-off resistance of the transistor,and the TMR is magnified 36000 times.Compared with the traditional current mirror,the integral nonlinearity(INL)of the CMF is reduced by 67.9%.The sampling period of the 3-bit/4-bit precision adjustable ADC is 4ns,and the conversion power consumption is 108 f J~135 f J.When MBAD-CIM completes the operation of 4-bit activation,1-bit weight,and 4-bit output(4b IN-1b W-4b OUT),the energy efficiency is 8.71 TOPS/W~56.5 TOPS/W,which is 6.6 times higher than traditional MRAM array,and the calculation delay is reduced by 78.4%.In the MNIST recognition task,the calculated energy efficiency of the 4b IN-4b W-8b OUT(signed weight)operation is 11.1 TOPS/W,and the inference accuracy of the network is 92.8%.The layout area of the 4Mb MRAM is 10.23 mm~2,and the write yield of a single bit-cell is 99.99%.The write power consumption of AWDC is reduced by 28.9%compared to traditional designs.
【Key words】 STT-MRAM; computing-in-memory; analog domain calculation; linearity; tunneling magnetoresistance ratio;
- 【网络出版投稿人】 东南大学 【网络出版年期】2024年 12期
- 【分类号】TP18;TP333