← 返回 JSSC 论文列表JSSC 2024第1期Memory22nmEmerging MemoryNeural Network Accelerator
An 8b-Precision 8-Mb STT-MRAM Near-Memory-Compute Macro Using Weight-Feature and Input-Sparsity Aware Schemes for Energy-Efficient Edge AI Devices
提出一种基于STT-MRAM的8位精度近存计算宏,采用系统-电路协同设计解决能效和延迟问题。
22nm STT-MRAM, 436GB/s读取带宽, 20ns计算延迟, 53.6-190.2 TOPS/W能效
近存计算STT-MRAM能效优化人工智能边缘计算
▸权重特征感知读取方案(WFAR)
▸切换感知权重调谐方案(TAWT)
▸差分电荷累积增强电压敏感放大器(DCME-VSA)
Abstract
Nonvolatile near-memory-compute (nvNMC) macros are promising candidates for edge artificial intelligence (AI) devices requiring high energy efficiency, short wakeup- to-compute latency, and robust inference accuracy with high precision of inputs (IN), weights ( W ), and outputs (OUT). Nonetheless, the practical application of nvNMC macros is hindered by inherent design challenges: 1) high energy consumption in reading repetitious weight data, 2) low energy efficiency due to high bitstream toggling rate in digital multiply- and-accumulate (MAC) circuits, 3) narrow signal margin and high memory readout latency, and 4) redundant MAC computation in digital MAC circuits under input-stationary flow. In addressing these challenges, we developed a number of schemes based on the concept of system-circuit co-design, including 1) a weight-feature aware read (WFAR) scheme; 2) a toggling-aware weight-tuning (TA WT) scheme; 3) a differential charge-accumulating margin-enhanced voltage-sensing amplifier (DCME-VSA); and 4) an input-sparsity aware pre-computing unit (ISAPU). An 8-Mb spin-transfer torque magnetic random access memory (STT-MRAM) nvNMC macro fabricated using foundry-provided 22 nm STT-MRAM achieved read bandwidth of 436 GB/s, MAC computing latency of 20 ns, and energy efficiency of 53.6–190.2 TOPS/W when performing eight-bit input and eight-bit weight MAC operations with 26-bit output and 576 accumulations.