← 返回 JSSC 论文列表
📄 下载 JSSC 原文 PDF
JSSC 2024第10期Memory28nmSRAMCIM

A 28-nm 36 Kb SRAM CIM Engine With 0.173 µm2 4T1T Cell and Self-Load-0 Weight Update for AI Inference and Training Applications Chenyang Zhao , Jinbei Fang, Xiaoli Huang, Deyang Chen, Zhiwang Guo, Jingwen Jiang , Jiawei Wang

28纳米36Kb SRAM存内计算引擎,采用0.173µm² 4T1T单元,支持自加载0权重更新。
28nm CMOS, 263.1/412.1 TOPS/W (FF/BP), 2.5/4.9 TOPS/mm² (FF/BP)
存内计算SRAM神经网络训练能效优化权重更新
4T1T SRAM单元实现非破坏性读取和最小面积
共享路径双模读取减少电路开销
IR-drop感知自适应钳位器提升电压精度
Abstract
Computing-in-memory (CIM) promises high energy efficiency (EE) and performance in accelerating the feed-forward (FF) and back-propagation (BP) processes of deep neural net- works (DNNs) with less data movement and high parallelism. However, challenges still lie in large memory cells, network mapping, and IR-drop variation to realize efficient CIM imple- mentation. In this work, a 28-nm 36 Kb static random-access memory (SRAM) CIM engine with nondestructive-read (NDR) cell and weight update energy saving is used for multiply- accumulate (MAC) acceleration in artificial intelligence (AI) inference and train applications. A 4T1T SRAM bit-cell is pro- posed with NDR and records the smallest cell size of 0.173 µm2. The power-on self-load-0 feature of the 4T1T cell saves the weight update energy and latency for writing 0. The shared- path dual-mode read (SPDMR) brings fewer circuit overheads to support both FF and BP paths. The bit-interleaving weight mapping (BIWM) speeds up the BP path without slowing FF. IR-drop-aware adaptive clampers (IRDAA-Cs) with hierarchical read word-lines (RWLs) and read bit-lines (RBLs) apply possibly accurate voltages on near/far cells. The engine achieves an EE of 263.1/412.1 TOPS/W, as well as an area efficiency (AE) of 2.5/4.9 TOPS/mm 2 for FF/BP process @1-bit weight/activation with 74.4%–78.3% reduction in weight update energy.