← 返回 JSSC 论文列表JSSC 2022第4期Memory40nmEmerging Memory
CHIMERA: A 0.92-TOPS, 2.2-TOPS/W Edge AI Accelerator With 2-MByte On-Chip Foundry Resistive RAM for Efficient Training and Inference Kartik Prabhu , Albert Gural, Student Member , IEEE, Zainab F. Khan
CHIMERA是一款基于40nm CMOS工艺的非易失性边缘AI加速器,支持训练和推理,采用RRAM技术。
40nm CMOS, 0.92-TOPS峰值性能, 2.2-TOPS/W能效
边缘AIRRAM非易失性内存DNN加速器低秩训练
▸首个使用RRAM宏的非易失性DNN芯片,支持边缘AI训练和推理
▸低秩训练算法显著减少RRAM权重更新步骤和能耗
▸ENDURER硬件模块提升RRAM的写入耐久性
Abstract
Implementing edge artificial intelligence (AI) infer- ence and training is challenging with current memory technolo- gies. As deep neural networks (DNNs) grow in size, this problem is only getting worse. This article presents CHIMERA, the first non-volatile DNN chip for both edge AI training and inference using foundry on-chip resistive RAM (RRAM) macros and no off-chip memory, fabricated in 40-nm CMOS. CHIMERA’s DNN accelerator is specifically optimized for RRAM and achieves 0.92-TOPS peak performance and 2.2-TOPS/W energy effi- ciency. We scale inference up to 6 × larger DNNs by con- necting six CHIMERAs in an illusion system with just 4% overhead in measured execution time and 5% in energy, enabled by communication-sparse DNN mappings that exploit RRAM non-volatility through quick chip wake-up and shutdown (<33 µs). Our incremental edge AI training algorithm, called low-rank training, overcomes RRAM write energy, speed, and endurance challenges and achieves the same accuracy as tradi- tional algorithms with up to 283 × fewer RRAM weight update steps and 340 × better energy-delay product. Combined with ENDUrance REsiliency using random Remapping (ENDURER), a hardware module that provides resilience to write endurance failures, we enable ten years of 20-samples/min incremental edge AI training. Manuscript received August 17, 2021; revised November 24, 2021; accepted December 24, 2021. Date of publication January 25, 2022; date of current ver- sion March 28, 2022. This article wa