⚡ 本页包含 AI 生成的分析内容,仅供参考
本文提出了一种基于广播对齐的浮点存内计算宏,支持非二进制补码MAC运算,在28nm工艺上实现了17.83至62.84 TFLOPS/W的能效,适用于高精度CNN推理与训练。
Yuhui Shi1, Yuchen Tang1, Jinwu Chen1, Zhican Zhang1, Zhichao Liu1, Bo Liu1, Weiwei Shan1, Xin Wang3, Hao Cai1, Wenwu Zhu3, Jun Yang1,2, Xin Si1 Southeast University, Nanjing, China National Center of Technology Innovation for EDA, Nanjing, China 3 Tsinghua University, Beijing, China 1 2 *Equally Credited Authors (ECAs) The rapid advancement of artificial-intelligence (AI) models has increased demand for high-precision and energy-efficient edge-AI chips. Floating-point (FP) support is essential for high-precision neural-network (NN) training and inference; yet FP incurs higher energy and area overhead due to complex FP multiplication and accumulation (MAC) operations. Digital compute-in-memory (DCIM) and floating-point CIM (FP-CIM) [1-10] have emerged as promising techniques to improve energy efficiency with higher accuracy. Previous FP-CIM implementations [1-7] achieved good performance through various alignment
Xing Wang*1,2, Tianhui Jiao*1, Yi Yang1, Shaochen Li1, Dongqi Li1, An Guo1,