⚡ 本页包含 AI 生成的分析内容,仅供参考
该论文提出了一种采用28nm工艺的64kb SRAM计算内存宏,集成了数字域浮点计算单元和双位6T-SRAM单元,实现了31.6 TFLOPS/W的能效。它解决了传统SRAM-CIM只能高效支持整数精度MAC运算的局限,将计算能力扩展到浮点域,从而更适合检测、分割等复杂AI任务。
Yongliang Zhou, Lizheng Ren, Yeyang Xue, Xueshan Dong, Hui Gao, Yiran Zhang, Jingmin Zhang, Yuyao Kong, Tianzhu Xiong, Bo Wang, Hao Cai, Weiwei Shan, Jun Yang Southeast University, Nanjing, China SRAM-based computing-in-memory (SRAM-CIM) has been intensively studied and developed to improve the energy and area efficiency of AI devices. SRAM-CIMs have effectively implemented high integer (INT) precision multiply-and-accumulate (MAC) operations to improve the inference accuracy of various image classification tasks [13,5,6]. To realize more complex AI tasks, such as detection and segmentation, and to support on-chip training for better inference accuracy, floating-point MAC (FP-MAC) operations with high-energy efficiency are required. However, most SRAM-CIMs that previously used digital [5,6] or analog [1-4] in-memory computing cannot effectively support FP-MACs: e.g., Brain Float16 (BF16) datatype. Since supporting high floatingpoint input (IN), weight (W) and output (OUT) precision for
An Guo, Xin Si, Xi Chen, Fangyuan Dong, Xingyu Pu, Dongqi Li,