← 返回 JSSC 论文列表
📄 下载 JSSC 原文 PDF
JSSC 2024第1期Memory22nmSRAM

A Floating-Point 6T SRAM In-Memory-Compute Macro Using Hybrid-Domain Structure for Advanced AI

提出一种混合域结构的浮点6T SRAM内存计算宏,实现高能效和高精度
72.14 TFLOPS/W (BF16输入/权重, FP32输出, 128次累加)
浮点计算内存计算混合域结构6T SRAM人工智能边缘设备
混合域宏结构实现指数和尾数在同一CIM宏内计算
时间域计算方案实现高能效指数计算
基于乘积指数的输入尾数对齐方案实现同列尾数累加
位值依赖的数字-模拟混合计算方案实现高精度尾数计算
Abstract
Advanced artificial intelligence edge devices are expected to support floating-point (FP) multiply and accumu- lation operations while ensuring high energy efficiency and high inference accuracy. This work presents an FP compute-in- memory (CIM) macro that exploits the advantages of computing in the time, digital, and analog-voltage domain for high energy efficiency and accuracy. This work employs: 1) a hybrid-domain macrostructure to enable the computation of both the exponent and mantissa within the same CIM macro; 2) a time-domain computing scheme for energy-efficient exponent computation; 3) a product-exponent-based input-mantissa alignment scheme to enable the accumulation of the product mantissa in the same column; and 4) a place-value-dependent digital–analog-hybrid computing scheme to enable energy-efficient mantissa computa- tions of sufficient accuracy. A 22-nm 832-kB FP-CIM macro fab- ricated using foundry-provided compact 6T-static random access memory (SRAM) cells achieved a high energy efficiency of 72.14 tera-floating-point operations per second (TFLOPS)/W while performing FP-multiply-and-accumulate (MAC) operations involving BF16-input, BF16-weight, FP32-output, and 128 accumulations.