← 返回论文列表 📄 下载原文 PDF  ISSCC 2025 · 14.5
ISSCC 2025Session 14 · COMPUTE-IN-MEMORYAI / ML28nm

A 28nm 192.3TFLOPS/W Accurate/Approximate Dual-Mode-Transpose Digital 6T-SRAM CIM Macro for Floating-Point Edge Training and Inference connection using the 3rd-metal layer and connect the corresponding diagonal in the 4th layer; the row connection is in the 5th layer. This method circumvents the need for numerous MAC circuits and read ports as is the case for previous T-CIM works [7,8], resulting in a reduction in area and power consumption.

⚡ 本页包含 AI 生成的分析内容,仅供参考

📋 论文概要

本文提出了一种28nm工艺下的精确/近似双模式转置数字6T-SRAM存算一体宏单元,支持FP8、BF16、INT4和INT8等多种数据格式,实现了192.3 TFLOPS/W的高能效,旨在解决边缘训练中浮点运算的能效瓶颈。通过引入FP8新型浮点数据格式,相比BF16提升了训练效率。

💡 主要创新点

核心指标
192.3 TFLOPS/W
工艺节点
28nm
重要性
发表年份
ISSCC 2025

🏷 关键词

存算一体浮点训练边缘训练FP8高能效

📄 原文摘要

Shidong Lv3, Hao Wu1,2, Cailian Ma1,2, Ming Li1,2, Jinshan Yue1, Xinghua Wang3, Guozhong Xing1, Pui-In Mak4, Xiaoran Li3, Feng Zhang1 Figure 14.5.4 depicts the DCIM architecture supporting FP8, BF16, INT4, and INT8 formats. FP8 is proposed as a novel FP data format to achieve higher training efficiency than BF16

👥 作者与机构

Yiyang Yuan1,2, Bingxin Zhang1,2, Yiming Yang3, Yishan Luo1,2, Qirui Chen3,

分类:AI / ML · 年份:ISSCC 2025