← 返回论文列表 📄 下载原文 PDF  ISSCC 2026 · 31.7
ISSCC 2026Session 31 · AI ACCELERATORSOther28nm FD-SOI

LUT-SSM: A 99.3TFLOPS/W LUT-Based State-Space Model Accelerator Using Energy-Efficient Element-Wise Layer Fusion and LUT-Friendly Weight-Only Quantization

⚡ 本页包含 AI 生成的分析内容,仅供参考

📋 论文概要

提出一种基于查找表(LUT)的状态空间模型(SSM)加速器LUT-SSM,通过多对多LUT-ACC和逐元素层融合技术高效支持INT-FP的通用矩阵乘法和顺序逐元素运算,解决SSM部署中的存储访问和能效问题。在28nm FD-SOI工艺上实现99.3TFLOPS/W峰值能效。

💡 主要创新点

工艺节点
28nm FD-SOI
重要性
发表年份
ISSCC 2026

🏷 关键词

状态空间模型查找表加速器逐元素融合权重量化高能效

📄 原文摘要

Ulsan National Institute of Science and Technology, Ulsan, Korea, 4Naver Cloud, Seongnam, Korea *Equally Credited Authors (ECAs) 1 3 Abstract State-space models (SSMs) and weight-only quantization alleviate huge external memory access with a minimal accuracy degradation. To support both efficiently, we propose LUTSSM, a LUT-based SSM accelerator with a many-to-many LUT-ACC and an element-wise (EW) layer fusion. LUT-SSM efficiently supports INT-FP GEMM and sequential EW operations in the SSM blocks. Fabricated in 28nm FD-SOI, LUT-SSM achieves 99.3TFLOPS/W peak efficiency and 96.1TFLOPS/W on Mamba 1.4B, improving energy efficiency by 1.28× over recent designs. Recently, attention-based models [1] have been deployed in a variety of applications, e.g., AI agents, autonomous driving, robotics, and healthcare. However, they suffer from significant external memory access (EMA) for weights and intermediate caches, leading to

👥 作者与机构

Sunwoo Yoo*1,2, Dongyun Kam*3, Gunho Park4, Soonhyun Kwon1, Dongsoo Lee4, Youngjoo Lee1

Korea Advanced Institute of Science and Technology, Daejeon, Korea, 2Pohang University of Science and Technology, Pohang, Korea,

分类:Other · 年份:ISSCC 2026