← 返回论文列表 📄 下载原文 PDF  ISSCC 2022 · 29.3
ISSCC 2022Session 29 · ML CHIPS FOR EMERGING APPLICATIONSAI / ML28nm CMOS

A 28nm 15.59µJ/Token Full-Digital Bitline-Transpose CIM-Based Sparse Transformer Accelerator with Pipeline/Parallel Reconfigurable Modes

⚡ 本页包含 AI 生成的分析内容,仅供参考

📋 论文概要

该论文提出了一种基于全数字位线转置计算-in-memory(CIM)的稀疏Transformer加速器,采用28nm工艺,实现了15.59µJ/Token的高能效。针对Transformer注意力机制中的动态矩阵乘法和稀疏性问题,通过流水线/并行恢复技术减少了数据移动和计算开销。

💡 主要创新点

核心指标
15.59µJ/Token
工艺节点
28nm CMOS
重要性
发表年份
ISSCC 2022

🏷 关键词

Transformer加速器计算-in-memory (CIM)稀疏计算位线转置流水线并行

📄 原文摘要

state-of-the-art results in many fields, like natural language processing and computer vision, but their large number of matrix multiplications (MM) result in substantial data movement and computation, causing high latency and energy. In recent years, computing-in-memory (CIM) has been demonstrated as an efficient MM architecture, but a Transformer’s attention mechanism of raises new challenges for CIM in both memory access and computation aspects (Fig. 29.3.1): 1a) Unlike conventional static MM with pre-trained weights, the attention layers introduce dynamic MM (QKT, A’V), whose weights and inputs are both generated at runtime, leading to redundant off-chip memory access for intermediate data. 1b) A CIM pipeline architecture can mitigate the above problem, but produces a new challenge. Since the K

👥 作者与机构

Fengbin Tu1,2, Zihan Wu1, Yiqi Wang1, Ling Liang2, Liu Liu2, Yufei Ding2,

Leibo Liu1, Shaojun Wei1, Yuan Xie2, Shouyi Yin1 Tsinghua University, Beijing, China University of California, Santa Barbara, CA 1 2 Transformer models have achieved

分类:AI / ML · 年份:ISSCC 2022