⚡ 本页包含 AI 生成的分析内容,仅供参考
该论文提出了一种基于全数字位线转置计算-in-memory(CIM)的稀疏Transformer加速器,采用28nm工艺,实现了15.59µJ/Token的高能效。针对Transformer注意力机制中的动态矩阵乘法和稀疏性问题,通过流水线/并行恢复技术减少了数据移动和计算开销。
state-of-the-art results in many fields, like natural language processing and computer vision, but their large number of matrix multiplications (MM) result in substantial data movement and computation, causing high latency and energy. In recent years, computing-in-memory (CIM) has been demonstrated as an efficient MM architecture, but a Transformer’s attention mechanism of raises new challenges for CIM in both memory access and computation aspects (Fig. 29.3.1): 1a) Unlike conventional static MM with pre-trained weights, the attention layers introduce dynamic MM (QKT, A’V), whose weights and inputs are both generated at runtime, leading to redundant off-chip memory access for intermediate data. 1b) A CIM pipeline architecture can mitigate the above problem, but produces a new challenge. Since the K
Fengbin Tu1,2, Zihan Wu1, Yiqi Wang1, Ling Liang2, Liu Liu2, Yufei Ding2,
Leibo Liu1, Shaojun Wei1, Yuan Xie2, Shouyi Yin1 Tsinghua University, Beijing, China University of California, Santa Barbara, CA 1 2 Transformer models have achieved