⚡ 本页包含 AI 生成的分析内容,仅供参考
本文提出MulTCIM,一种基于28nm工艺的注意力-令牌-位混合稀疏数字存内计算加速器,用于高效处理多模态Transformer推理。通过利用多模态信号中token重要性的差异以及注意力稀疏性,实现了2.24µJ/Token的低能耗。
natural language, speech, etc. Multimodal Transformer (MulT, Fig. 16.1.1) models introduce a cross-modal attention mechanism to vanilla transformers to learn from different modalities, achieving excellent results on multimodal AI tasks like video question answering and multilingual image retrieval. Transformers require specialized hardware for efficient inference [1]. Prior work demonstrates that a Compute-In-Memory (CIM) accelerator with attention sparsity can efficiently process vanilla transformers [2]. Multimodal signals like video and audio exhibit diverse token significance, providing new opportunities for token sparsity via runtime pruning [3]. Additionally, activation functions like GELU and softmax produce many near-zero values that expose bit sparsity in the
Fengbin Tu, Zihan Wu, Yiqi Wang, Weiwei Wu, Leibo Liu, Yang Hu,
Shaojun Wei, Shouyi Yin Tsinghua University, Beijing, China Human perception is multimodal and able to comprehend a mixture of vision,