⚡ 本页包含 AI 生成的分析内容,仅供参考
本文提出了一种基于子空间旋转的双量化大语言模型加速器,通过并行哈达玛转置器降低片上旋转功耗和面积,并采用融合缩放激活单元和重排位片LUT计算,实现了每token 51.6μJ的低能耗,相比现有技术降低32.6%。
Abstract A 51.6μJ/token accelerator for rotation-based dual-quantized LLMs is presented. A subspace-rotation method with parallel Hadamard transposer reduces on-chip rotation power by 62.3% and area by 59.7%. A fused scale-activation unit lowers energy by 61.5% vs. the naive FP design. Rearranged bit-slice LUT computation achieves 2.28× better energy efficiency compared to a direct bit-parallel MAC implementation, while supporting flexible bit-width. The chip reduces per-token energy by 32.6% over SOTA under an equal accuracy constraint. Large language models (LLMs) have demonstrated outstanding performance in natural language processing (NLP) tasks, pushing the boundaries of artificial intelligence (AI) applications [1-4]. However, this performance is achieved by scaling parameter counts into the billions, which greatly limits their deployment on cloud and edge devices [5]. As shown in Fig. 31.3.1, recent advanced quantization algorithms, such as group-wise and rotationbased quantizat
Bo Liu, Zihan Zou, Xinming Yan, Xilong Kang, Xinyang Chen, Bo Hu, Jiaming Lin, Haoran Du, Jun Yang, Xin Si, Hao Cai
Southeast University, Nanjing, China