⚡ 本页包含 AI 生成的分析内容,仅供参考
本文提出Revolver,一种面向边缘设备的低比特GenAI加速器,通过相位感知精度选择、局部旋转和谐波对齐置换以及切片整数反量化单元,解决了多步推理和多轮对话中的内存和功耗瓶颈,实现了3.99倍能效提升和2.10倍加速。
Abstract Revolver is a low-bit GenAI accelerator that enables reasoning and multi-turn chat on edge devices under tight memory and power budgets. It introduces Phase-Aware Precision Selection (PAPS) with Multi-Precision Residual Encoding (MPRE) for memory-efficient multi-phase execution, Local Rotation with Harmonic-Aligned Permutation (LR-HAP) for low-cost rotation, and a Sliced Integer-based Dequantization Unit (SIDU) for efficient dequantization, achieving 3.99× energy savings and 2.10× speedup. Multi-step reasoning [1, 2] and multi-turn chatting [3] are becoming dominant edge-AI tasks, requiring larger models with higher computation and memory capacity than conventional single-turn AI assistants [4–6]. However, on-device deployment is almost impossible due to limited memory capacity (≤~12GB [7]) and computing capability. To realize these tasks under limited model size, two trends have emerged: (i) knowledge
Sangjin Kim, Jungjun Oh, Byeongcheol Kim, Yuseon Choi, Gwangtae Park, Hoi-Jun Yoo
KAIST, Daejeon, Korea