⚡ 本页包含 AI 生成的分析内容,仅供参考
本文提出了一款基于12nm工艺的稀疏Transformer处理器,实现了18.1 TFLOPs/W的高能效。通过引入熵值提前退出机制、混合精度预测和细粒度电源管理,有效降低了大型语言模型推理的计算和能耗开销。
Thierry Tambe1, Jeff Zhang1, Coleman Hooper1, Tianyu Jia2, Paul N. Whatmough1,3, Joseph Zuckerman4, Maico Cassel Dos Santos4, Erik Jens Loscalzo4, Davide Giri4, Kenneth Shepard4, Luca Carloni4, Alexander Rush5, David Brooks1, Gu-Yeon Wei1 Harvard University, Cambridge, MA Peking University, Beijing, China 3 ARM, Boston, MA 4 Columbia University, New York, NY 5 Cornell University, New York, NY 1 2 Large language models have substantially advanced nuance and context understanding in natural language processing (NLP), further fueling the growth of intelligent conversational interfaces and virtual assistants. However, their hefty computational and memory demands make them potentially expensive to deploy on cloudless edge platforms with strict latency and energy requirements. For example, an inference pass using the state-of-the-art BERT-base model must serially traverse through 12 computationally intensive transformer layers, each layer containing 12 parallel attention
Entropy-Based Early Exit, Mixed-Precision Predication and, Fine-Grained Power Management