⚡ 本页包含 AI 生成的分析内容,仅供参考
本文提出了一款面向大语言模型(LLM)推理的5nm张量收缩处理器RNGD,通过结合高内存带宽与高密度计算,实现了显著的能效提升。该芯片旨在解决LLM推理中计算密集与功耗高昂的矛盾,提供低功耗的加速方案。
Byeongwook Bae1, Yojung Cha1, Wooyoung Choe1, Jonguk Choi1, Younggeun Choi1, Ki Jin Han2, Seokha Hwang1, Kiseok Jang1, Jaewoo Jeon1, Hyunmin Jeong1, Yeonsu Jung1, Hyewon Kim1, Sewon Kim1, Suhyung Kim1, Won Kim1, Yongseung Kim1, Youngsik Kim1, Hyukdong Kwon1, Jeong Ki Lee1, Juyun Lee1, Kyungjae Lee1, Seokho Lee1, Minwoo Noh1, Junyoung Park1, Jimin Seo1, June Paik1 FuriosaAI, Seoul, Korea Dongguk University, Seoul, Korea 1 2 There is a need for an AI accelerator optimized for large language models (LLMs) that combines high memory bandwidth and dense compute power while minimizing power consumption. Traditional architectures [1-4] typically map tensor contractions, which is the core computational task in machine learning models, onto matrix multiplication units. However, this approach often falls short in fully leveraging the parallelism and data locality inherent in tensor contractions. In this work, tensor contraction is used as a primitive instead
Sang Min Lee1, Hanjoon Kim1, Jeseung Yeon1, Minho Kim1, Changjae Park1,