⚡ 本页包含 AI 生成的分析内容,仅供参考
该论文实现了一款28nm工艺的视觉自回归(VAR)加速器,通过差分视觉注意力放大器推测关键令牌实现选择性执行,并采用全路径优化的MXINT处理单元适应偏态数据分布,以及链式反应式并行生成利用空间相关性,在47.3TFLOPs/W能效和小于0.6%的FID损失下取得37.6倍加速。
*Equally Credited Authors (ECAs) Abstract To accelerate Visual Autoregressive (VAR) applications, this work implements a 28nm VAR accelerator achieving 47.3TFLOPs/W and <0.6% FID loss. A differential visual attention amplifier speculates critical tokens for selective execution; a full-path optimized MXINT PE adapts to biased data distribution; and, a chain reaction-like parallel generation exploits spatial correlation. The 5.76mm2 chip runs at 400MHz, accelerating DeiT/ViT VAR by 37.6× with 2.75TFLOPs/mm2 area efficiency. In the field of generative AI, the Visual Autoregressive (VAR) model exhibits promising intelligence and has surpassed other generative models due to its scalability and versatility [1-4]. The former refers to its general performance scaling across models, and the latter points to the model’s adaptability to diverse downstream tasks. These benefits stem from its visual context learning capability, mirroring the large language model (LLM) [5-7].
Zhiheng Yue*, Xujiang Xiang*, Jiamu Fu, Shaojun Wei, Yang Wang, Yang Hu, Shouyi Yin
Tsinghua University, Beijing, China