← 返回论文列表 📄 下载原文 PDF  ISSCC 2026 · 31.4
ISSCC 2026Session 31 · AI ACCELERATORSOther22nm

VARSA: A Visual Autoregressive Generation Accelerator Using Performance-Scalable Multi-Precision PE-LUT and Grid-Similarity Attention Compression

⚡ 本页包含 AI 生成的分析内容,仅供参考

📋 论文概要

本文提出VARSA,一种22nm工艺的视觉自回归生成加速器,用于高效文本到图像生成。通过性能可扩展的混合PE-LUT核心、多精度并行处理以及注意力图压缩技术,实现了503mJ/inference的能效,比现有扩散模型加速器提升2.7-8.9倍。

💡 主要创新点

核心指标
503mJ/inference for 512×512 image generation
工艺节点
22nm
重要性
发表年份
ISSCC 2026

🏷 关键词

视觉自回归生成加速器多精度PE-LUT网格相似性

📄 原文摘要

Abstract This paper presents VARSA, a 22nm visual autoregressive accelerator for efficient text-toimage generation, featuring: 1) a performance-scalable hybrid PE-LUT core; 2) multi-precision parallel processing with runtime precision management; 3) attention map compression leveraging inter-grid similarity. The innovations enable VARSA to achieve 503mJ/inference for 512×512 image generation, which is 2.7-to-8.9× better than prior diffusion-based SOTA accelerators. AI-generated content (AIGC) has become an essential technique to support various content generation, e.g., images, videos. In recent years, the model architecture of visual AIGC has rapidly evolved from the Diffusion Transformer (DiT) [1-3], to the autoregressive (AR) Transformer [4-6] to the visual autoregressive model (VAR) [7-10]. As shown in Fig. 31.4.1, DiT or AR models require either extensive denoising timesteps or inefficient serial rasterscan, leading to high inference latency. For example, generating a 512×512 imag

👥 作者与机构

Jiaqi Zhou, Hongou Li, Keyao Jiang, Yiyang Sun, Tianyu Jia

Peking University, Beijing, China

分类:Other · 年份:ISSCC 2026