⚡ 本页包含 AI 生成的分析内容,仅供参考
该论文提出了一款基于16nm FinFET工艺的变压器扩散模型处理器Tiamat,支持class-conditional DiT-XL/2和text-to-image PixArt-α,实现每步98ms和134ms的图像生成时间。通过分类器无关引导批处理、权重重排序量化和混合微尺度阻塞等创新,显著减少外部内存访问和特征图缓冲区大小,同时保持接近浮点的精度。
Abstract A 16nm FinFET transformer-based diffusion model processor chip is fabricated for supporting class-conditional DiT-XL/2 and text-to-image PixArt-α with 98ms and 134ms generation time per step with 7.37TOPS and 1477mW at 400MHz. Classifier-free guidance batching reduces external memory access (EMA) for weights by 50%, and weight-reordered quantization brings another 37% reduction. Hybrid microscaling blocking reduces featuremap buffer size by 37%, while approaching floating-point quality. The rapid growth of generative AI drives the demand for real-time, high-quality image generation on edge devices. Recently, transformer-based diffusion models (DMs) [1-2] have emerged as strong candidates, delivering superior image quality with 1.5-to-4.3× smaller model sizes compared to traditional U-Net-based models [3-4], as depicted in Fig. 2.7.1. These models replace the U-Net backbone with stacked transformer blocks and
Po-Yen Lu, Po-Wei Chen, Shau-Hua Yeh, Zhi-Jun Zhang, Yong-Tai Chen, Ting-Yu Wang, Chao-Tsung Huang
National Tsing Hua University, Hsinchu, Taiwan