⚡ 本页包含 AI 生成的分析内容,仅供参考
提出了一款基于异构存内计算(CIM)的加速器,利用去噪相似性优化扩散模型推理,在28nm工艺下实现74.34TFLOPS/W的能效,解决了量化激活导致图像质量下降及GPU推理延迟高功耗大的问题。
China 3 Shanghai AI Lab, Shanghai, China 1 2 Diffusion models (DMs) have emerged as a powerful category of generative models with record-breaking performance in image synthesis [1]. A noisy image created from pure Gaussian random variables needs to be denoised by iterative DMs to ensure generative quality. For DMs, quantizing activations to integers (INT) degrades image quality due to changes in activation distributions and the accumulation of quantization errors across iterations. A GPU (Nvidia A100) requires 2560ms and 250W to generate a 256×256 image through 50 iterations of a floating-point (FP) DM. Two adjacent denoised images bring similar visual effects, where the difference between pixels at the same position is very small. As a result, for two adjacent DMs, most input differences within the same
Ruiqi Guo1, Lei Wang1, Xiaofeng Chen1, Hao Sun1, Zhiheng Yue1, Yubin Qin1,
Huiming Han1, Yang Wang1, Fengbin Tu2, Shaojun Wei1, Yang Hu1, Shouyi Yin1,3 Tsinghua University, Beijing, China Hong Kong University of Science and Technology, Hong Kong,