Initially applied for image synthesis [1], Diffusion Models (DMs) have been rapidly expanded into many content-generation tasks, e.g. 3D scenes [2-3] or video [4], and deliver exceptional performance. Figure 37.6.1 provides an overview of DM architecture, which typically processes random noisy input through multiple, i.e. 20-50, denoising steps to generate the desired output content. Each denoising step incorporates a U-Net structure with a down-sampling encoder and an up-sampling decoder, which contains repetitive transformer blocks. To support diverse content generation, multi-view [5] or temporal attention blocks [6] are integrated to enhance 3D scene or video-frame consistency. Due to the large number of denoising steps, generating a single piece of content consumes
Yiqi Jing1, Jiaqi Zhou1, Yiyang Sun1, Siyuan He1, Peiyu Chen2,3, Ru Huang1, Le Ye1,2, Tianyu Jia1
Peking University, Beijing, China Advanced Institute of Information Tecnhology of Peking University, Hangzhou, China 3 Nano Core Chip Electronic Technology, Hangzhou, China 1 2