⚡ 本页包含 AI 生成的分析内容,仅供参考
本文提出HuMoniX,一种面向文本到运动生成任务的专用处理器,通过利用迭代间输出稀疏性和帧间关节相似性,实现了57.3fps的帧率和12.8TFLOPS/W的能效,解决了传统运动生成速度慢、功耗高的问题。
media applications, such as film production and AR/VR. This process involves creating human joint movements and constructing detailed 3D meshes, like human skin, for each joint (see Fig. 23.10.1). It used to require hours or days of using motion capture suits or manually modeling each joint frame by frame. However, transformer-based diffusion models now enable the generation of joint movements within seconds from inputs such as text [1-2] or music [3]. These models start with random noise and iterate through denoising steps to produce the desired output while integrating provided inputs. Despite this advancement, two major challenges remain for real-time execution. The first challenge is the high computational cost caused by excessive kernel iterations in diffusion models, where each iteration involves multiple transformer blocks, and the total
Jaehoon Heo, Adiwena Putra, Sungwoong Yune, Jieon Yoon, Hangyeol Lee,
Jihoon Kim, Joo-Young Kim KAIST, Daejeon, Korea Recently, 3D human motion generation has become crucial for entertainment and