⚡ 本页包含 AI 生成的分析内容,仅供参考
针对非稀疏神经网络训练需求,提出一款采用40nm工艺的8位浮点训练处理器,通过共享指数偏置技术实现4.81TFLOPS/W高能效,解决了现代非稀疏激活函数下传统稀疏加速方法失效的问题。
*Equally Credited Authors (ECAs) Recent works on mobile deep-learning processors have presented designs that exploit sparsity [2, 3], which is commonly found in various neural networks. However, due to the shift in the machine learning community towards using non-sparse activation functions such as Leaky ReLU or Swish for better training convergence, state-of-theart models no longer exhibit the sparsity found in conventional ReLU-based models (Fig. 9.3.1, top). Moreover, contrary to error-tolerant image classification tasks, more difficult tasks such as image super-resolution require higher precision than plain 8b integers not just for training, but for inference without large accuracy degradation (Fig. 9.3.1, bottom). These changes offer new challenges faced by mobile deep-learning processors: they must process non-sparse networks efficiently and maintain higher precision for more challenging tasks.
Jeongwoo Park*, Sunwoo Lee*, Dongsuk Jeon
Seoul National University, Seoul, Korea