← 返回论文列表 📄 下载原文 PDF  ISSCC 2019 · 7.7
ISSCC 2019Session 7 · MACHINE LEARNINGAI / ML

LNPU: A 25.3TFLOPS/W Sparse Deep-Neural-Network Learning Processor with Fine-Grained Mixed Precision of FP8-FP16

⚡ 本页包含 AI 生成的分析内容,仅供参考

📋 论文概要

本文提出LNPU,一种支持片上稀疏深度神经网络学习的处理器,通过细粒度混合精度FP8-FP16实现高能效。解决了边缘设备上本地DNN学习能效低的问题,达到了25.3TFLOPS/W的性能。

💡 主要创新点

核心指标
25.3TFLOPS/W
重要性
发表年份
ISSCC 2019

🏷 关键词

稀疏神经网络混合精度学习处理器能效

📄 原文摘要

for energy-efficient deep learning (DL) acceleration [1-6]. Most prior DNN inference accelerators are trained in the cloud using public datasets; parameters are then downloaded to implement AI [1-5]. However, local DNN learning with domain-specific and private data is required meet various user preferences on edge or mobile devices. Since edge and mobile devices contain only limited computation capability with battery power, an energy-efficient DNN learning processor is necessary. Only [6] supported on-chip DNN learning, but it was not energy-efficient, as it did not utilize sparsity which represents 37%-61% of the inputs for various CNNs, such as VGG16, AlexNet and ResNet-18, as shown in Fig. 7.7.1. Although [3-5] utilized the sparsity, they only considered the inference phase with inter-channel accumulation in Fig. 7.7.1, and did not support intrachannel accumulation for the weight-gradient

👥 作者与机构

Jinsu Lee, Juhyoung Lee, Donghyeon Han, Jinmook Lee, Gwangtae Park, Hoi-Jun Yoo

KAIST, Daejeon, Korea Recently, deep neural network (DNN) hardware accelerators have been reported

分类:AI / ML · 年份:ISSCC 2019