⚡ 本页包含 AI 生成的分析内容,仅供参考
该论文提出了一种在8nm旗舰移动SoC中集成的双核稀疏感知神经网络处理单元(NPU),采用蝶形结构集成1024个MAC单元,实现了11.5TOPS/W的高能效。通过稀疏性感知技术有效利用剪枝后的网络稀疏性,解决了移动平台上深度神经网络高性能与低功耗的矛盾。
widely applied for image and speech recognition. Response time, connectivity, privacy and security drive applications towards mobile platforms rather than cloud. For mobile systems-on-a-chip (SoCs), energyefficient neural processing units (NPU) have been studied for performing the convolutional layers (CLs) and fully-connected layers (FCLs) [2-5] in deep neural networks. Moreover, considering that neural networks are getting deeper, the NPU needs to integrate 1K or even more multiply/accumulate (MAC) units. For energy efficiency, compression of neural networks has been studied by pruning neural connections and quantizing weights and features with 8b or even lower fixed-point precision without accuracy loss [1]. A hardware accelerator exploited
Jinook Song1, Yunkyo Cho1, Jun-Seok Park1, Jun-Woo Jang2,
Sehwan Lee2, Joon-Ho Song2, Jae-Gon Lee1, Inyup Kang1 Samsung Electronics, Hwaseong, Korea Samsung Advanced Institute of Technology, Suwon, Korea 1 2 Deep learning has been