⚡ 本页包含 AI 生成的分析内容,仅供参考
该论文提出了一种在28nm工艺下实现的硬件利用率感知的神经网络加速器,通过动态数据流架构解决不同网络结构下MAC单元利用率下降的问题,实现了11.2TOPS/W的能效。
Research Institute, Hsinchu, Taiwan 1 With the rapid evolution of AI technology, various neural network structures have been developed for diverse applications. As a typical ease, Fig. 22.4.1 shows that the convolution (Conv) layer used in the convolutional neural networks (CNNs) features distinct shapes and types. Neural network accelerators with high peak energy efficiency have been demonstrated [1-4]. However, they usually suffer decreased hardware (mainly multiply-accumulate (MAC) units) utilization for various network structures, which reduces the attainable energy efficiency accordingly. To improve the MAC utilization, the Nvidia deep learning accelerator (NVDLA) [5] applies hardware parallelism along the channel direction, but the MAC utilization is still low for the shallow layers. According to
Cheng-Yan Du1, Chieh-Fu Tsai2, Wen-Ching Chen3, Liang-Yi Lin3,
Nian-Shyang Chang3, Chun-Pin Lin3, Chi-Shi Chen3, Chia-Hsiang Yang1 National Taiwan University, Taipei, Taiwan 2 Delta Electronics, Taipei, Taiwan 3 Taiwan Semiconductor