← 返回 JSSC 论文列表JSSC 2021第1期Clocking & PLLs65nmNeural Network Accelerator
A Dynamic Timing Enhanced DNN Accelerator With Compute-Adaptive Elastic Clock
提出一种基于弹性时钟链的动态时序增强DNN加速器,通过多域时钟管理提升性能与能效
6×8 PE阵列,MNIST和CIFAR-10数据集上运行频率提升19%,能耗降低34%
DNN加速器弹性时钟链动态时序裕量多域时钟管理能效优化
▸弹性时钟链技术实现动态时序裕量利用
▸16个时钟域的多域时钟管理方案
▸基于运行时指令和操作数的动态时钟周期调整
Abstract
This article presents a deep neural network (DNN) accelerator using an adaptive clocking technique (i.e., elastic clock chain) to exploit the dynamic timing margin for the 2-D processing element (PE) array-based DNN accelerator. To address two major challenges on exploiting dynamic timing margin for modern deep learning accelerators (i.e., diminishing dynamic timing margin on a large array and strong timing dependence on runtime operands), in this work, we proposed an elastic clock chain scheme to provide a flexible multi-domain clock management scheme for in situ compute adaptability. More specifically, a total of 16 clock domains have been created for the 2-D PE array with the clock periods dynamically adjusted based on both runtime instructions and operands. The multi- domain clock sources are generated from a multi-phase delay- locked loop (DLL) and delivered by a global clock bus. The clock offsets between neighboring domains are deliberately man- aged to maintain the synchronization among clock domains. A1 6 × 8 PE array that supports different DNN dataflows and bit-precisions was fabricated using a 65-nm CMOS process. The measurement results on MNIST and CIFAR-10 data sets showed that the effective operating frequency was improved by up to 19% for a single instruction multiple data (SIMD) data flow by enabling the operation of the proposed elastic clock chain. The performance improvement was converted into up to 34% energy saving. Compared with SIMD data flow, the systolic