← 返回 JSSC 论文列表JSSC 2024第1期Digital Circuits
C-DNN: An Energy-Efficient Complementary Deep-Neural-Network Processor With Heterogeneous CNN/SNN
提出一种结合CNN和SNN的互补DNN处理器,实现高效能推理与训练。
85.8 TOPS/W (CIFAR-10), 79.9 TOPS/W (CIFAR-100)
深度神经网络卷积神经网络脉冲神经网络能效优化ASIC
▸结合CNN和SNN的异构核心架构,支持互补推理与训练
▸集成CNN-SNN工作负载分配器和注意力模块,优化能耗
▸采用分布式L1缓存和FDWSG技术,减少内存访问和训练操作
Abstract
In this article, we propose a complementary deep- neural-network (C-DNN) processor by combining convolutional neural network (CNN) and spiking neural network (SNN) to take advantage of them. The C-DNN processor can support both complementary inference and training with heterogeneous CNN and SNN core architecture. In addition, the C-DNN processor is the first DNN accelerator application-specific integrated circuit (ASIC) that can support CNN–SNN workload division by using their magnitude–energy tradeoff. The C-DNN processor integrates the CNN–SNN workload allocator and attention module to find a more energy-efficient network domain for each workload in DNN. They enable the C-DNN processor to operate at the energy optimal point. Moreover, the SNN processing element (PE) array with distributed L1 cache can reduce the redundant memory access for SNN processing, resulting in a 42.2%–49.1% reduction. For high energy-efficient DNN training, the C-DNN processor integrates the global counter and local delta-weight (LDW) unit to eliminate power-consuming counters for a forward delta-weight generation. Furthermore, the forward delta-weight-based sparsity generation (FDWSG) is proposed to reduce the number of operations for training by 31%–79%. The C-DNN processor achieves an energy efficiency of 85.8 and 79.9 TOPS/W for inference with CIFAR-10 and CIFAR-100, respectively (VGG-16). Moreover, the C-DNN processor achieves ImageNet classification with state-of-the-art energy efficiency of 2