← 返回 JSSC 论文列表JSSC 2019第1期Memory28nm
An Always-On 3.8 μJ/86% CIFAR-10 Mixed-Signal Binary CNN Processor With All Memory on Chip in 28-nm CMOS
一款超低功耗混合信号二进制CNN处理器,实现CIFAR-10图像分类任务3.8µJ/分类的能效
28nm CMOS, 0.6V/0.8V供电, 237FPS, 0.9mW功耗, 3.8µJ/分类, 86%准确率
混合信号处理器二进制CNN超低功耗开关电容神经元CIFAR-10分类
▸采用BinaryNet算法将权重和激活值约束为+1/-1,简化乘法运算为XNOR操作
▸权重静止数据并行架构结合输入复用技术,大幅降低内存访问能耗
▸创新的开关电容神经元设计,包含1024位温度计编码CDAC和9位二进制加权偏置模块
Abstract
The trend of pushing inference from cloud to edge due to concerns of latency, bandwidth, and privacy has created demand for energy-efficient neural network hardware. This paper presents a mixed-signal binary convolutional neural network (CNN) processor for always-on inference applications that achieves 3.8 µJ/classification at 86% accuracy on the CIFAR-10 image classification data set. The goal of this paper is to establish the minimum-energy point for the representative CIFAR-10 inference task, using the available design tradeoffs. The BinaryNet algorithm for training neural networks with weights and activations constrained to +1a n d −1 drastically simplifies multiplications to XNOR and allows integrating all memory on-chip. A weight-stationary, data-parallel architec- ture with input reuse amorti zes memory access across many computations, leaving wide vector summation as the remain- ing energy bottleneck. This design features an energy-efficient switched-capacitor (SC) neuron that addresses this challenge, employing a 1024-bit thermometer-coded capacitive digital-to- analog converter (CDAC) section for summing pointwise products of CNN filter weights and activations and a 9-bit binary- weighted section for adding the filter bias. The design occupies 6m m 2 in 28-nm CMOS, contains 328 kB of on-chip SRAM, operates at 237 frames/s (FPS), and consumes 0.9 mW from 0.6 V/0.8 V supplies. The corresponding energy per classification (3.8 µJ) amounts to a 40 × improvement over the previous