← 返回 JSSC 论文列表
📄 下载 JSSC 原文 PDF
JSSC 2021第4期Memory10nmNeural Network Accelerator

A 617-TOPS/W All-Digital Binary Neural Network Accelerator in 10-nm FinFET CMOS Phil C. Knag , Member , IEEE, Gregory K. Chen , Member , IEEE, H.

一款采用10nm FinFET CMOS工艺的全数字二进制神经网络加速器,峰值能效达617 TOPS/W。
10nm FinFET CMOS, 峰值能效617 TOPS/W, 计算密度418 TOPS/mm², 内存密度414 KB/mm²
二进制神经网络全数字加速器计算近内存近阈值电压能效优化
采用计算近内存(CNM)架构减少互连和数据移动开销
利用近阈值电压(NTV)轻量级流水线降低时序元件开销
通过优化的数据访问模式减少卷积和池化操作的内存访问
Abstract
A binary neural network (BNN) chip explores the limits of energy efficiency and computational density for an all-digital deep neural network (DNN) inference accelerator. The chip intersperses data storage and computation using compu- tation near memory (CNM) to reduce interconnect and data movement costs. It performs wide inner product operations to leverage parallelism inherent in DNN computations. The BNN chip leverages lightweight pipelining at a near-threshold voltage (NTV) to reduce the overhead of sequential elements. It employs optimized data access patterns to reduce memory accesses for convolutional operation with pooling layers. The combination of these techniques enables the BNN chip to achieve a peak energy efficiency of 617 TOPS/W. The digital BNN chip approaches the energy efficiency of analog in-memory techniques while also ensuring deterministic, scalable, and bit-accuracy operation. Moreover, the all-digital design leverages process scaling and does not require additional memory transistors or passive devices to attain a peak compute density of 418 TOPS/mm 2 and a memory density of 414 KB/mm 2. The binary design is extended to enable bit-serial integer precision operation with a reconfigurable 1-b multiplication circuit and element-wise partial sum shift and accumulate. This technique allows for fine-grain mixed precision and retains energy efficiency by exploiting parallelism inherent in DNNs. The bit-serial binary operation allows for bit-accurate operation and h