← 返回 JSSC 论文列表
📄 下载 JSSC 原文 PDF
JSSC 2019第6期Memory65nmNeural Network Accelerator

A 64-Tile 2.4-Mb In-Memory-Computing CNN Accelerator Employing Charge-Domain Compute

一款采用电荷域混合信号操作的64片24Mb内存计算CNN加速器,提升计算信噪比和可扩展性。
65nm CMOS, HLs/FL能效866/1.25 TOPS/W, 吞吐量18876/43.2 GOPS
内存计算CNN加速器电荷域操作混合信号DNN
电荷域混合信号操作提升计算信噪比和可扩展性
支持模拟/二进制输入激活/权重的第一层和二进制/二进制输入激活/权重的隐藏层
采用8T位单元与覆盖金属-氧化物-金属电容器的内存计算结构
Abstract
Large-scale matrix-vector multiplications, which dominate in deep neural networks (DNNs), are limited by data movement in modern VLSI technologies. This paper addresses data movement via an in-memory-computing accelerator that employs charged-domain mixed-signal operation for enhancing compute SNR and, thus, scalability. The architecture supports analog/binary input activation (IA)/weight first layer (FL) and binary/binary IA/weight hidden layers (HLs), with batch nor- malization and input–output (IO) (buffering) circuitry to enable cascading, if desired, for realizing different DNN layers. The architecture is arranged as 8 × 8 = 64 in-memory-computing neuron tiles, supporting up to 512, 3 × 3 × 512-input HL neurons and 64, 3 × 3 × 3-input FL neurons, configurable via tile-level clock gating. In-memory computing is achieved using an 8T bit cell with overlaying metal-oxide-metal (MOM) capacitor, yielding a structure having 1.8× the area of a standard 6T bit cell. Imple- mented in 65-nm CMOS, the design achieves HLs/FL energy effi- ciency of 866/1.25 TOPS/W and throug hput of 18876/43.2 GOPS (1498/3.43 GOPS/mm 2), when implementing convolution layers; and 658/0.95 TOPS/W, 9438/10.47 GOPS (749/0.83 GOPS/mm 2), when implementing convolution followed by batch normalization layers. Several large-scale neural networks are demonstrated, showing performance on standard benchmarks (MNIST, CIFAR- 10, and SVHN) equivalent to ideal digital computing.