← 返回 JSSC 论文列表
📄 下载 JSSC 原文 PDF
JSSC 2020第4期Memory40nm

A 7.3 M Output Non-Zeros/J, 11.7 M Output Non-Zeros/GB Reconfigurable Sparse Matrix–Matrix Multiplication Accelerator Dong-Hyeon Park , Student Member , IEEE, Subhankar Pal , Student Member , IEEE, Siying Feng, Student Member , IEEE, Paul Gao, Student Member , IEEE,J i e l u nT a n ,Student Member , IEEE, Austin Rovinski , Student Member , IEEE

一款40nm CMOS工艺的可重构稀疏矩阵乘法加速器,具有48个异构核心和可重构内存层次结构。
40nm CMOS, 12.6×能效提升, 11.7×带宽效率提升, 17.1×计算密度提升
稀疏矩阵乘法异构计算可重构内存能效优化交叉开关
异构核心设计(Arm Cortex-M0和Cortex-M4)
可重构内存层次结构
合成可聚合交叉开关
Abstract
A sparse matrix–matrix multiplication (SpMM) accelerator with 48 heterogeneous cores and a reconfigurable memory hierarchy is fabricated in 40-nm CMOS. The compute fabric consists of dedicated floating-point multiplication units, and general-purpose Arm Cortex-M0 and Cortex-M4 cores. The on-chip memory reconfigures scratchpad or cache, depending on the phase of the algorithm. The memory and compute units are interconnected with synthesizable coalescing crossbars for efficient memory access. The 2.0-mm × 2.6-mm chip exhibits 12.6× (8.4×) energy efficiency gain, 11.7 × (77.6×) off-chip bandwidth efficiency gain, and 17.1 × (36.9×) compute density gains against a high-end CPU (GPU) across a diverse set of synthetic and real-world power-law graph-based sparse matrices.