← 返回 JSSC 论文列表JSSC 2024第3期Digital Circuits16nm
Amber: A 16-nm System-on-Chip With a Coarse- Grained Reconfigurable Array for Flexible Acceleration of Dense Linear Algebra
Amber是一款16nm SoC,集成粗粒度可重构阵列,用于加速密集线性代数应用。
538 INT16 GOPS/W, 483 BFloat16 GFLOPS/W
系统级芯片粗粒度可重构阵列密集线性代数动态部分重配置能效优化
▸动态部分重配置(DPR)技术提升资源利用率
▸支持仿射访问模式的流式内存控制器
▸低开销超越和复数算术运算
Abstract
Amber is a system-on-chip (SoC) with a coarse- grained reconfigurable array (CGRA) for acceleration of dense linear algebra applications, such as machine learning (ML), image processing, and computer vision. It is designed using an agile accelerator–compiler codesign flow; the compiler updates automatically with hardware changes, enabling con- tinuous application-level evaluation of the hardware–software system. To increase hardware utilization and minimize recon- figurability overhead, Amber features the following: 1) dynamic partial reconfiguration (DPR) of the CGRA for higher resource utilization by allowing fast switching between applications and partitioning resources between simultaneous applications; 2) streaming memory controllers supporting affine access patterns for efficient mapping of dense linear algebra; and 3) low- overhead transcendental and complex arithmetic operations. The physical design of Amber features a unique clock distribution method and timing methodology to efficiently layout its hier- archical and tile-based design. Amber achieves a peak energy efficiency of 538 INT16 GOPS/W and 483 BFloat16 GFLOPS/W. Compared with a CPU, a GPU, and a field-programmable gate array (FPGA), Amber has up to 3902×, 152×, and 107× better energy-delay product (EDP), respectively.