← 返回论文列表 📄 下载原文 PDF  ISSCC 2019 · 7.5
ISSCC 2019Session 7 · MACHINE LEARNINGAI / ML

A 65nm 0.39-to-140.3TOPS/W 1-to-12b Unified NeuralNetwork Processor Using Block-Circulant-Enabled Transpose-Domain Acceleration with 8.1× Higher TOPS/mm2 and 6T HBST-TRAM-Based 2D Data-Reuse Architecture

⚡ 本页包含 AI 生成的分析内容,仅供参考

📋 论文概要

本文提出一种65nm工艺的统一神经网络处理器,通过块循环矩阵使能的转置域加速技术,支持CNN、FC、RNN三种网络,实现0.39-140.3 TOPS/W的宽能效范围和1-12b精度可配置,解决了异构架构面积效率低和能效低的问题。

发表年份
ISSCC 2019

📄 原文摘要

Yung-Ning Tu2, Yi-Ju Chen2, Ao Ren3, Yanzhi Wang3, Meng-Fan Chang2, Xueqing Li1, Huazhong Yang1, Yongpan Liu1 Tsinghua University, Beijing, China National Tsing Hua University, Hsinchu, Taiwan 3 Northeastern University, Boston, MA 1 2 Energy-efficient neural-network (NN) processors have been proposed for batterypowered deep-learning applications, where convolutional (CNN), fully-connected (FC) and recurrent NNs (RNN) are three major workloads. To support all of them, previous solutions [1-3] use either area-inefficient heterogeneous architectures, including CNN and RNN cores, or an energy-inefficient reconfigurable architecture. A block-circulant algorithm [4] can unify CNN/FC/RNN workloads with transpose-domain acceleration, as shown in Fig. 7.5.1. Once NN weights are trained using the block-circulant pattern, all workloads are transformed into consistent matrix-vector multiplications (MVM), which can potentially achieve 8to-128× storage savings and a O(n2)-to-O(nlog(n)) computation compl

👥 作者与机构

Jinshan Yue1, Ruoyang Liu1, Wenyu Sun1, Zhe Yuan1, Zhibo Wang1,

分类:AI / ML · 年份:ISSCC 2019