← 返回论文列表 📄 下载原文 PDF  ISSCC 2023 · 16.4
ISSCC 2023Session 16 · EFFICIENT COMPUTE-IN-MEMORY BASED PROCESSORS FOR MLDigital Processors28nm CMOS

TensorCIM: A 28nm 3.7nJ/Gather and 8.3TFLOPS/W FP32 Digital-CIM Tensor Processor for MCM-CIM-Based Beyond-NN Acceleration

⚡ 本页包含 AI 生成的分析内容,仅供参考

📋 论文概要

本文提出TensorCIM,一款基于28nm工艺的数字CIM张量处理器,用于加速推荐模型等超越神经网络的稀疏聚集和稀疏代数操作,通过MCM-CIM架构有效缓解数据移动瓶颈。

💡 主要创新点

工艺节点
28nm CMOS
重要性
发表年份
ISSCC 2023

🏷 关键词

数字CIM张量处理器稀疏加速推荐模型MCM超越NN加速

📄 原文摘要

Recommendation Models (DLRMs) have computational and data-movement requirements beyond those seen in typical NN processing. Such beyond-NN applications typically consist of Sparse Gathering (SpG) and Sparse Algebra (SpA). SpG comprises gathering and reducing tensors from sparsely distributed addresses (in GCN’s aggregation phase and DLRM’s embedding layer). SpA refers to NN-based sparse tensor multiplication for the gathered tensors (in GCN’s combination phase and DLRM’s fullyconnected layer). Due to the large application size, data movement is the main bottleneck for beyond-NN acceleration. Digital Computing-In-Memory (CIM) is an efficient and precise architecture for reducing data movement [1-3]. Large-scale beyond-NN acceleration motivates the demand for scaling out digital CIM processors. However, a

👥 作者与机构

Fengbin Tu, Yiqi Wang, Zihan Wu, Weiwei Wu, Leibo Liu, Yang Hu,

Shaojun Wei, Shouyi Yin Tsinghua University, Beijing, China Applications such as Graph Convolutional Networks (GCNs) and Deep Learning

分类:Digital Processors · 年份:ISSCC 2023