⚡ 本页包含 AI 生成的分析内容,仅供参考
提出了一种基于计算在存储边界(COMB)的神经网络处理器,通过双极位稀疏性优化实现高效计算,并支持可扩展多芯片架构,解决了传统CIM宏中权重数据访问的瓶颈问题。
Tianchan Guan2, Shengcheng Wang2, Dimin Niu2, Hongzhong Zheng2, Chixiao Chen1, Mingyu Wang1, Lihua Zhang1, Xiaoyang Zeng1, Qi Liu1, Yuan Xie2, Ming Liu1 Fudan University, Shanghai, China Alibaba DAMO Academy, Shanghai, China 1 2 *Equally Credited Authors (ECAs) Recently, computing-in-memory (CIM) macros, originally designed to reduce the intensive memory accesses of AI tasks, have been employed in low-power machine learning SoCs due to their ultra-high computing efficiency [1, 2, 3]. These CIM macros still access weight data through on/off-chip memories, similar to processing elements in near-memory-computing architectures. The implementation poses challenges when counting the overall SoC energy efficiency (Fig. 15.3.1). First, the memory wall issue is unsolved. The weight updates affect overall system performance when large networks are deployed and massive off-chip weight data transfer occurs. Even for tiny machine
Haozhe Zhu*1, Bo Jiao*1, Jinshan Zhang*1, Xinru Jia1, Yunzhengmao Wang1,