⚡ 本页包含 AI 生成的分析内容,仅供参考
该论文提出了一种基于张量训练和位级稀疏性优化的存内计算处理器,通过位级权重重用和变精度计算,实现了5.99到691.1 TOPS/W的能效范围,并采用计算宏聚合技术减少70.32%的内存访问。
position of M macros. The CMA reduces memory accesses by 70.32%, on average, in various NNs. Ruiqi Guo1, Zhiheng Yue1, Xin Si2, Te Hu1, Hao Li1, Limei Tang1, Yabing Wang1, Leibo Liu1, Meng-Fan Chang3, Qiang Li2, Shaojun Wei1, Shouyi Yin1 Figure 15.4.4 shows the SCCB in BSO-CIM, including 8 Wbw-CAs, 8 groups of double 5b SAs, an RWCU and a DRU. Each Wbw-CA consists of 16×64 6T cells and 64 local computing cells (LCCs). To optimize cell current (IC) with bit-level weight-sparsity, the local bit-line (LBL) reflecting the 1b weight (W1b) in the 6T cell controls the MN2 of an LCC. If W1b=0, MN2 is turned off and never induces IC. In each cycle, a 4b input is split into 2b input (INh2b, INl2b). The MN2-MN3 pair multiplies INh2b and W1b, and outputs the product to HGBL. The voltage at HGBL represents the Wbw-MACV, which is converted to digital by a 5b SA. Each weight is split by bit, and stored in the same row-column
balancing. In each macro, 16 VGBLs and 8 HGBLs are activated simultaneously. 2 TCs, are divided into tiles. Each tile has 16×8×M weights, mapped into the same row-column