⚡ 本页包含 AI 生成的分析内容,仅供参考
该论文提出了一种电容重构的存内计算(CIM)宏,能够统一加速CNN和Transformer两种不同精度需求的神经网络架构。通过重构电容阵列,实现了818-4094 TOPS/W的宽能效范围,适应不同计算精度要求。
In the rapidly evolving landscape of machine learning, workloads using diverse neuralnetwork architectures must be covered: including CNNs for image processing, transformers for natural language processing (NLP), and hybrid architectures that blend CNNs and transformers for audio processing. As illustrated in Fig. 34.5.1, these varied architectures have unique computational precision requirements. While CNNs achieve satisfactory accuracy even with low-computational precision, compute SNR or CSNR[1], transformers require higher CSNR to reach their full potential. This diversity amplifies the need for versatile hardware accelerators that can efficiently handle both CNNs and transformers, while meeting the multifaceted demands of modern machine-learning applications. Figure 34.5.3 shows our resource-efficient multi-bit row driver, which enables bit-parallel computation. In CNN mode, where lower CSNR is permissible, we exploit analog multibit inputs to facilitate bit-parallel computation to a
Keio University, Yokohama, Japan