⚡ 本页包含 AI 生成的分析内容,仅供参考
本文在3nm工艺下实现了一款完全数字式存内计算宏,支持INT12×INT12乘法累加操作,通过并行MAC方案提升了吞吐量并降低了能耗,同时分析了数据模式依赖性。
Cheng-En Lee1, Xiaochen Peng2, Vineet Joshi3, Chao-Kai Chuang1, Shu-Huan Hsu1, Takeshi Hashizume4, Toshiaki Naganuma4, Chen-Hung Tien1, Yao-Yi Liu1, Yen-Chien Lai1, Chia-Fu Lee1, Tan-Li Chou1, Kerem Akarvardar2, Saman Adham3, Yih Wang1, Yu-Der Chih1, Yen-Huei Chen1, Hung-Jen Liao1, Tsung-Yung Jonathan Chang1 TSMC, Hsinchu, Taiwan TSMC, San Jose, CA 3 TSMC, Ottawa, Canada 4 TSMC, Yokohama, Japan 1 The parallel MAC scheme not only achieves a higher throughput but also a lower energy consumption. Figure 34.4.3 explains the data pattern dependency in real workloads and the data toggle rate difference between serial and parallel MAC operations. We analyzed the data pattern and toggle rate difference on AlexNet, ResNet-50, MobileNetV2 and Inception-v1 running inference on the ImageNet dataset. As shown on Fig. 34.4.3(left), the input data toggle rate for the MSBs is lower than that for the LSBs, for all CNNs we
Hidehiro Fujiwara1, Haruki Mori1, Wei-Chang Zhao1, Kinshuk Khare1,