⚡ 本页包含 AI 生成的分析内容,仅供参考
本文提出了一款28nm工艺的CNN-Transformer混合加速器,通过内存计算强度感知和混合注意力层融合技术,实现了0.22µJ/Token的高能效,解决了语义分割任务中计算与内存开销大的问题。
Luhong Liang2, Yitong Zhou2, Di Pang2, Man-To Yung2, Dong Zhang1,2, Xijie Huang1,2, Shih-Yang Liu1,2, Yongkun Wu1,2, Fengshi Tian1,2, Chi-Ying Tsui1,2, Fengbin Tu1,2, Kwang-Ting Cheng1,2 The Hong Kong University of Science and Technology, Hong Kong, China 2 AI Chip Center for Emerging Smart System, Hong Kong, China 1 *Equally Credited Authors (ECAs) Recently, hybrid models integrating a CNN and a Transformer (ConvFormer), shown in Fig. 23.2.1, have achieved significant advancements in semantic segmentation tasks [1-4], which are critical for autonomous driving and embodied intelligence. The CNN enhances the multi-scale feature extraction ability of the Transformer to achieve pixel-level classification, but the large token length (TL) demand of semantic segmentation (>16K TL) incurs significant computation and memory overheads. Prior NN accelerators [5-12] demonstrate that sparse computing and pruning can effectively reduce computation and
Pingcheng Dong*1,2, Yonghao Tan*1,2, Xuejiao Liu2, Peng Luo2, Yu Liu2,