⚡ 本页包含 AI 生成的分析内容,仅供参考
本文提出了一种多模式8K-MAC硬件利用率感知的神经处理单元,采用统一多精度数据路径,在4nm旗舰工艺上实现。该设计旨在满足实时应用中不同性能需求,包括高精度计算、多种深度学习层类型的高效处理以及极低功耗的始终在线运行。
Taeho Jeon1, Yesung Kang1, Heonsoo Lee1, Dongwoo Lee1, James Kim1, YoungJong Lee1, Sangkyu Park 1, Jun-Woo Jang2, SangHyuck Ha1, MinSeong Kim1, Jihoon Bang1, Suk Hwan Lim1, Inyup Kang1 Samsung Electronics, Hwaseong, Korea Samsung Advanced Institute of Technology, Suwon, Korea 1 2 Recent work on neural-network accelerators has focused on obtaining high performance in order to meet the needs of real-time applications with vastly different performance requirements, including high precision computation, efficiency for various Deep Learning (DL) layer types, and extremely low power to run always-on applications. Applying a single mode or datatype uniformly across these different scenarios would be less efficient than using different operating modes according to different operating scenarios. For example, super-resolution typically requires FP16 precision for higher image quality, while NNs for face-detection need only INT4 or INT8 precision. Using higher precision
Jun-Seok Park1, Changsoo Park1, Suknam Kwon1, Hyeong-Seok Kim1,