⚡ 本页包含 AI 生成的分析内容,仅供参考
该论文提出了一种支持任意量化精度(1-8比特)的可扩展精度加速器,用于通用深度学习推理,能效达127.8 TOPS/W。它解决了不同网络层在稀疏性和精度要求上的差异问题,通过减少存储、逻辑和延迟浪费实现高效处理。
deep learning accelerators has focused on inference tasks to improve performance by means of maximally utilizing sparsity and quantization. Unlike CNN-only networks, however, recent state-of-the-art (SOTA) models consist of multiple blocks of various layers with different layer-by-layer characteristics in sparsity and required precision. This trend presents challenges in building a general accelerator architecture to maximize the benefits from sparsity and quantization, while supporting efficient processing for various models ranging from traditional CNNs to the new models to come in the future. First, there are multiple considerations that include the bottleneck in data bandwidth, as well as the trade-off between sparsity and required precision. The required
with Reduction of Storage, Logic and Latency Waste
Seunghyun Moon1, Han-Gyeol Mun1, Hyunwoo Son2, Jae-Yoon Sim1 Pohang University of Science and Technology, Pohang, Korea Gyeongsang National University, Jinju, Korea 1 2 Research on