⚡ 本页包含 AI 生成的分析内容,仅供参考
本文提出了一种全数据通路(Full-Datapath)的SRAM浮点计算存内(CIM)宏,针对复合AI(Compound AI)系统实现高效边缘部署。该宏能效达51.6 TFLOPs/W,接近稀疏性理论界限,且精度损失小于2-30,解决了传统大模型在边缘设备上部署的成本和尺寸问题。
with exceptional performance, but their prohibitive size and cost limits deployment on edge devices. The compound-AI combines several specialized small models to achieve matched or even superior accuracy on target downstream tasks [1,2]: e.g., RevCol [3] fuses multiple convolutional models and surpasses the 7.2B-parameter monolithic LLM model [4] on the ImageNet classification task by 1.64%. Therefore, the shift to compound systems opens opportunities for edge deployment in addition to model-size scaling. SRAM-based floating point (FP) CIM promises accelerated edge-AI models [5-14]. Figure 14.4.1 shows a conventional FP CIM macro, which faces three challenges for compound-AI acceleration: (1) Previous FP CIMs rely on a Gaussian data distribution to reduce alignment loss, where data is distributed around its mean [11].
Zhiheng Yue*, Xujiang Xiang*, Yang Wang, Ruiqi Guo, Huiming Han,
Shaojun Wei, Yang Hu, Shouyi Yin Tsinghua University, Beijing, China *Equally Credited Authors (ECAs) Large-language models (LLM) are widely used