← 返回论文列表 📄 下载原文 PDF  ISSCC 2026 · 31.6
ISSCC 2026Session 31 · AI ACCELERATORSAI / ML

Tri-Oracle: A 17.78μJ/Token Vision-Language Model Accelerator with Token-Attention-Weight Redundancy Prediction

⚡ 本页包含 AI 生成的分析内容,仅供参考

📋 论文概要

本文提出Tri-Oracle,一种视觉-语言模型(VLM)加速器,利用token、注意力和权重的冗余进行加速。通过Token合并单元合并68%冗余token,注意力头预测单元跳过44%注意力延迟,以及条带式近零值计算单元减少45%前馈网络延迟。芯片在28nm CMOS工艺下实现17.78μJ/token的能效和1.19-3.99TOPS/mm2的面积效率,无需微调。

💡 主要创新点

发表年份
ISSCC 2026

📄 原文摘要

Abstract This paper presents Tri-Oracle, a VLM accelerator exploiting token, attention, and weight redundancies. A Token Merging Unit (TMU) merges 68% of redundant tokens. An Attention Head Prediction Unit (AHPU) predicts streaming heads and skips them, reducing attention latency by 44% and balancing load. A Strip-wise SNZV Computation Unit (SSCU) predicts near-zero values, cutting FFN latency by 45% and EMA by 57%. Fabricated in 28nm CMOS, Tri-Oracle achieves 17.78µJ/token and 1.19 to 3.99TOPS/mm2 without fine tuning. Recently, the success of transformer-based large language models (LLMs) has expanded into multimodal domains, particularly vision-language models (VLMs) [1] that combine visual perception with natural language understanding. As shown in Fig. 31.6.1, VLMs are increasingly utilized in vision questioning and reasoning tasks by co-processing image and text tokens with an LLM. Its emerging applications include e-commerce and security

👥 作者与机构

Seungjae Yoo, Hangyeol Kim, Muyoung Son, Yi Chen, Suheon Jeong, Joo-Young Kim

KAIST, Daejeon, Korea

分类:AI / ML · 年份:ISSCC 2026