⚡ 本页包含 AI 生成的分析内容,仅供参考
该论文提出了SoulMate,一款全设备端移动智能系统芯片,集成检索增强生成(RAG)和LLM微调功能,采用混合秩token处理与相似性感知序列处理架构,在28nm CMOS工艺下实现9.8mW低功耗实时交互,显著提升能效。
Abstract This work presents SoulMate, a fully on-device mobile intelligence system-on-chip, integrating retrieval-augmented generation (RAG) and fine tuning of a personal LLM. SoulMate is fabricated in 28nm CMOS with a novel mixed-rank token processing and similarity-aware sequence processing architecture. It demonstrates real-time user interaction consumining only 9.8-to-180.5mW power, and state-of-the-art energy efficiency, such as 26.3μJ/token for inference and 56.8μJ/token for fine-tuning. Mobile intelligence systems with large language models (LLMs) provide personalized conversational assistance tailored to each user’s characteristics [1-3]. They respond to user queries with awareness of individual preferences as well as factual knowledge, providing more personalized and relevant answers. However, the LLMs of existing systems require >10B parameters and >8GB RAM, while executing >1T operations per query, far exceeding
Seongyon Hong, Jiwon Choi, Jeonggyu So, Nayeong Lee, Wooyoung Jo, Zhamaliddin Kalzhan, Woojin Chin, Hoi-Jun Yoo
KAIST, Daejeon, Korea