⚡ 本页包含 AI 生成的分析内容,仅供参考
该论文提出了一种基于数字计算存储一体(DCIM)的流式多说话人语音识别(ASR)加速器,通过上下文感知冗余跳过和二维可写CIM技术,在22nm工艺上实现了1.87ms/帧的低延迟和高能效(37.50 TFLOPS/W),解决了流式多说话人ASR的实时性和能效挑战。
*Equally Credited Authors (ECAs) Abstract This paper presents a DCIM-based accelerator for on-device streaming multi-speaker ASR (MS-ASR), featuring: 1) a context-aware redundancy skipping scheme with online sparse block prediction, 2) a 2D-writable CIM with static pages and dynamic pages, 3) a similarity- aware TCAM for rapid search of similar speakers and N-best semantic results. The test chip achieves a system energy efficiency of 37.50TFLOPS/W in MXFP8 and shows superior performance on MS-ASR tasks with 1.87-7.51ms/frame and 0.158-1.26mJ/frame efficiency. The rising popularity of AI-assisted devices has sparked heightened interest in real-time automatic speech recognition (ASR). Prior single-pass ASR models, such as connectionist temporal classification (CTC) [1-2], recurrent neural-network transducer (RNN-T) [3-4] and attention-based encoder-decoder (AED) [5-6], exhibit high accuracy in terms of word error
Wenjie Ren*, Mingxuan Li*, Zhenghao Jin, Yifan Ding, Ruohuang Xu, Xiangjun Ye, Yufei Ma, Tianyu Jia, Le Ye
Peking University, Beijing, China