⚡ 本页包含 AI 生成的分析内容,仅供参考
该论文提出了一种基于加权有限状态转换器(WFST)模型的低功耗实时语音识别硬件加速器,在6mW功耗下实现了5K词库的实时解码,解决了可穿戴设备等场景下云端语音识别功耗和带宽过高的问题。
Hardware-accelerated speech recognition is needed to supplement today’s cloud-based systems in power- and bandwidth-constrained scenarios such as wearable electronics. With efficient hardware speech decoders, client devices can seamlessly transition between cloud-based and local tasks depending on the availability of power and networking. Most previous efforts in hardware speech decoding [1–2] focused primarily on faster decoding rather than low-power devices operating at real-time speed. More recently, [3] demonstrated real-time decoding using 54mW and 82MB/s memory bandwidth, though their architectural optimizations are not easily generalized to the weighted finite-state transducer (WFST) models used by state-of-the-art software decoders. This paper presents a 6mW speech recognition ASIC that uses WFST search networks and performs end-to-end decoding from audio input to text output. Algorithms and data structures developed for software speech decoders are also
Michael Price, James Glass, Anantha P. Chandrakasan
Massachusetts Institute of Technology, Cambridge, MA