⚡ 本页包含 AI 生成的分析内容,仅供参考
该论文提出了一款面向物联网设备的25mm² SoC,通过贝叶斯语音去噪和基于注意力的序列到序列学习,实现了18ms的噪声鲁棒语音到文本延迟。解决了之前ASR芯片在噪声和多说话人场景下性能差的问题。
Marco Donato2, Paul N. Whatmough1,3, Alexander M. Rush4, David Brooks1, Gu-Yeon Wei1 Harvard University, Cambridge, MA Tufts University, Medford, MA 3 ARM, Boston, MA 4 Cornell University, New York, NY 1 2 Automatic speech recognition (ASR) using deep learning is essential for user interfaces on IoT devices. However, previously published ASR chips [4-7] do not consider realistic operating conditions, which are typically noisy and may include more than one speaker. Furthermore, several of these works have implemented only small-vocabulary tasks, such as keyword-spotting (KWS), where context-blind deep neural network (DNN) algorithms are adequate. However, for large-vocabulary tasks (e.g., >100k words), the more complex bidirectional RNNs with an attention mechanism [1] provide context learning in long sequences, which improve ASR accuracy by up to 62% on the 200kwords LibriSpeech dataset, compared to a simpler unidirectional RNN (Fig. 9.8.1).
Thierry Tambe1, En-Yu Yang1, Glenn G. Ko1, Yuji Chai1, Coleman Hooper1,